Repository navigation
VALIDATION: independent review of the 2026-08-06 migration work #45
Description
Activity
Independent validation — results
Ran the full checklist without
scripts/build_audit.pyexcept where a check explicitly asks for it: fresh HTTP fetches, independent hashing, a cell-by-cell workbook diff with openpyxl, re-executing the lectures' data cells under pandas 3.0.5, re-reading every manifestread_asrecipe against the committed bytes, and a headless-Chromium CORS test from the wasm site's own origin. Validation scripts were written from scratch for this pass.Verdict: the work holds. Every byte-identity, provenance, manifest, tracker, and blind-spot claim verified — several came out stronger than claimed. One reader-facing regression was found that the checklist's own checks would not have caught (the wasm site cannot fetch the repointed URLs from the browser — CORS, filed as #46), one claim needs a caveat updated (
data.quantecon.orgnow fronts a third-party TLS cert before its 503), and two checks could not be completed as specified (Wayback figure compare — archive.org was down all session; an in-browser Colab session — verified by proxy instead).The one real finding — wasm #52/#53 broke in-browser data loading (#46)
lecture-wasmexecutes cells in the reader's browser (JupyterLite/Pyodide viapyodide_http.patch_all()— XHR, CORS-enforced per hop). The repoints changed the wasm URL form from directraw.githubusercontent.com(servesaccess-control-allow-origin: *) togithub.com/…/raw/(a 302 whose response carries an empty ACAO header, failing the CORS check before the redirect is followed). Confirmed empirically with headless Chromium fetching fromquantecon.github.io/lecture-wasm/long-run-growth/: both repointed URLs fail withFailed to fetch; the pre-repoint form returns 200 with the exact 1,765,204 bytes. The pre-repoint wasm URLs were a deliberate wasm-specific adaptation that the repoint normalised away. Since the wasm build bakes no outputs, figures on those two wasm pages exist only after in-browser execution, which now dies at the first data cell.inequality.mdandmle.mdhave the same failure pre-existing (viahigh_dim_datagithub.com/raw URLs);heavy_tails.mdis CORS-clean (media.githubusercontent.com). The repoints were correctly sequenced — the old URLs point at files #825 has since deleted — only the form is wrong, and the fix is a two-lecture URL flip plus a repoint rule. Details, evidence table, and the Phase 4 implication (data.quantecon.org must serve ACAO*before wasm can cut over) in #46.So the claim "no reader-visible output changed" is true for the intro site, downloaded notebooks, and Colab — and false for wasm interactive execution.
1. Reader-facing outcomes
Check Result intro long_run_growth.html/inflation_history.html✅ both 200; figure _images/*.pngspot-checks 200Same two pages on the wasm site ⚠️ pages 200 at/long-run-growth/,/inflation-history/, and their sources point at data-lectures — but in-browser execution broken, see #46Download each notebook and run it ✅ fetched intro.quantecon.org/_notebooks/*.ipynb; both reference only data-lectures URLs; executed the data cells live under pandas 3.0.5 — Full data (21682, 5), Regional (23, 18), longprices (401, 4) 1600–2000, all 11 chapter_3 sheets readOpen each in Colab ✅ by proxy: fetched the exact bytes Colab loads ( lecture-python-intro.notebooks@main) — data-lectures URLs only — and executed those fetches successfully; no live Colab session run{download}chapter_3.xlsx link✅ 200, 73,281 bytes, opens as xlsx (all 11 sheets parsed) "hosted on GitHub" link ✅ → data-lectures/blob/main/CATALOG.md, 200Figure spot-check vs Wayback ⚠️ archive.org CDX was 503/429 for page snapshots all session. Substitute proof: the repoint commits' lecture diffs are pure URL-string swaps (git show 3ccd49c8,69307fff— no code-logic change) and the data bytes are identical (§2), so the figures are determined-identical. The_images/<content-hash>.pngfilename comparison against a pre-migration snapshot remains worth doing when IA recoversSet 1 404 window closed ✅ published notebooks now fetch data-lectures URLs, all 200 2. Data integrity — byte-identity
Three-way match on all three files: freshly fetched
raw/mainbytes = sidecarintegrity.sha256= pre-deletion blob in intro's history (3ccd49c8^for mpd2020,3f715b84^for the set 2 pair).File sha256 (all three sources identical) mpd2020.xlsxf67af0fd…e3bceblongprices.xls1160f9d3…b344f7chapter_3.xlsxee17d003…d30923The "not simply copied" concern is answered by construction: the git blob and the live server were hashed independently of the manifests and agree with them. mpd2020's add date also checks out —
e23ae056, 2023-03-23, PR #120, blob never modified since, exactly as the manifest says.3. mpd2020 provenance — reproduced exactly, and stronger
- Cell-by-cell diff (openpyxl, all six sheets, 379,701 cells): exactly 3 differ, all on
Regional datarow 1 —GDP pc 2011 prices→gdppc_2011,Population→pop, andgdppc_2011added over the unlabeled world column. All 21,683 rows ofFull dataidentical. Matches the manifest'sdeltarecord verbatim. - IA digest
4OWZWHTE5HXBBF4XLCCQTGDXY4HXOGOKis constant across all 10 snapshots, 2021-01-10 → 2026-04-01 — beyond the claimed 2026-01-02 — bracketing our 2023-03-23 commit. Better: the base32-SHA-1 of today's fresh rug.nl download equals that digest exactly, so upstream has served one byte-identical file for 5½ years and the edits are unambiguously ours. - Consequence check: the lecture's
header=(0,1,2)read against the pristine upstream raisesKeyError: 'gdppc_2011'— the local edits are load-bearing, soconstructedis the right class and a "clean upstream refresh" would break the lecture outright, not subtly change it.
4. Manifests
- All 18 sidecars parse; every
filenamematches its sidecar's name — no silently-invisible manifest. - Classifications match the claimed set: mpd2020/chapter_3/caron/nom_balances
constructed, longprices/assignat/dette/fig_3verbatim; both.npycarrybuilder_status: unrecovered. .npyre-test, done independently: scanned every cell of every sheet of the three workbooks. caron.npy: neither terminal value (96.696, 0.431) appears anywhere, no contiguous column matches — claim confirmed. nom_balances.npy: its terminal values 90.0 and 33,555.59 do occur in assignat.xlsx (Fig6, Data, Data2, Denomina, Dom-nat, Ramel, Ramel2, seignor) — and the manifest already records precisely that partial overlap, notes no contiguous 81-value column exists (confirmed) and that the series max 37,540.933 appears nowhere (confirmed). The manifest is accurate to the decimal; "overlaps without being extractable" is the right description.- Every
read_asrecipe re-executed against the committed bytes reproduces the recorded shape and the recordedknown_nulls_total— all 10 ranges across assignat (4), dette (5), fig_3 (1), plus both npy shape/dtype blocks and the mpd2020/longprices sheet reads (via the lecture's full recipe incl.index_col/iloc). positional_reads: truepresent on exactly the three French Revolution workbooks and no others.
5. Tracker consistency
migration.yml: 18 records, 13repointed/ 5landed, and the 5 are exactly assignat, dette, fig_3, caron.npy, nom_balances.npy. Everyrepointedrecord names its PRs; spot-checked that intro#823/#824 and wasm#52/PLAN: correct the stale figures, widen rule 1 to the org, restate rule 6 #53 really contain the corresponding repoints (they do — subjects and diffs match).build_catalog.py→ zero diff, CATALOG.md current.build_audit.py all --stricton today'smain(i.e. post-#825): exit 0.
6. The two blind spots
- The audit cannot see prose references — 10 {download}/link refs to files the migration will delete #42: the audit's refs for the french_rev files contain only the code reads — the
blob/prose links atfrench_rev.md:60-62appear nowhere inaudit.json(the two "prose" strings in it are annotation notes), so the blind spot is genuinely real, not just asserted. Independent re-sweep by a different method (all{download}directives + allgithub.com/QuantEcon/lecture-*/(raw|blob|tree)URLs in both repos' sources, classified against code cells) reproduces the ledger exactly: 12 = 2 chapter_3 downloads (fixed → data-lectures) + 2 directory links (fixed → CATALOG.md) + 6 french_rev blob links still outstanding in both repos at lines 60–62 + 2 life-expectancy downloads (wave 4), plus the 2data.ipynblinks tracked separately in the The audit cannot see prose references — 10 {download}/link refs to files the migration will delete #42 comment. Nothing beyond the documented set was found. - Repoint rule 3: verified from workflow files — exactly 7 of 9 repos publish on a
publish*tag push (continuous_time_mcs, dp, jax, advanced.myst, intro, programming, python.myst); wasm and data-lectures publish on push-to-main. The lag is real: lecture-dp's newest tag ispublish-2026jul07against a jul20 main head — 29–30 days stale at the time the claim was made. The twopublish-2026aug06/…btags exist on intro. HTML-survives/notebook-breaks confirmed mechanically: published HTML embeds content-hashed_images/*.png; published notebooks carry runtime fetch code — which is also what Repoint rule for wasm: in-browser reads need raw.githubusercontent.com — the github.com/*/raw/ form fails CORS, breaking long_run_growth and inflation_history on the wasm site #46 exploits in the other direction on wasm.
7. Decisions
Claim Result SCF minis fit plain git ✅ from the LFS pointers: 32,853,734 B = 31.33 MiB and 75,902,999 B = 72.39 MiB, both under the 104,857,600 B hard block (the larger one is past the 50 MiB warning threshold, worth knowing) SCF_plus.dta99.1 MiB, no lecture consumer✅ 103,934,093 B = 99.12 MiB, 0.88 MiB under the block; grep across all nine workspace repos finds no lecture read — only PLAN.md mentions it generating_mini.mdreads over the network✅ line 37: pd.read_stata('https://github.com/QuantEcon/high_dim_data/blob/main/SCF_plus/SCF_plus.dta?raw=true')pandas 3 text dtype is str✅ pandas 3.0.5: str(dtype)→str(repr isStringDtype(storage='python', na_value=nan)) — notstring, notobjectdata.quantecon.org= empty AWS LB✅ with a caveat to record: behind TLS it is awselb/2.0returning 503 (http → 301 → https), and rDNS of 52.64.86.66 still saysdata.quantecon.org— but the LB now presents a certificate for*.dev.cloud.payok.com.au, so a browser hits a cert error before any 503. Not a live service either way, but the third-party cert is new relative to the recorded finding and worth a line on #37 — if that LB is no longer QuantEcon-controlled, the DNS record is dangling on someone else's infrastructure8. Deliberately not done — all confirmed as decisions
intro#825 was open when this issue was filed and has since merged (2026-08-06 08:06 UTC) — the gate (publish first) was honoured, and §2/§5 above were run against the post-#825 tree. #35 (licence policy deferral), #13 (business_cycle dumps, recommendation posted, no decision), and ws#23 (step 6 write-backs) are open and say what the checklist says they say. Schema drift is real and measurable:
doi(6 manifests),file_url(7),sheets(8),positional_reads(3) all appear in manifests and nowhere inmanifest-schema.yml. The delta register #39 exists with mpd2020 as its sole entry, correctly distinguishingdivergedfrom the elevenunverifiables. #20 closed, #37/#40 open as stated.Still open after this pass
- Repoint rule for wasm: in-browser reads need raw.githubusercontent.com — the github.com/*/raw/ form fails CORS, breaking long_run_growth and inflation_history on the wasm site #46 — the wasm CORS regression: two-lecture URL-form fix + repoint rule + Phase 4 ACAO acceptance criterion.
- Wayback
_imageshash comparison for one lecture, when archive.org recovers — the determinism argument above is solid, but the checklist wanted the published bytes compared and that half is still owed. - A true in-browser Colab run of each notebook (verified here at the bytes-Colab-loads + fetches-execute level only).
- The
data.quantecon.orgcert observation → Track Y (deferred): data.quantecon.org DNS + custom domain — retired in favor of qeld (D11); reopenable, NXDOMAIN as of 2026-08-10 #37.
Method note for the record: this validation was performed by a different agent from the one that did the migration work, with its own scripts (openpyxl cell diff, pandas re-reads of every
read_as, LFS-pointer size reads, headless-Chromium CORS harness) rather than the repo's audit tooling, per the bias warning at the top of this issue.- Cell-by-cell diff (openpyxl, all six sheets, 379,701 cells): exactly 3 differ, all on
Closing: the validation ran in full and every checklist item has a verdict in the report above — all boxes ticked accordingly.
Where the three caveats landed, so nothing rides on this issue staying open:
- wasm in-browser execution — the one real finding — is tracked in Repoint rule for wasm: in-browser reads need raw.githubusercontent.com — the github.com/*/raw/ form fails CORS, breaking long_run_growth and inflation_history on the wasm site #46, with the code fix in Use CORS-clean hosts for in-browser data reads lecture-wasm#54 and the rule in Add repoint rule 5: wasm reads need a CORS-clean URL form #47, both open and browser-verified. Since
lecture-wasmpublishes on push, the site self-heals when Audit: assert on lfs_media and on ref/path resolvability — strict exits 0 on a repoint that 404s every read #54 merges. - data.quantecon.org third-party-cert observation and the CORS acceptance criterion are recorded on Track Y (deferred): data.quantecon.org DNS + custom domain — retired in favor of qeld (D11); reopenable, NXDOMAIN as of 2026-08-10 #37.
- Accepted residuals, documented in the report and deliberately not re-tracked: the Wayback
_imageshash comparison (archive.org was down; the URL-only-diff + byte-identity argument covers it) and a true in-browser Colab session (verified at the exact-bytes-Colab-loads + fetches-execute level).
- wasm in-browser execution — the one real finding — is tracked in Repoint rule for wasm: in-browser reads need raw.githubusercontent.com — the github.com/*/raw/ form fails CORS, breaking long_run_growth and inflation_history on the wasm site #46, with the code fix in Use CORS-clean hosts for in-browser data reads lecture-wasm#54 and the rule in Add repoint rule 5: wasm reads need a CORS-clean URL form #47, both open and browser-verified. Since
- added a commit that references this issue
on Aug 6, 2026
Independent validation of the migration work landed on 2026-08-06. Everything below was verified during the work by the same agent that did it — this issue exists so it can be checked by someone who was not.
Bias to test for: most claims here were verified by running
scripts/build_audit.py, which is also the tool the work was built around. A validation pass should confirm reader-facing outcomes independently of that tool wherever possible — fetch the URLs, open the notebooks, read the published pages.What landed
data-lectureslecture-python-introlecture-wasmworkspace-lecturesIssues opened: #35, #37, #39, #40, #42 here; workspace-lectures#23. Closed: #20. Rewritten: #8, workspace-lectures#14.
Publishes:
publish-2026aug06(set 1),publish-2026aug06b(set 2).1. Reader-facing outcomes — check these first, without the audit tool
Two datasets moved. The claim is that no reader-visible output changed and nothing 404s.
intro.quantecon.org/long_run_growth.htmlandinflation_history.htmlrender, figures intactdata-lecturesand succeed{download}link forchapter_3.xlsxininflation_history— it should serve the spreadsheet, not 404inflation_history— should reach data-lectures'CATALOG.mdKnown live-site history worth confirming is now clean: between set 1 merging and
publish-2026aug06, the publishedlong_run_growthnotebook pointed at a deleted file and returned 404. That window is closed, but it is the single clearest thing to verify independently.2. Data integrity — the claim is byte-identity
Every repoint asserted the served bytes are identical to what the lecture used before.
mpd2020.xlsx,longprices.xls,chapter_3.xlsx: fetch fromgithub.com/QuantEcon/data-lectures/raw/main/lectures/<file>and comparesha256againstintegrity.sha256in its sidecar manifestlecture-python-intro's git history (git show <pre-merge-sha>:lectures/datasets/<file> | shasum -a 256)sha256values were not simply copied from the file they describe without an independent check3. The
mpd2020.xlsxprovenance finding — the highest-value claim to re-testClaim: the file is not a pristine Maddison release. All 21,682 data rows match upstream, but three header labels on the
Regional datasheet were edited locally, andlong_run_growthdepends on them viaheader=(0,1,2).rug.nl/ggdc/historicaldevelopment/maddison/data/mpd2020.xlsxand diff cell-by-cell againstlectures/mpd2020.xlsx4OWZWHTE5HXBBF4XLCCQTGDXY4HXOGOK) should be unchanged across snapshots from 2021-01-10 to 2026-01-02, i.e. predating the 2023-03-23 commit that added our copylong_run_growth's regional-data outputThis one matters because it determines the dataset's
class(constructed, notverbatim) and is recorded as the first entry in the upstream-delta register (#39).4. Manifests — 8 written in #38
filenamematches the sidecar's own name (noteload_manifests()keys onfilename, so a typo makes a manifest silently invisible)mpd2020.xlsxconstructed;longprices.xls/assignat/dette/fig_3verbatim;chapter_3.xlsxconstructed(hand-transcribed from print);caron.npy/nom_balances.npyconstructedwithbuilder_status: unrecovered.npyfiles claim no recoverable provenance. Re-test that: they should not be extractable from any column of the committed Sargent–Velde workbookssheets:blocks match what the lectures actually read —read_asshould reproduce the recorded shapepositional_reads: trueis set on the three French Revolution workbooks, which are read withheader=Noneplususecols/skiprows/nrows5. Tracker consistency
migration.ymlreads 13repointed/ 5landed; the 5 are exactly thefrench_revfilesrepointedrecord names the PRs that actually did the repointingCATALOG.mdis current (python scripts/build_catalog.pythengit diff --exit-code)scripts/build_audit.py all --strictpasses onmain6. The two blind spots — verify they are real, and that the counts are right
#42 — the audit cannot see prose references. Claim: 12 references (
{download}directives, markdown file links, directory links) point at files the migration will delete, and none fails any build.build_audit.pygenuinely does not report themfrench_rev.md:60-62(both repos) are still outstanding — they are set 3's problemRepoint rule 3 — the published site lags
main. Claim: 7 of 9 manifest repos publish on apublish*tag, with gaps up to 29 days.7. Decisions to sanity-check before they are built on
These were settled today and Track A's remaining work depends on them.
sources/-plus-plain-git storage design changesSCF_plus.dtais 99.1 MiB — within 0.9 MB of the hard block, and has no lecture consumergenerating_mini.mdreads its input fromhigh_dim_dataover the network — confirm, because archiving that repo without repointing it re-introduces a legacy-repo dependency silentlystr— notstring, notobject. Affects every future manifestdata.quantecon.orgis an AWS load balancer with no healthy targets — 503 on every hostname, not a live service8. Things deliberately not done
Confirm each was a decision rather than an oversight.
data.quantecon.orgis promoted (Licensing: record-and-track policy, and the per-dataset inventory for migration #35)business_cycle.mddumps (Do the business_cycle .md dumps belong in the published lectures/ tree? #13) — recommendation posted, decision not takenconsumers[].repo,source.doi,source.file_url,schema.sheetsand others are used by manifests but undocumented inmanifest-schema.ymlWhere the reasoning lives
PLAN.md(tracks, repoint rules, phases) ·AGENTS.md(working rules) · #8 (scaffolding checklist) · workspace-lectures#14 (workspace tracker) · workspace-lectures#23 (next work plan). PR descriptions carry the per-change reasoning and are the best record of why each decision was made.