Skip to content

VALIDATION: independent review of the 2026-08-06 migration work #45

Description

@mmcky

Independent validation of the migration work landed on 2026-08-06. Everything below was verified during the work by the same agent that did it — this issue exists so it can be checked by someone who was not.

Bias to test for: most claims here were verified by running scripts/build_audit.py, which is also the tool the work was built around. A validation pass should confirm reader-facing outcomes independently of that tool wherever possible — fetch the URLs, open the notebooks, read the published pages.

What landed

Repo PRs merged
data-lectures #36, #38, #41, #43, #44
lecture-python-intro #823, #824 (#825 open)
lecture-wasm #52, #53
workspace-lectures #22

Issues opened: #35, #37, #39, #40, #42 here; workspace-lectures#23. Closed: #20. Rewritten: #8, workspace-lectures#14.
Publishes: publish-2026aug06 (set 1), publish-2026aug06b (set 2).

1. Reader-facing outcomes — check these first, without the audit tool

Two datasets moved. The claim is that no reader-visible output changed and nothing 404s.

  • intro.quantecon.org/long_run_growth.html and inflation_history.html render, figures intact
  • The same two pages on the wasm site
  • Download each lecture's notebook and run it — the data cells should fetch from data-lectures and succeed
  • Open each in Colab — this is the failure mode the whole migration exists to fix
  • Click the {download} link for chapter_3.xlsx in inflation_history — it should serve the spreadsheet, not 404
  • Follow the "hosted on GitHub" link in inflation_history — should reach data-lectures' CATALOG.md
  • Spot-check a figure against the pre-migration published version (e.g. via the Wayback Machine) — the bytes are identical, so the figures must be too

Known live-site history worth confirming is now clean: between set 1 merging and publish-2026aug06, the published long_run_growth notebook pointed at a deleted file and returned 404. That window is closed, but it is the single clearest thing to verify independently.

2. Data integrity — the claim is byte-identity

Every repoint asserted the served bytes are identical to what the lecture used before.

  • For each of mpd2020.xlsx, longprices.xls, chapter_3.xlsx: fetch from github.com/QuantEcon/data-lectures/raw/main/lectures/<file> and compare sha256 against integrity.sha256 in its sidecar manifest
  • Compare the same against the pre-deletion blob in lecture-python-intro's git history (git show <pre-merge-sha>:lectures/datasets/<file> | shasum -a 256)
  • Confirm the three manifests' sha256 values were not simply copied from the file they describe without an independent check

3. The mpd2020.xlsx provenance finding — the highest-value claim to re-test

Claim: the file is not a pristine Maddison release. All 21,682 data rows match upstream, but three header labels on the Regional data sheet were edited locally, and long_run_growth depends on them via header=(0,1,2).

  • Re-download from rug.nl/ggdc/historicaldevelopment/maddison/data/mpd2020.xlsx and diff cell-by-cell against lectures/mpd2020.xlsx
  • Confirm exactly three cells differ, and that they are header labels rather than data
  • Independently confirm the edits are ours: the Internet Archive digest for the upstream file (4OWZWHTE5HXBBF4XLCCQTGDXY4HXOGOK) should be unchanged across snapshots from 2021-01-10 to 2026-01-02, i.e. predating the 2023-03-23 commit that added our copy
  • Sanity-check the consequence: replacing the file with a clean upstream copy should visibly change long_run_growth's regional-data output

This one matters because it determines the dataset's class (constructed, not verbatim) and is recorded as the first entry in the upstream-delta register (#39).

4. Manifests — 8 written in #38

  • Every manifest parses, and filename matches the sidecar's own name (note load_manifests() keys on filename, so a typo makes a manifest silently invisible)
  • Provenance classifications are defensible: mpd2020.xlsx constructed; longprices.xls / assignat / dette / fig_3 verbatim; chapter_3.xlsx constructed (hand-transcribed from print); caron.npy / nom_balances.npy constructed with builder_status: unrecovered
  • The two .npy files claim no recoverable provenance. Re-test that: they should not be extractable from any column of the committed Sargent–Velde workbooks
  • The sheets: blocks match what the lectures actually read — read_as should reproduce the recorded shape
  • positional_reads: true is set on the three French Revolution workbooks, which are read with header=None plus usecols/skiprows/nrows

5. Tracker consistency

  • migration.yml reads 13 repointed / 5 landed; the 5 are exactly the french_rev files
  • Each repointed record names the PRs that actually did the repointing
  • CATALOG.md is current (python scripts/build_catalog.py then git diff --exit-code)
  • scripts/build_audit.py all --strict passes on main

6. The two blind spots — verify they are real, and that the counts are right

#42 — the audit cannot see prose references. Claim: 12 references ({download} directives, markdown file links, directory links) point at files the migration will delete, and none fails any build.

  • Confirm build_audit.py genuinely does not report them
  • Confirm the count. It was revised 10 → 12 mid-session after Copilot found a directory link the original sweep missed — the sweep shared the blind spot it was documenting, so 12 should not be taken on trust either
  • Confirm the six in french_rev.md:60-62 (both repos) are still outstanding — they are set 3's problem

Repoint rule 3 — the published site lags main. Claim: 7 of 9 manifest repos publish on a publish* tag, with gaps up to 29 days.

  • Verify the trigger for each repo from its workflow files
  • Verify that rendered HTML survives a deleted data file (figures baked at build time) while notebooks do not — this is why nothing reported the set 1 breakage

7. Decisions to sanity-check before they are built on

These were settled today and Track A's remaining work depends on them.

  • Both SCF minis fit plain git — 31.3 MiB and 72.4 MiB against GitHub's 100 MiB limit. If wrong, the whole sources/-plus-plain-git storage design changes
  • SCF_plus.dta is 99.1 MiB — within 0.9 MB of the hard block, and has no lecture consumer
  • generating_mini.md reads its input from high_dim_data over the network — confirm, because archiving that repo without repointing it re-introduces a legacy-repo dependency silently
  • pandas 3 reports text dtype as str — not string, not object. Affects every future manifest
  • data.quantecon.org is an AWS load balancer with no healthy targets — 503 on every hostname, not a live service

8. Things deliberately not done

Confirm each was a decision rather than an oversight.

Where the reasoning lives

PLAN.md (tracks, repoint rules, phases) · AGENTS.md (working rules) · #8 (scaffolding checklist) · workspace-lectures#14 (workspace tracker) · workspace-lectures#23 (next work plan). PR descriptions carry the per-change reasoning and are the best record of why each decision was made.

Activity

  1. mmcky commented on Aug 6, 2026

    @mmcky
    ContributorAuthor

    Independent validation — results

    Ran the full checklist without scripts/build_audit.py except where a check explicitly asks for it: fresh HTTP fetches, independent hashing, a cell-by-cell workbook diff with openpyxl, re-executing the lectures' data cells under pandas 3.0.5, re-reading every manifest read_as recipe against the committed bytes, and a headless-Chromium CORS test from the wasm site's own origin. Validation scripts were written from scratch for this pass.

    Verdict: the work holds. Every byte-identity, provenance, manifest, tracker, and blind-spot claim verified — several came out stronger than claimed. One reader-facing regression was found that the checklist's own checks would not have caught (the wasm site cannot fetch the repointed URLs from the browser — CORS, filed as #46), one claim needs a caveat updated (data.quantecon.org now fronts a third-party TLS cert before its 503), and two checks could not be completed as specified (Wayback figure compare — archive.org was down all session; an in-browser Colab session — verified by proxy instead).

    The one real finding — wasm #52/#53 broke in-browser data loading (#46)

    lecture-wasm executes cells in the reader's browser (JupyterLite/Pyodide via pyodide_http.patch_all() — XHR, CORS-enforced per hop). The repoints changed the wasm URL form from direct raw.githubusercontent.com (serves access-control-allow-origin: *) to github.com/…/raw/ (a 302 whose response carries an empty ACAO header, failing the CORS check before the redirect is followed). Confirmed empirically with headless Chromium fetching from quantecon.github.io/lecture-wasm/long-run-growth/: both repointed URLs fail with Failed to fetch; the pre-repoint form returns 200 with the exact 1,765,204 bytes. The pre-repoint wasm URLs were a deliberate wasm-specific adaptation that the repoint normalised away. Since the wasm build bakes no outputs, figures on those two wasm pages exist only after in-browser execution, which now dies at the first data cell. inequality.md and mle.md have the same failure pre-existing (via high_dim_data github.com/raw URLs); heavy_tails.md is CORS-clean (media.githubusercontent.com). The repoints were correctly sequenced — the old URLs point at files #825 has since deleted — only the form is wrong, and the fix is a two-lecture URL flip plus a repoint rule. Details, evidence table, and the Phase 4 implication (data.quantecon.org must serve ACAO * before wasm can cut over) in #46.

    So the claim "no reader-visible output changed" is true for the intro site, downloaded notebooks, and Colab — and false for wasm interactive execution.

    1. Reader-facing outcomes

    Check Result
    intro long_run_growth.html / inflation_history.html ✅ both 200; figure _images/*.png spot-checks 200
    Same two pages on the wasm site ⚠️ pages 200 at /long-run-growth/, /inflation-history/, and their sources point at data-lectures — but in-browser execution broken, see #46
    Download each notebook and run it ✅ fetched intro.quantecon.org/_notebooks/*.ipynb; both reference only data-lectures URLs; executed the data cells live under pandas 3.0.5 — Full data (21682, 5), Regional (23, 18), longprices (401, 4) 1600–2000, all 11 chapter_3 sheets read
    Open each in Colab ✅ by proxy: fetched the exact bytes Colab loads (lecture-python-intro.notebooks@main) — data-lectures URLs only — and executed those fetches successfully; no live Colab session run
    {download} chapter_3.xlsx link ✅ 200, 73,281 bytes, opens as xlsx (all 11 sheets parsed)
    "hosted on GitHub" link ✅ → data-lectures/blob/main/CATALOG.md, 200
    Figure spot-check vs Wayback ⚠️ archive.org CDX was 503/429 for page snapshots all session. Substitute proof: the repoint commits' lecture diffs are pure URL-string swaps (git show 3ccd49c8, 69307fff — no code-logic change) and the data bytes are identical (§2), so the figures are determined-identical. The _images/<content-hash>.png filename comparison against a pre-migration snapshot remains worth doing when IA recovers
    Set 1 404 window closed ✅ published notebooks now fetch data-lectures URLs, all 200

    2. Data integrity — byte-identity

    Three-way match on all three files: freshly fetched raw/main bytes = sidecar integrity.sha256 = pre-deletion blob in intro's history (3ccd49c8^ for mpd2020, 3f715b84^ for the set 2 pair).

    File sha256 (all three sources identical)
    mpd2020.xlsx f67af0fd…e3bceb
    longprices.xls 1160f9d3…b344f7
    chapter_3.xlsx ee17d003…d30923

    The "not simply copied" concern is answered by construction: the git blob and the live server were hashed independently of the manifests and agree with them. mpd2020's add date also checks out — e23ae056, 2023-03-23, PR #120, blob never modified since, exactly as the manifest says.

    3. mpd2020 provenance — reproduced exactly, and stronger

    • Cell-by-cell diff (openpyxl, all six sheets, 379,701 cells): exactly 3 differ, all on Regional data row 1 — GDP pc 2011 prices→gdppc_2011, Population→pop, and gdppc_2011 added over the unlabeled world column. All 21,683 rows of Full data identical. Matches the manifest's delta record verbatim.
    • IA digest 4OWZWHTE5HXBBF4XLCCQTGDXY4HXOGOK is constant across all 10 snapshots, 2021-01-10 → 2026-04-01 — beyond the claimed 2026-01-02 — bracketing our 2023-03-23 commit. Better: the base32-SHA-1 of today's fresh rug.nl download equals that digest exactly, so upstream has served one byte-identical file for 5½ years and the edits are unambiguously ours.
    • Consequence check: the lecture's header=(0,1,2) read against the pristine upstream raises KeyError: 'gdppc_2011' — the local edits are load-bearing, so constructed is the right class and a "clean upstream refresh" would break the lecture outright, not subtly change it.

    4. Manifests

    • All 18 sidecars parse; every filename matches its sidecar's name — no silently-invisible manifest.
    • Classifications match the claimed set: mpd2020/chapter_3/caron/nom_balances constructed, longprices/assignat/dette/fig_3 verbatim; both .npy carry builder_status: unrecovered.
    • .npy re-test, done independently: scanned every cell of every sheet of the three workbooks. caron.npy: neither terminal value (96.696, 0.431) appears anywhere, no contiguous column matches — claim confirmed. nom_balances.npy: its terminal values 90.0 and 33,555.59 do occur in assignat.xlsx (Fig6, Data, Data2, Denomina, Dom-nat, Ramel, Ramel2, seignor) — and the manifest already records precisely that partial overlap, notes no contiguous 81-value column exists (confirmed) and that the series max 37,540.933 appears nowhere (confirmed). The manifest is accurate to the decimal; "overlaps without being extractable" is the right description.
    • Every read_as recipe re-executed against the committed bytes reproduces the recorded shape and the recorded known_nulls_total — all 10 ranges across assignat (4), dette (5), fig_3 (1), plus both npy shape/dtype blocks and the mpd2020/longprices sheet reads (via the lecture's full recipe incl. index_col/iloc).
    • positional_reads: true present on exactly the three French Revolution workbooks and no others.

    5. Tracker consistency

    • migration.yml: 18 records, 13 repointed / 5 landed, and the 5 are exactly assignat, dette, fig_3, caron.npy, nom_balances.npy. Every repointed record names its PRs; spot-checked that intro#823/#824 and wasm#52/PLAN: correct the stale figures, widen rule 1 to the org, restate rule 6 #53 really contain the corresponding repoints (they do — subjects and diffs match).
    • build_catalog.py → zero diff, CATALOG.md current.
    • build_audit.py all --strict on today's main (i.e. post-#825): exit 0.

    6. The two blind spots

    7. Decisions

    Claim Result
    SCF minis fit plain git ✅ from the LFS pointers: 32,853,734 B = 31.33 MiB and 75,902,999 B = 72.39 MiB, both under the 104,857,600 B hard block (the larger one is past the 50 MiB warning threshold, worth knowing)
    SCF_plus.dta 99.1 MiB, no lecture consumer ✅ 103,934,093 B = 99.12 MiB, 0.88 MiB under the block; grep across all nine workspace repos finds no lecture read — only PLAN.md mentions it
    generating_mini.md reads over the network ✅ line 37: pd.read_stata('https://github.com/QuantEcon/high_dim_data/blob/main/SCF_plus/SCF_plus.dta?raw=true')
    pandas 3 text dtype is str ✅ pandas 3.0.5: str(dtype) → str (repr is StringDtype(storage='python', na_value=nan)) — not string, not object
    data.quantecon.org = empty AWS LB ✅ with a caveat to record: behind TLS it is awselb/2.0 returning 503 (http → 301 → https), and rDNS of 52.64.86.66 still says data.quantecon.org — but the LB now presents a certificate for *.dev.cloud.payok.com.au, so a browser hits a cert error before any 503. Not a live service either way, but the third-party cert is new relative to the recorded finding and worth a line on #37 — if that LB is no longer QuantEcon-controlled, the DNS record is dangling on someone else's infrastructure

    8. Deliberately not done — all confirmed as decisions

    intro#825 was open when this issue was filed and has since merged (2026-08-06 08:06 UTC) — the gate (publish first) was honoured, and §2/§5 above were run against the post-#825 tree. #35 (licence policy deferral), #13 (business_cycle dumps, recommendation posted, no decision), and ws#23 (step 6 write-backs) are open and say what the checklist says they say. Schema drift is real and measurable: doi (6 manifests), file_url (7), sheets (8), positional_reads (3) all appear in manifests and nowhere in manifest-schema.yml. The delta register #39 exists with mpd2020 as its sole entry, correctly distinguishing diverged from the eleven unverifiables. #20 closed, #37/#40 open as stated.

    Still open after this pass

    1. Repoint rule for wasm: in-browser reads need raw.githubusercontent.com — the github.com/*/raw/ form fails CORS, breaking long_run_growth and inflation_history on the wasm site #46 — the wasm CORS regression: two-lecture URL-form fix + repoint rule + Phase 4 ACAO acceptance criterion.
    2. Wayback _images hash comparison for one lecture, when archive.org recovers — the determinism argument above is solid, but the checklist wanted the published bytes compared and that half is still owed.
    3. A true in-browser Colab run of each notebook (verified here at the bytes-Colab-loads + fetches-execute level only).
    4. The data.quantecon.org cert observation → Track Y (deferred): data.quantecon.org DNS + custom domain — retired in favor of qeld (D11); reopenable, NXDOMAIN as of 2026-08-10 #37.

    Method note for the record: this validation was performed by a different agent from the one that did the migration work, with its own scripts (openpyxl cell diff, pandas re-reads of every read_as, LFS-pointer size reads, headless-Chromium CORS harness) rather than the repo's audit tooling, per the bias warning at the top of this issue.

  2. mmcky commented on Aug 6, 2026

    @mmcky
    ContributorAuthor

    Closing: the validation ran in full and every checklist item has a verdict in the report above — all boxes ticked accordingly.

    Where the three caveats landed, so nothing rides on this issue staying open:

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions