Skip to content

VALIDATION: independent review of the 2026-08-18 Track C wave C2 + Track D migration and prune work #100

Description

@mmcky

Everything below was verified during the work, by the same session that did it — same tools, same URL forms, same environment, same mental model. This issue exists to break that circularity: run these checks in a fresh session with a diversified toolchain and report what survives.

Bias to test for: this session's verification monoculture was curl-based HTTP status/byte probes plus basename greps, and a single pinned venv (python 3.13, pandas 2.3.3, scipy 1.16.3) for every numeric claim; scripts/build_audit.py was run by the session itself for all four strict/dry-run measurements. Confirm reader-facing outcomes independently of all of that: open the published pages and notebooks in a browser, execute code rather than grepping it, use a different Python/pandas where a check involves parsing, and re-derive counts with your own parser rather than confirming the session's numbers.

What landed (2026-08-18, all merged the same day)

repo landed
QuantEcon/lecture-python-advanced.myst #374 (prune jb clean . --html in ci+publish, 6ff5ab3), #375 (wave C2 repoint, 878b87a), #376 (wave C2 deletion, 0c96791), tag publish-2026aug18 (rerun of run 32087672065 after a 75-min hang on the texlive apt step; deploy had not run when cancelled)
QuantEcon/data-lectures #98 (land wave C2 + Track D: 4 datasets, 2 builders, sources/dataBHS.mat under LFS, f42eeaf), #99 (flip all four to repointed, f09485c)
QuantEcon/lecture-python-programming #612 (Track D repoint, 55c87c9) — whose merge auto-opened translation-sync PRs lecture-python-programming.zh-cn#92, lecture-python-programming.fr#33, lecture-python-programming.fa#153 (all still open, deliberately)
QuantEcon/lecture-tools-techniques issue #11 filed (byte-identical dataBHS.mat, same broken-downloadable-notebook bug)
QuantEcon/workspace-lectures comments: #41 (prune repo 1 of 9 landed), #40 (re-audit booked for ~2026-08-24 with checklist)

1. Reader-facing outcomes

  • The clearest incident to confirm fixed: https://python-advanced.quantecon.org/_notebooks/five_preferences.ipynb downloads, and executes end-to-end in a clean environment with no dataBHS.mat on disk (Colab or local jupyter). Before this work the notebook called loadmat('dataBHS.mat') against a file the site serves at 404, so it could not run. The session verified the URL swap by grep, never by executing the published notebook — this check is the unexercised surface.
  • Same execution check for /_notebooks/risk_aversion_or_mistaken_beliefs.ipynb (live FRED-shaped read from data-lectures) and /_notebooks/match_transport.ipynb — at minimum their data-read cells.
  • The three published pages (risk_aversion_or_mistaken_beliefs.html, match_transport.html, five_preferences.html) render their data-driven figures — in particular five_preferences' consumption-growth histogram with two density curves (its figure was claimed shape-preserved through the loadmat→read_csv conversion).
  • https://github.com/QuantEcon/data-lectures/raw/main/lectures/<f> serves 200 with sizes 27963 / 14365 / 10160 / 793 for fred_data.csv / acs_data_summary.csv / dataBHS.csv / test_pwt.csv, and a never-existed control path 404s.
  • The old raw blob https://raw.githubusercontent.com/QuantEcon/lecture-python-advanced.myst/main/lectures/_static/lecture_specific/risk_aversion_or_mistaken_beliefs/fred_data.csv is 404, and the published site still serves /_static/.../fred_data.csv at 200 — expected under the settle policy until the next cache+publish cycle, and the distinction is the point: runtime reads were repointed before the blob left main, stale-serving clears later (re-audit on QuantEcon/workspace-lectures#40).

2. Artifact integrity

  • lectures/dataBHS.csv (sha256 13116a3d90ddc7f8b272b3ca903a136552147b21b8e9b3829472daa3d0d09c63) parses back bit-exactly from sources/dataBHS.mat (sha256 28c5f85286718e70b205f6a3fb269ebb49bd635194e2d0d488409b017be5e890) under float_precision='round_trip', and the lecture's 30-bin histogram of c[1:]-c[:-1] has identical counts and edges under pandas' default parser. Use a pandas other than 2.3.3.
  • sources/dataBHS.mat is a real LFS object (pointer oid = the sha256 above), and git check-attr filter -- sources/dataBHS.mat prints lfs.
  • The four lectures/<f> files are byte-identical to the copies the lectures read before migration: check against advanced.myst @ 878b87a^ paths and lecture-python-programming @ 55c87c9^ (git history, not the session's recorded hashes).

3. Highest-value claim: fred_data.csv reproduces from live FRED byte-for-byte

Everything about this file's committed builder status rests on it. Re-run builders/fred_data.py (pinned env per requirements.txt, then once more on a different pandas) and diff against the committed file. The claim's load-bearing details, each falsifiable: DFII5/DFII10 must be fetched with fredgraph's fq=Monthly&fam=avg aggregation (the bare series is daily); fredgraph now serves header observation_date where the file says DATE (the builder renames on read); the window is pinned 1953-04-01..2024-12-01. If the diff is non-empty, check whether FRED revised a value before concluding the builder is wrong — that distinction decides between fixing the builder and re-recording integrity.upstream.

  • Three fresh runs reproduce sha256 45a4fd41aeadf55072ea8e753c12dcb50ccb89bc34fb7eeb5e451ea0796811bc on separate days if possible (the session's three runs were minutes apart — stability over time is the unexercised margin).

4. Records written

  • The four new sidecar manifests parse (PyYAML and one non-Python YAML parser), their schema blocks match the measured bytes (row counts 861/351/236/8, the declared null structure — DFII pair 597 nulls each before 2003-01, zero nulls elsewhere), and each integrity.sha256 matches the committed file.
  • lectures/test_pwt.csv.yml's central negative claim: the committed values match no downloadable PWT vintage — re-derive for at least PWT 7.0 (pwt70_06032011version.zip member pwt70_w_country_names.csv) and PWT 6.3 (pwt63_nov182009version.zip): Argentina/Australia POP match 6.3 but cc/cg do not, and 7.0 differs on tcgdp/cc/cg for every row. The archive sha256s are in the manifest header; re-derive rather than trust.
  • sources/README.md's dataBHS un-refetchability trail: tomsargent.com/source_code.html 404s, larspeterhansen.org lists no code/data for the 2009 JET paper, the article page shows no supplement, and a GitHub-wide code search for dataBHS finds only QuantEcon-descended copies — with NEWQDATA as the search's positive control.

5. Tracker consistency

  • migration.yml parsed (never grep): 37 records, all status: repointed, zero landed/pending/final; the four new records cite Land wave C2 + Track D — the four remaining static datasets #98 and the repoint PRs #375/#612 with date 2026-08-18.
  • Strict audit green on data-lectures main against current clones, and CATALOG.md regenerates with an empty diff.
  • scripts/audit_annotations.yml no longer contains entries for fred_data.csv, acs_data_summary.csv, test_pwt.csv, or dataBHS.mat — but does still contain the lecture-dp orphan entry for acs_data_summary.csv under committed_unreferenced: (deleting that one would have been a mistake; confirm it survived).
  • builders/README.md's re-derived coverage counts (28 constructed, 18 with builders across 16 files, 10 unrecovered) match a fresh parse of the manifests.

6. Known blind spots

  • The deletion-time org sweep covered 278/278 default branches by Trees API. Its documented margins: non-default branches, fork content, binary contents. Probe at least advanced.myst's branch list for the three deleted basenames (the C1 validation found a .pkl variant on origin/hansen this way).
  • The session claims lecture-python-programming.ml holds test_pwt.csv as an orphan (no pandas.md/polars.md lectures exist there). Confirm from its tree, and confirm the manifest's consumer list (8 entries, three translation repos × two files plus the source repo × two) against your own sweep.
  • The .notebooks mirror self-heal claim: lecture-python-advanced.notebooks/{risk_aversion_or_mistaken_beliefs,five_preferences,match_transport}.ipynb on main each carry exactly one data-lectures/raw/main read and zero refs/heads / loadmat residue.

7. Decisions and mechanisms settled today

  • The prune discriminated on #376's preview: Monday's restored cache (built 2026-08-17 03:20, before the C2 deletion) contained the three files, and the preview 404'd all three while serving risk_aversion_or_mistaken_beliefs.html and eggs_backus.png at 200. Note the preview URL form: only the deploy-id-prefixed https://<id>--lustrous-melomakarona-3ee73e.netlify.app serves content; the deploy-preview-N-- form 404s everything including controls.
  • The five_preferences repoint's shape argument: with a pandas Series, c[1:] - c[:-1] misaligns on index and the figure silently changes; the merged cell uses data[['c']].to_numpy() to keep (236,1). Execute both variants and confirm the Series variant is actually wrong (the session asserted but did not run the broken variant).
  • Translation sync mechanics: the three sync PRs (zh-cn#92, fr#33, fa#153) carry all four URL swaps each with surrounding translations preserved, and the sync opens PRs rather than committing (auto-merge off) — which is why test_pwt.csv deletion is gated and was deliberately not done.

8. Deliberately not done

Where the reasoning lives

AGENTS.md and PLAN.md in this repo (URL forms, sources/ rules, repoint ordering); the work plan and its decisions on QuantEcon/workspace-lectures#45; the PR bodies of #98/#99 here and #374/#375/#376 in advanced.myst and #612 in lecture-python-programming; the re-audit checklist on QuantEcon/workspace-lectures#40.

For the validator: work in a session that did not do this work. Do not use the tool named under "Bias to test for" except where a check explicitly says to run it. Re-derive counts rather than confirming them. Where a check can be run against a surface the original session did not exercise, do that too — the margin beyond the checklist is where regressions hide. Deliver: one comment on this issue with a per-item verdict (confirmed / confirmed with caveat / refuted / not completable, with evidence), a new issue for any regression found (do not bury findings in the comment), and leave the checkboxes to the issue owner unless told otherwise.

Activity

  1. mmcky commented on Aug 18, 2026

    @mmcky
    ContributorAuthor

    Validation report

    Run in a fresh session that did no part of this work, 2026-08-18. 28 of 30 checks confirmed, 1 confirmed-with-refuted-detail, 1 not completable. One defect found and filed separately as #102. No regression in the shipped migration.

    Toolchain diversification. No curl anywhere — HTTP probes were python urllib and wget. No scripts/build_audit.py except in check 5.2, which asks for it by name. Numeric work ran in a fresh venv at python 3.12.13 / pandas 2.2.3 / scipy 1.14.1 / PyYAML 6.0.3, not the session's 3.13.9 / 2.3.3 / 1.16.3 (that env was used only for FRED run 1, which the check specifies). YAML was parsed twice — PyYAML and Ruby's Psych. Counts were re-derived from my own parsers, never read off the session's numbers. Where the checklist said "grep", I executed instead.

    1. Reader-facing outcomes

    1.1 five_preferences executes end-to-end — CONFIRMED. Downloaded /_notebooks/five_preferences.ipynb (200, 99852 B) and executed all 66 cells with nbclient in a scratch directory asserted empty of .mat/.csv: 0 errors, 60 s. The read cell is pd.read_csv('https://github.com/QuantEcon/data-lectures/raw/main/lectures/dataBHS.csv'); the notebook contains zero occurrences of loadmat, .mat, _static and refs/heads. The incident is fixed on the surface that was previously unexercised.

    1.2 The other two — CONFIRMED, and stronger than asked. Both ran in full, not just their data cells: risk_aversion_or_mistaken_beliefs 78 cells / 0 errors, match_transport 117 cells / 0 errors, same clean directory. Each carries exactly one data-lectures/raw/main read. risk_aversion's six _static hits are all <img> references to fig2_tom.png / eggs_backus.png / eggs_backus2.png — images, not data. match_transport's three .mat hits are substrings of .match_perfect_pairs.

    1.3 Figures render — CONFIRMED, by image comparison. Pulled the built PNGs off all three published pages and compared them against the figures my execution produced. five_preferences' consumption-growth histogram is identical to the published one — two density curves, LHS max 30 / RHS max 90, same bars. I re-derived the 30-bin counts from the .mat and from the CSV under pandas' default parser independently: [1,0,1,1,3,2,0,7,3,11,14,9,24,29,26,20,26,13,15,9,12,2,2,0,3,0,0,1,0,1], identical both ways, and the rendered bars match. risk_aversion's yield/spread panel renders with recession shading and the TIPS series starting 2003; match_transport's wage-dispersion figure renders from the ACS file.

    1.4 Raw URLs — CONFIRMED. Via urllib: 200 / 27963, 200 / 14365, 200 / 10160, 200 / 793 for fred_data / acs_data_summary / dataBHS / test_pwt, sha256 prefixes 45a4fd41, 719a3353, 13116a3d, 229316b3. A never-existed control path under the same prefix returns 404.

    1.5 Settle-policy distinction — CONFIRMED, and extended. Old blob at raw.githubusercontent.com/.../risk_aversion_or_mistaken_beliefs/fred_data.csv → 404; published site /_static/.../fred_data.csv → 200 / 27963. Beyond the checklist: the site also still serves /_static/lecture_specific/match_transport/acs_data_summary.csv (200 / 14365), the same stale-cache case; dataBHS.mat 404s at both the site root and the _static path and is gone from main — it was never served, which is the original bug.

    2. Artifact integrity

    2.1 Conversion gate — CONFIRMED under pandas 2.2.3. All 708 values parse back bit-exactly from lectures/dataBHS.csv against scipy.io.loadmat of sources/dataBHS.mat under float_precision='round_trip' (max abs diff 0.0, zero mismatches in all three columns). Under the default parser, exactly 18 values differ — matching the "18 of 708" figure recorded in the builder — and the 30-bin histogram of c[1:]-c[:-1] has identical counts and identical bin edges either way. Growth moments reproduce to printed precision: mean 0.004952, std 0.005050. Unasked bonus: builders/dataBHS.py's transform reproduces the committed CSV to the byte, sha256 13116a3d90ddc7f8b272b3ca903a136552147b21b8e9b3829472daa3d0d09c63.

    2.2 LFS — CONFIRMED. git cat-file -p HEAD:sources/dataBHS.mat is a pointer with oid sha256:28c5f85286718e70b205f6a3fb269ebb49bd635194e2d0d488409b017be5e890, size 5588; the working-tree file hashes to that same value at that size; git check-attr filter -- sources/dataBHS.mat prints lfs.

    2.3 Byte identity against git history — CONFIRMED. Read from history, not from recorded hashes:

    file historical blob data-lectures copy identical
    fred_data.csv advanced.myst 878b87a^:lectures/_static/.../risk_aversion_or_mistaken_beliefs/ lectures/ yes (45a4fd41…)
    acs_data_summary.csv advanced.myst 878b87a^:lectures/_static/.../match_transport/ lectures/ yes (719a3353…)
    test_pwt.csv lecture-python-programming 55c87c9^:lectures/_static/.../pandas/data/ lectures/ yes (229316b3…)
    dataBHS.mat advanced.myst 878b87a^:lectures/dataBHS.mat sources/ yes (28c5f852…)

    dataBHS.csv is a format conversion and correctly makes no byte-identity claim; its committed input is byte-identical to the historical blob, which is the right thing to check.

    3. fred_data.csv reproduces from live FRED — CONFIRMED (one caveat)

    Three fresh builder runs, two environments, each writing to a temp tree and diffing against the committed file:

    run environment result
    1 python 3.13.9 / pandas 2.3.3 (the requirements.txt pin, as the check specifies) byte-identical, 27963 B
    2 python 3.12.13 / pandas 2.2.3 byte-identical
    3 python 3.12.13 / pandas 2.2.3, ~40 min later byte-identical

    All three land on 45a4fd41aeadf55072ea8e753c12dcb50ccb89bc34fb7eeb5e451ea0796811bc. Caveat: all three are same-day. The stability-over-time margin cannot be exercised by a validator working the day the issue was filed; it needs a re-run in a week.

    The load-bearing details were falsified rather than read. Bare DFII5 over the same window is daily — 5717 rows, of which only 171 fall on a first-of-month stamp — and cannot fill the column; with fq=Monthly&fam=avg it is 264 rows and matches the committed DFII5 exactly. fredgraph serves header observation_date today (confirmed on both DFII5 and GS1) while the committed file's header line is DATE,GS1,GS5,GS10,DFII5,DFII10,USREC — the builder's rename-on-read is doing real work.

    4. Records written

    4.1 Sidecars — CONFIRMED. All four parse under PyYAML 6.0.3 and under Ruby's Psych (YAML.safe_load), as do migration.yml and scripts/audit_annotations.yml. Measured against the bytes with my own parser: row counts 861 / 351 / 236 / 8, declared column lists equal to the measured ones, and every integrity.sha256 matches. Null structure re-derived independently of the builder's asserts: DFII5 and DFII10 have 597 nulls each, all 597 fall before 2003-01 and none from 2003-01 onward (264 complete months), and every other column has zero; the monthly grid is unbroken 1953-04-01 to 2024-12-01 on first-of-month stamps; USREC is {0, 1}.

    4.2 The test_pwt negative claim — CONFIRMED at the centre, one detail refuted → #102. Downloaded all four archives and re-derived. The header's archive hashes are all correct (7.0 member 4623b92a…, 6.3 member f9609c42…, 6.2 a3293337…, 6.1 1477a118…). PWT 7.0 differs on tcgdp, cc and cg for every row and renames the isocode column, as recorded. PWT 6.3 has no tcgdp column and matches no cc or cg value. 6.2 and 6.1, which I checked as extra margin, are eliminated harder still — POP 0 of 8 and cc/cg 0 of 8 in both. The committed values match no downloadable PWT vintage; the record's conclusion stands. The refuted detail: 6.3 matches four POP values (ARG, AUS, ISR, ZAF) and four XRAT values exactly, not "Argentina and Australia POP but nothing else" — filed as #102, a one-sentence header correction. Separately I confirmed the naming claim the header rests on: country isocode is the 6.1/6.2-era column name on those releases' Data sheets, and both 6.3 and 7.0 renamed it to isocode.

    4.3 The dataBHS un-refetchability trail — CONFIRMED, with one part not completable. tomsargent.com/source_code.html returns 404 over http and the https endpoint resets the connection — exactly as sources/README.md describes. GitHub-wide code search for dataBHS returns QuantEcon-descended copies only (data-lectures, lecture-python-advanced.myst, lecture-python-advanced.notebooks, lecture-tools-techniques, lecture-mapping) plus the token-collision noise in unrelated JavaScript the README predicts; the NEWQDATA positive control returns hits, so the search was live. Not completable: the ScienceDirect article page 403s any scripted request, so "no supplementary material" needs a human browser. Caveat: larspeterhansen.org/research/papers/ returns 200 / 90 KB with zero .zip, .mat, "source code" or "replication" markers anywhere on it, but the 2009 JET paper is not listed on that index at all — consistent with "lists no code or data for it", not a direct per-paper confirmation.

    5. Tracker consistency

    5.1 migration.yml — substance CONFIRMED, the count REFUTED. Parsed, never grepped. datasets is a mapping of 40 records, not 37. Forty is the right number and the repo agrees with itself: CATALOG.md line 9 states "40 datasets", and every file in lectures/ except business_cycle_data.csv has a record. The count was never 37 — walking the last twelve commits that touched the file gives 26 → 26 → 26 → 31 → 31 → 33 → 33 → 36 → 36 → 40 → 40 → 40. So this is an error in the checklist, not in the tracker. Everything else in the item holds: all 40 are status: repointed, zero landed / pending / final (at f42eeaf the four sat at landed, and f09485c flipped them), pending holds only the P4 dynamic-snapshot placeholder, and the four new records each cite QuantEcon/data-lectures#98 dated 2026-08-18 with repoints to QuantEcon/lecture-python-advanced.myst#375 (the three C2 files) and QuantEcon/lecture-python-programming#612 (test_pwt.csv), same date.

    5.2 Strict audit and CATALOG — CONFIRMED. All nine clones fast-forwarded to current main first. scripts/build_audit.py scan --strict exits 0: "40 static files, 25 orphans, 23 live-API lectures". scripts/build_catalog.py regenerated CATALOG.md with an empty git diff (run under python 3.12.13, not the session's interpreter).

    5.3 audit_annotations.yml — CONFIRMED exactly, including the trap. Parsed with PyYAML. datasets: holds one entry and it is none of the four; no key or value anywhere mentions fred_data.csv, acs_data_summary.csv, test_pwt.csv or dataBHS.mat. The single surviving hit is under committed_unreferenced: (24 entries) — lecture-dp:lectures/_static/lecture_specific/match_transport/acs_data_summary.csv → {kind: orphan, note: "inherited copy, no consuming lecture in this repo"}. It survived.

    5.4 builders/README.md coverage — CONFIRMED. Fresh parse of all 40 manifests: 28 constructed (12 verbatim), 18 carrying a builder = 13 committed + 5 committed-frozen, across 16 distinct builder files (generating_mini.md and webscrape_forbes.ipynb produce two each — 18 − 16 = 2), and 10 unrecovered. No verbatim manifest carries a builder.

    6. Known blind spots

    6.1 Non-default branches — CONFIRMED, and the margin is not hypothetical. advanced.myst has 55 remote heads. 54 still carry lectures/dataBHS.mat, 39 carry acs_data_summary.csv, and origin/hansen carries fred_data.csv — the same branch the C1 validation caught a .pkl on. Mostly style-guide/additive_functionals-* and copilot/fix-* staleness. Nothing reads through those refs and the deletion sweep documented this exclusion, so it is a recorded margin rather than a defect — but if the goal is that the bytes stop being reachable, branch cleanup is the remaining lever.

    6.2 The .ml orphan and the consumer list — CONFIRMED by my own sweep. lecture-python-programming.ml's tree is 114 paths; it holds lectures/_static/lecture_specific/pandas/data/test_pwt.csv, and its only lectures/*.md is intro.md — no pandas.md, no polars.md. Orphan confirmed. My own wget sweep over the four live default branches:

    repo pandas.md polars.md old-path residue
    lecture-python-programming 1 3 0
    lecture-python-programming.zh-cn 1 3 0
    lecture-python-programming.fr 1 3 0
    lecture-python-programming.fa 1 3 0

    Exactly 1 data-lectures URL per pandas.md and 3 per polars.md, zero lecture_specific/pandas/data/ residue. The manifest's consumers list is exactly those eight entries.

    6.3 The .notebooks mirror — CONFIRMED. All three notebooks on lecture-python-advanced.notebooks main carry exactly one data-lectures/raw/main read and zero refs/heads, zero loadmat, zero _static CSV paths. Mirror head 39c1508, 2026-08-18T02:43:56Z. Extra: the mirror files are byte-identical to what the site serves (105676 / 99852 / 129259 B) — the self-heal is complete, not partial.

    7. Decisions and mechanisms

    7.1 The prune — PARTIALLY COMPLETABLE. The mechanism landed and is verifiable: 6ff5ab3 adds a jb clean . --html step after the cache restore in both workflows, present today at .github/workflows/ci.yml:56 and .github/workflows/publish.yml:74, before the PDF/notebook steps. The #376 preview deploy is gone, so the 404-vs-200 discrimination and the deploy-id-prefixed URL-form finding cannot be re-observed by anyone now. Recording them here as they stand: they are the only claims in this issue that rest on a surface that no longer exists, and the next prune rollout PR is the chance to re-derive the URL-form detail live.

    7.2 The shape argument — CONFIRMED by running the broken variant. With a pandas Series, c[1:] - c[:-1] index-aligns rather than shifts: the result is 236 rows, 2 NaN and 234 exact zeros, and its 30-bin histogram collapses to a single spike of 234 in one bin. The merged data[['c']].to_numpy() gives (236, 1), the same shape loadmat returns, and reproduces the published figure exactly. The assertion was correct and is now executed.

    7.3 Translation sync — CONFIRMED with a correction. The mechanism holds: sync opens PRs rather than committing, and auto-merge was off. But fr#33 was closed unmerged, superseded by QuantEcon/lecture-python-programming.fr#35 (same title, merged 04:00:18Z); zh-cn#92 merged 03:14:55Z and fa#153 merged 03:37:00Z. So all three repos did receive the swap, but for .fr it arrived on a different PR number than any record here names — worth knowing because QuantEcon/data-lectures#101's commit message says "zh-cn, fr and fa all merged their sync PRs", which is true of the change and not of #33. Translations are preserved around the reads: zh-cn's pandas.md keeps its Chinese prose either side of the read cell, fr's polars.md keeps "Rechargement du jeu de données" and "Petit jeu de données -- Penn World Tables (~8 lignes)" on all three.

    8. Deliberately not done — CONFIRMED, and the state has moved

    test_pwt.csv is still on lecture-python-programming main and still served by both hosts (200 / 793 from raw and from python-programming.quantecon.org) — not deleted, as intended. QuantEcon/lecture-tools-techniques#11, QuantEcon/data-lectures#87, QuantEcon/data-lectures#84 and QuantEcon/workspace-lectures#41 are all open. publish-2026aug18 points at 878b87a, the repoint commit — the publish preceded the deletion, as claimed. All six PRs are merged at the stated commits.

    State change since filing: the sentence "all still open, deliberately" about the three sync PRs is no longer true — all three landed within hours (fr via #35). The deletion gate for test_pwt.csv is therefore down to republish lag, which QuantEcon/data-lectures#101 already records.

    What I would do next

    The one thing worth another pass is the fred_data.csv stability-over-time margin — three same-day runs are three samples of one network moment. A single re-run of builders/fred_data.py in a week, folded into the QuantEcon/workspace-lectures#40 re-audit, closes it. The 54 stale branches carrying dataBHS.mat are the other loose end, and belong with the org-wide sweep rather than with this migration.

    Checkboxes left unticked for the issue owner.

    Validator environment: python 3.12.13 / pandas 2.2.3 / scipy 1.14.1 / numpy via that venv / PyYAML 6.0.3 / xlrd 2.0.2 / openpyxl 3.1.5 / nbclient; Ruby Psych for the second YAML parse; urllib and wget for HTTP; macOS 25.6.0.

  2. mmcky commented on Aug 18, 2026

    @mmcky
    ContributorAuthor

    Closing: all 27 checks are ticked, and every loose end from the validation report above now has a resolution on record.

    open at report time resolution
    §4.2's refuted detail — the PWT 6.3 overlap understated in test_pwt.csv.yml Fixed and merged in #103 (with Copilot's wording refinement); #102 auto-closed on merge
    §4.3's "not completable" — ScienceDirect 403s scripted requests Human browser check 2026-08-18: the JET article page shows no supplementary material. The dataBHS un-refetchability trail is now confirmed on all four legs
    §8's Colab run for #84 Run by a human 2026-08-18. It surfaced a distinct, pre-existing blocker — usetex=True requires a system LaTeX Colab does not have, failing ~50 cells before the data read — filed as QuantEcon/lecture-python-advanced.myst#377, with an addendum on QuantEcon/lecture-tools-techniques#11 (its copy sets it twice). Not a migration regression: the setting predates the work and my local execution of all three notebooks (0 errors) stands
    §3's same-day caveat — FRED reproduction stability over time Delegated to the ~08-24 re-audit on QuantEcon/workspace-lectures#40, which now carries the builder re-run line with the expected sha256 45a4fd41… and the revised-value-vs-broken-builder decision rule
    §7.1's expired preview surface Accepted as recorded; QuantEcon/workspace-lectures#41 now documents that #376's preview held the prune's only isolating evidence and gives the remaining 8 rollout repos the method for capturing it in-thread while their previews are alive
    §5.1's "37 records" A miscount in this checklist, not in the tracker — migration.yml has 40 records and the repo agrees with itself; no action
    §7.3's fr#33 Closed unmerged, superseded by QuantEcon/lecture-python-programming.fr#35 same-day; the correction is recorded here and on QuantEcon/workspace-lectures#40 so future audits follow the right PR number

    Also added to QuantEcon/workspace-lectures#40 while wrapping up: the concrete C2 checklist rows the booking asked for when the deletion merged (derived from 0c96791's tree, with deletion-day baselines), and the Track D gate state.

    Still live, tracked elsewhere by design: the test_pwt.csv deletion (republish-gated, QuantEcon/workspace-lectures#40), the phase-2 renames (#87), the tools-techniques repoint (QuantEcon/lecture-tools-techniques#11), the 8 remaining prune repos (QuantEcon/workspace-lectures#41), and the new QuantEcon/lecture-python-advanced.myst#377.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions