Skip to content

Flip wave P4 to repointed — the four business_cycle snapshots - #153

Draft
mmcky wants to merge 1 commit into
mainfrom
wasm/business-cycle-snapshots
Draft

mmcky wants to merge 1 commit into
mainfrom
wasm/business-cycle-snapshots

Conversation

@mmcky

@mmcky mmcky commented Sep 30, 2026

Copy link
Copy Markdown
Contributor

This is block C of #118, this repo's half of lecture-wasm's adoption of the four business_cycle snapshots. When lecture-wasm's adoption PR (QuantEcon/lecture-wasm#85, for QuantEcon/lecture-wasm#70) merges, lectures/business_cycle.md there reads gdp_growth_annual.csv, unemployment_rate_annual.csv, private_credit_to_gdp.csv and us_business_cycle_monthly.csv from raw/main. This PR records that on the same day: lecture-wasm becomes the consumer of all four, their migration.yml records go from landed to repointed, and P4 is ticked in PLAN.md.

Part of #146 and #118. #146 is done when the next audit-dashboard run on main after this merges is green and deploys, so it gets closed by hand at that point rather than on merge.

Before merging

What changes

File Change
The four manifests: lectures/gdp_growth_annual.csv.yml, unemployment_rate_annual.csv.yml, private_credit_to_gdp.csv.yml, us_business_cycle_monthly.csv.yml consumers becomes one entry, {repo: QuantEcon/lecture-wasm, file: lectures/business_cycle.md, on_refresh: rebuild}, in place of each "None yet" comment. In us_business_cycle_monthly.csv.yml nothing else changes, so #152's licence and source edits to that file merge cleanly
lectures/gdp_growth_annual.csv.yml The three comments that described the inherited bytes, which the first refresh (#112) replaced, are vintage-neutral now: source.version, the comment above retrieved and the comment above integrity.upstream.status. The Integrity header said "no lecture has ever read this file", which this PR makes false, so it is reworded. The reason lecture-wasm had excluded the lecture is stated precisely: neither wbgapi nor pandas_datareader ships with Pyodide, and FRED's fredgraph.csv sends no CORS header
migration.yml The four records go from landed to repointed, each with one repoints entry for lecture-wasm, and the comment on the gdp_growth_annual.csv record names the consumer. The stale pending: P4 wave ("a maintained UNRATE twin") is removed, so the dashboard draws "Wave P4" once; pending is now []. The closing comment says final means qeld adoption, instead of routing it through #15 and Phase 4 DNS
scripts/audit_annotations.yml lecture-wasm:business_cycle:wbgapi and lecture-wasm:business_cycle:pandas_datareader are deleted: B shows those calls in non-executing blocks, and the scan finds no API marker in the adopted lecture. Intro's pandas_datareader note called the snapshot pipeline unadopted and named the World Bank file; it now names the FRED twin, us_business_cycle_monthly.csv, and says intro keeps its live call by decision
PLAN.md P4 in Phase 8 is ticked, with the flip, the dry-run and the canary proof (#147, runs 36654189187 and 36654357802). The consumer fan-out line in Phase 5 records that a consumer which executes at read time needs no dispatch, as #118 sets out under "Not in scope"
.github/workflows/refresh-snapshots.yml The header no longer says no snapshot has a consumer. The first one reads raw/main in the reader's browser and needs no dispatch
scripts/snapshots.py The refresh PR body lists a rebuild consumer as "→ dispatch a rebuild (not wired yet: PLAN Phase 5)" instead of implying that a dispatch happens
CATALOG.md Regenerated. The four files leave "awaiting repoint": 44 datasets, 44 read by lectures today

A parsed comparison with main confirms that scope. Each manifest differs only in consumers, plus source.version in gdp_growth_annual.csv.yml. migration.yml differs only in the four records' status and repoints and in pending, and audit_annotations.yml only in the two deletions and the one note.

The decision it records

on_refresh: rebuild for lecture-wasm on all four files. No prose in B's lecture quotes a value a refresh can move: the only numbers outside its code are dates and periods of historical events, such as 2010-2011 or the 1970s, and its ranges run "to the present". I checked this on B's commit, so it needs checking again if B's prose changes in review. The value sets nothing off today, because consumer fan-out is out of scope: lecture-wasm runs the lecture in the reader's browser and reads raw/main when it runs.

The placeholder to fill

B is QuantEcon/lecture-wasm#85, already filled in. Its merge date is unknown until it merges. The placeholder parses as a YAML string, so the audit and the dashboard accept it and print it as text; the migration page reads "completed B_MERGE_DATE" until it is filled.

File Where (line numbers as of this PR's first push) Placeholder
migration.yml the repoints entry of gdp_growth_annual.csv (line 855), unemployment_rate_annual.csv (873), private_credit_to_gdp.csv (886) and us_business_cycle_monthly.csv (899) date: B_MERGE_DATE
PLAN.md the P4 line **Complete B_MERGE_DATE**
PLAN.md "Consumer half and the flip" B_MERGE_DATE in its heading

This fills all six (macOS sed; on Linux drop the ''). Keep the date unquoted, as in the other records, so YAML reads it as a date:

sed -i '' -e 's/B_MERGE_DATE/YYYY-MM-DD/g' migration.yml PLAN.md

Checks run (2026-09-30)

Check How Result
validate-datasets scripts/validate_datasets.py, then with --builders, as the workflow runs them, in Python 3.12.13 venvs with PyYAML 6.0.3 and wbgapi 1.0.12, under pandas 2.3.3 and 3.0.5 44 of 44 manifests pass, and the builder layer is all green. The same as on main
consumed-file-check .github/scripts/check_consumed_files.py, then scripts/build_catalog.py and git diff --exit-code -- CATALOG.md 44 manifests, 47 files hash-checked, 0 errors. CATALOG.md is current
The refresh stamp and PR body scripts/snapshots.py stamp and pr-body on a scratch copy Only the stamped lines change, the rewritten comments survive, and the stamp's read-back passes. The PR body lists lecture-wasm with the new wording
The strict audit in both directions Below As #146 expects
The dashboard pages Below The four count as migrated, "Wave P4" appears once, and no count is negative

The strict audit dry-run

This is python scripts/build_audit.py all --strict over shallow clones of the eight lecture repos, cloned as audit-dashboard.yml clones them, with #151's annotation applied. The scan reads each clone's origin/main, so for "B's branch" the lecture-wasm clone's origin/main was set to B's head, rebased on lecture-wasm main at ff9196d.

data-lectures tree lecture-wasm Exit Warnings
main (records at landed) B's branch 1 exactly 4, "marked landed but some consumer already reads data-lectures", one per file
this PR (records at repointed) B's branch 0 none
this PR main (ff9196d) 1 4 × "marked repointed but no lecture reads it at all", and 2 × missing_api_annotations for lecture-wasm's business_cycle

Without #151's entry, main alone exits 1 on missing_api_annotations: lecture-python-intro:mle:yfinance, which is why #151 goes first.

The dashboard pages

From the exit-0 run (this PR against B's branch):

Page main today This PR
Overview, migration legend "migrated … (40)", "copied here, lectures not yet switched (4)", "not migrated (-4, of which 0 queued …)" "migrated … (44)", "not yet switched (0)", "not migrated (0 …)"
migration.html, datasets migrated 40 44
migration.html, "Wave P4" twice, both open: "4 datasets, consumed by 0 lecture series" and the pending entry once: "Wave P4 — 4 datasets, consumed by 1 lecture series — completed B_MERGE_DATE"
The four dataset records landed repointed, "✓ verified against today's scan", consumed by lecture-wasm · business_cycle

With #148 merged in as well, its legend correction holds after the flip: 44 / 0 / 0, "Wave P4" once, and the broad sweep "completed 2026-08-18".

Merging alongside the open PRs

PR Shared files Result
#151 none Independent. It must be on main before this PR's checks can pass
#152 us_business_cycle_monthly.csv.yml A three-way merge of that file is clean in both orders: the consumers block from here, the licence, source and schema text from #152. #152 leaves CATALOG.md unchanged
#148 none git merge-tree is clean in both orders
#145 PLAN.md, private_credit_to_gdp.csv.yml One conflict hunk in PLAN.md: this PR's P4 block against #145's edits to the two lines after it. Keep this PR's P4 block, then #145's two lines. The manifest merges cleanly
#144 CATALOG.md, migration.yml migration.yml merges cleanly, with #144's record after the four P4 records. CATALOG.md conflicts in two hunks; regenerate it, which gives "45 datasets · 45 read by lectures today"

If anything else that regenerates CATALOG.md lands first, such as the monthly refresh PR expected on Monday 2026-10-05, rebase and re-run python scripts/build_catalog.py.

Left for whichever of this PR and #145 lands second

#145 rewrites PLAN.md lines that this flip makes false. Whichever lands second applies these:

  1. The paragraph that says CATALOG.md and migrated both say 40, or (in docs: close-out tidy — status, convention pointers and stale facts after the migration #145) that the gap is the four Track E snapshots: CATALOG.md and migrated agree at 44 after the flip, or at 45 once Add us_household_net_worth_2022.csv: SCF 2022 household net worth with survey weights #144 lands, whose file is listed ahead of its lecture. The last gap was the four business_cycle snapshots, landed 2026-09-01 at consumers: [] and flipped on the day lecture-wasm adopted them.
  2. The PLAN: pilot migration — one dataset per hosting pattern to validate the datasets design meta#338 row: "P1–P3 complete, P4 (the dynamic-snapshot twin) in progress" becomes "P1–P4 complete (P4 on B_MERGE_DATE)".
  3. The Track E row: "landed 2026-09-01 at consumers: []" becomes "landed 2026-09-01, read by lecture-wasm since B_MERGE_DATE", and "Next: the lecture-wasm adoption …, then the flip; session plan lecture-wasm dynamic snapshots: business_cycle reads the four data-lectures snapshots (P4) #118" becomes "the lecture-wasm adoption and the flip landed B_MERGE_DATE (P4); project tracker lecture-wasm dynamic snapshots: business_cycle reads the four data-lectures snapshots (P4) #118".

Keep the header's "40 of 40 static". The header of private_credit_to_gdp.csv.yml still calls the filename provisional; #145 corrects it, so this PR leaves it alone. Once B's preview gate has passed, the Phase 8 line that says no in-browser Pyodide run is recorded can also note B's preview as the first, for these four files.

🤖 Generated with Claude Code

lecture-wasm's business_cycle now reads gdp_growth_annual.csv,
unemployment_rate_annual.csv, private_credit_to_gdp.csv and
us_business_cycle_monthly.csv from raw/main (QuantEcon/lecture-wasm#70),
so the records here say so on the same day. This is block C of #118,
worked from the checklist in #146.

- The four manifests record the consumer, {repo: QuantEcon/lecture-wasm,
  file: lectures/business_cycle.md, on_refresh: rebuild}, in place of
  each "None yet" comment. `rebuild`, because no prose in the lecture
  quotes a value a refresh can move: its numbers are years of historical
  events, and its ranges run "to the present".
- gdp_growth_annual.csv.yml: source.version, the comment above
  `retrieved` and the comment above integrity.upstream.status described
  the inherited bytes the first refresh (#112) replaced; they are
  vintage-neutral now. The Integrity header's "no lecture has ever read
  this file" goes too, since this commit makes it false, and the reason
  lecture-wasm had excluded the lecture is stated precisely: neither
  wbgapi nor pandas_datareader ships with Pyodide, and FRED's
  fredgraph.csv sends no CORS header.
- migration.yml: the four records go from landed to repointed, each with
  one lecture-wasm repoints entry. The stale pending P4 wave (the "UNRATE
  twin", which drew "Wave P4" twice on the dashboard) is removed, and the
  closing comment says `final` means qeld adoption instead of routing it
  through the retired data.quantecon.org cutover (#15).
- scripts/audit_annotations.yml: the lecture-wasm business_cycle wbgapi
  and pandas_datareader entries go, since the adopted lecture shows those
  calls in non-executing blocks. Intro's pandas_datareader note now names
  the twin lecture-wasm reads, not an "unadopted" pipeline.
- PLAN.md: the P4 box is ticked, with the flip and the canary proof
  (#147, runs 36654189187 and 36654357802). The Phase 5 fan-out line
  records that a consumer which executes at read time needs no dispatch.
- refresh-snapshots.yml's header no longer says no snapshot has a
  consumer, and scripts/snapshots.py's refresh PR body marks the rebuild
  dispatch as not wired yet (PLAN Phase 5) instead of implying it runs.
- CATALOG.md regenerated: 44 datasets, 44 read by lectures today.

lecture-wasm's adoption is QuantEcon/lecture-wasm#85. One placeholder
remains, to replace on the day it merges: B_MERGE_DATE, in each of the
four migration.yml repoints entries and on PLAN.md's P4 lines. It
parses as a string, so the audit and the dashboard accept it and print
it as text.

Dry-run of the strict audit (build_audit.py all --strict) over fresh
shallow clones of the eight lecture repos, on 2026-09-30. With
lecture-wasm at the adoption branch, main's landed records exit 1 with
one migration_inconsistency per file, and these repointed records exit
0. With lecture-wasm at its main, the repointed records exit 1. The
exit-0 case holds once intro's new yfinance read in mle.md
(QuantEcon/lecture-python-intro#849) is annotated, which #151 does.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant