diff --git a/AGENTS.md b/AGENTS.md index 72ce97b..214e0c4 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -4,7 +4,7 @@ Guidance for coding agents (and humans) making changes in this repository. Read ## What this repo is -The canonical home for **data consumed by the QuantEcon lecture series** (renamed from `QuantEcon/data` on 2026-07-16, per [meta#336](https://github.com/QuantEcon/meta/issues/336)). Its purpose is **stability**: it snapshots upstream sources — with attribution to each source carried in the manifest — so a lecture build never depends on a live API or a third-party host staying up. It is a **cache, not a content-distribution host**. The published tree is **flat** (`lectures/`, since 2026-07-16) and live on GitHub Pages; consumers fetch it over the raw GitHub forms below. **There is no pending host transition**: the `data.quantecon.org` custom domain was deferred indefinitely on 2026-08-12 in favor of the `qeld` consumer package (`PLAN-QELD-PACKAGE.md`, D11) — the stable interface lectures get is a package call, `qeld.url('')`, not a branded host ([#37](https://github.com/QuantEcon/data-lectures/issues/37), [#15](https://github.com/QuantEcon/data-lectures/issues/15)). The full convention lives in the draft manual page ([QuantEcon.manual#108](https://github.com/QuantEcon/QuantEcon.manual/pull/108)). +The canonical home for **data consumed by the QuantEcon lecture series** (renamed from `QuantEcon/data` on 2026-07-16, per [meta#336](https://github.com/QuantEcon/meta/issues/336)). Its purpose is **stability**: it snapshots upstream sources — with attribution to each source carried in the manifest — so a lecture build never depends on a live API or a third-party host staying up. It is a **cache, not a content-distribution host**. The published tree is **flat** (`lectures/`, since 2026-07-16) and live on GitHub Pages; consumers fetch it over the raw GitHub forms below. **There is no pending host transition**: the `data.quantecon.org` custom domain was deferred indefinitely on 2026-08-12 in favor of the `qeld` consumer package (`PLAN-QELD-PACKAGE.md`, D11) — the stable interface lectures get is a package call, `qeld.url('')`, not a branded host ([#37](https://github.com/QuantEcon/data-lectures/issues/37), [#15](https://github.com/QuantEcon/data-lectures/issues/15)). The convention's authoritative text is this file, `README.md` and `manifest-schema.yml`. A style-guide page for lecture authors is tracked in [QuantEcon.manual#153](https://github.com/QuantEcon/QuantEcon.manual/issues/153), which also settles the older draft ([QuantEcon.manual#108](https://github.com/QuantEcon/QuantEcon.manual/pull/108)); the `data-{scope}` repository type is proposed for QEP-3 in [QuantEcon/qeps#41](https://github.com/QuantEcon/qeps/issues/41). ## Rules @@ -33,9 +33,9 @@ A constructed dataset without its committed builder is a bug. Manifest fields: ` **Capture what the source gives you; never let a missing field block a useful dataset.** Rich provenance — DOI, upstream version, exact retrieval date, licence id — is always welcome and worth recording whenever it is available, because it makes the data auditable years later at almost no ongoing cost. But effort scales with what the source actually provides: where a field is genuinely unavailable, record it as an explicit, reasoned gap (see the inherited-file states below) rather than fabricating it or refusing the file. A clean, well-documented source should produce a short manifest; only genuinely messy provenance earns a long one. -#### Two inherited-file states that look like violations but are tracked, not hidden +#### Three inherited-file states that look like violations but are tracked, not hidden -The Feb 2025 migration left files that cannot fully satisfy the rules above. The manifest records each gap **explicitly** — visible in the generated catalog — rather than burying it by misclassification. Both are provisional decisions from the P1 pilot ([meta#338](https://github.com/QuantEcon/meta/issues/338)), to be folded into [manual#108](https://github.com/QuantEcon/QuantEcon.manual/pull/108). +The Feb 2025 migration left files that cannot fully satisfy the rules above. The manifest records each gap **explicitly** — visible in the generated catalog — rather than burying it by misclassification. `retrieved: null` and `unrecovered` are provisional decisions from the P1 pilot ([meta#338](https://github.com/QuantEcon/meta/issues/338)); `committed-frozen` joined them with the builders directory ([#61](https://github.com/QuantEcon/data-lectures/pull/61)). All three are to be folded into the style-guide page ([QuantEcon.manual#153](https://github.com/QuantEcon/QuantEcon.manual/issues/153)). - **`retrieved: null` — inherited-undated bytes.** `retrieved` is required, but may be `null` when the bytes were inherited (e.g. from a lecture repo) with **no recorded upstream-retrieval date**. Do **not** reconstruct one from git history — that records when QuantEcon acquired the file, not when it was retrieved from the source, and the false precision is worse than an honest null. A null `retrieved` must be paired with an `integrity.upstream` entry that **accounts for the gap** — never left bare. Any resolved status does that: `unverifiable` says the vintage cannot be established at all, and `diverged` or `verified` say something stronger, because a re-fetch that hashes identically pins the vintage by content rather than by date, which is what a `retrieved` date was a proxy for. What is not acceptable is `null` beside `unverified` — that is two unanswered questions, not one answered a different way. (Amended 2026-08-13, wave B1': the rule previously named `unverifiable` alone, which was the only case that had arisen. `life-expectancy-vs-gdp-per-capita.csv` and `mpd2020.xlsx` already paired `null` with `diverged`; B1' added five files where an upstream re-fetch succeeded byte-identically.) - **`builder_status: committed-frozen` — the builder is here, and deliberately will not run.** For a dataset built from a source that must not be refreshed: a frozen vintage, or a scraper we will not re-run. The artifact is kept as the record of what produced these bytes, so it is committed verbatim and not edited — editing it is what would destroy its value as provenance. Distinct from `committed`, which asserts a runnable four-stage builder, and from `unrecovered`, which says the builder is absent. @@ -120,7 +120,7 @@ Where one builder produces a **set** of files, name it for the set and let each `scripts/` is repo tooling — the audit dashboard and the catalog generator — and produces no dataset. Keep the two apart. -**Where a builder reads its input from.** The normal case is the third-party upstream, fetched at run time: eight of the nine `committed` builders here do that, and it is the fetch stage of the contract below. A builder reads from `sources/` **only when the input cannot be re-fetched** — the upstream is gone, unlocatable, or was inherited with no recoverable source. `sources/` is that exception layer, not a general input tree, and it is emphatically not "the big-file directory": the defining property is un-refetchability, not size. What it must never be is a network read from another QuantEcon repo — that is how a retired repo becomes load-bearing again. +**Where a builder reads its input from.** The normal case is the third-party upstream, fetched at run time: every `committed` builder here does that except the few that read `sources/`, and it is the fetch stage of the contract below. A builder reads from `sources/` **only when the input cannot be re-fetched** — the upstream is gone, unlocatable, or was inherited with no recoverable source. `sources/` is that exception layer, not a general input tree, and it is emphatically not "the big-file directory": the defining property is un-refetchability, not size. What it must never be is a network read from another QuantEcon repo — that is how a retired repo becomes load-bearing again. Builders follow four stages — **fetch → pre-process → validate → write** — and only write on validation pass. **The validate stage is shared**: `builders/_validate.py` reads the manifest's `schema` block as its spec (columns and `pattern` runs, dtype families, exact `known_nulls`, the `nulls:` placement rule, `row_count_floor`, `date_range`) and measures the overlap window against the previous vintage; a builder calls `validate(frame.reset_index(), manifest, previous)` and layers on only what the schema cannot say — value bands, a grid check, the per-series revision **bound** (a tracking snapshot is revised by its source, so the test is a tolerance plus a printed summary, never equality). The same function runs over every committed CSV on every PR (`scripts/validate_datasets.py`, `validate-datasets.yml`), so a manifest that drifts from its bytes fails the PR, not the next refresh. **A dynamic builder also exposes `check_committed()`** — its own `validate()` on the committed bytes, no network — which the same workflow runs (`--builders`) under both pandas majors; a builder-specific check that breaks on a pandas change fails the PR that introduces it (#128 was a week-old pandas-3 break the shared layer could not see). Lectures always read the last-good snapshot: an upstream outage may fail a refresh, it must never break a lecture build. @@ -193,7 +193,7 @@ scripts/ # repo tooling — NOT published, produces no dataset # a refresh, render the refresh PR body audit_annotations.yml # curated judgment for not-yet-migrated data refs migration.yml # migration lifecycle tracker (status + PR provenance per dataset) -manifest-schema.yml # per-dataset manifest schema (strawman) +manifest-schema.yml # per-dataset manifest schema — the authoritative field reference requirements.txt PLAN.md # roadmap — start here AGENTS.md # this file diff --git a/PLAN-QELD-PACKAGE.md b/PLAN-QELD-PACKAGE.md index 4a23afb..d9fe7e6 100644 --- a/PLAN-QELD-PACKAGE.md +++ b/PLAN-QELD-PACKAGE.md @@ -1,6 +1,6 @@ # PLAN — `qeld`, the consumer-side data package -**Status:** design settled, nothing implemented · **Last updated:** 2026-08-12 +**Status:** design settled, nothing implemented · **Last updated:** 2026-09-29 (facts in §5.3, §8.3–§8.6 and the Q2 row refreshed; no decision changed) **Relationship to `PLAN.md`:** that document migrates *bytes* into this repo. This one gives *consumers* a stable way to read them. They are independent — the migration completes with or without `qeld` — but the call-site convention here replaces repoint rules 5–6 for any lecture that adopts it. @@ -273,8 +273,11 @@ The best diff in the corpus is `lecture-python-advanced.myst/lectures/hansen_jag ### 5.3 `url()` alone covers the corpus -After the `.npy` conversion (§8.3), **`dataBHS.mat` is the only file in the endgame that cannot be read from -a URL**. That is what justifies dropping `fetch()`. +After the `.npy` conversion (§8.3), **every file in the endgame can be read from a URL**: `dataBHS.mat`, the +one exception when this was written, was converted to `dataBHS.csv` at migration +([#98](https://github.com/QuantEcon/data-lectures/pull/98), 2026-08-18; §8.4). That is what justifies +dropping `fetch()`. Until the conversion, the `.npy` pair is the only published format that cannot be passed +straight to a reader function — `np.load` needs `requests` plus `BytesIO`. ### 5.4 Two latent bugs found, worth fixing regardless @@ -333,7 +336,7 @@ or every migrated read classifies `local-path` and the dashboard inverts. | phase | work | gate | |---|---|---| | **Q1 — Audit first** | `build_audit.py` learns `qeld.url('X')` → pattern `qeld`, counted migrated **and terminal**. For `pattern == 'qeld'`, assert the key exists in `lectures/` and is not deprecated — otherwise the qeld path loses every assertion #55/#48/#47 added. `migration.yml`: `final` := every code read via qeld, with §4.1 carve-outs terminal on the direct form (the canonical-host arm was retired with D11) | `audit.json` `stats` and `problems` unchanged on today's repos (**not** "byte-identical" — the audit stamps `date.today()`) | -| **Q2 — Schema hygiene** | Document `read_as` (used in 6 manifests) and `sheets` (5) in `manifest-schema.yml` — both are in use and neither appears in the file `AGENTS.md` calls "the authoritative, commented field reference". Add `deprecated:` (new, used nowhere yet) since §3.3 warns on it. `shape` is already documented. Delete `then: "iloc[1:]"` from `longprices.xls.yml:70` by moving `iloc[1:]` into the lecture — a post-read transform encoded as a string to evaluate is exactly what D5 excludes | `manifest-schema.yml` covers every field any manifest uses. Needs none of #14's decisions — do not block on it | +| **Q2 — Schema hygiene** | Document `read_as` (used in 6 manifests) and `sheets` (5) in `manifest-schema.yml` — both are in use and neither appears in the file `AGENTS.md` calls "the authoritative, commented field reference". Add `deprecated:` (new, used nowhere yet) since §3.3 warns on it. `shape` is already documented. Delete `then: "iloc[1:]"` from `longprices.xls.yml:70` by moving `iloc[1:]` into the lecture — a post-read transform encoded as a string to evaluate is exactly what D5 excludes | `manifest-schema.yml` covers every field any manifest uses. Needs none of #14's decisions — do not block on it. *(2026-09-29: `read_as` and `sheets` are now described there; `deprecated:` and the `then:` cleanup remain.)* | | **Q3 — Package** | `packages/qeld/`: `url()`, `info()`, context detection, advisory catalog. Catalog compiler shares a freshness gate with `CATALOG.md`. Format tier-1 assertion. First release to PyPI via trusted publishing | Offline suite green on every PR: catalog compiles and is fresh; unknown key warns and still returns a URL; URL form correct per detected context; suffix fidelity incl. `.csv.gz`; `info()` fields present. CPython matrix | | **Q4 — Live leg** | Post-merge + scheduled job: fetch each served URL, compare to the manifest hash, open an issue on failure | Green on `main`; an induced failure opens an issue | | **Q5 — Browser session** | `%pip install qeld==` in a real `lecture-wasm` page (**`%pip` routes through piplite, not micropip** — a console `micropip.install` is a false pass); `pd.read_excel(qeld.url('mpd2020.xlsx'), sheet_name='Regional data', header=[0,1,2], index_col=0)`; a `.csv.gz` read; record observed Pyodide and pyodide-kernel versions | Written pass/fail. Fail ⇒ wasm keeps URLs and the plan proceeds for the CPython repos | @@ -374,29 +377,39 @@ columns (`date`/`specie_value`, `date`/`nominal_balances`). Converting to CSV de and two imports from `french_rev` in every consuming repo. **But the window closed.** When this was analysed both files had `consumers: []`; the A3 set has since been -repointed (#49) and both now have two consumers. So this is no longer a free replacement — it needs the +repointed (#49), and both now have six consumers each (intro, wasm, `lecture-intro.zh-cn`, the actions +canary, `tom-econ370-2025` and `python-lecture-sandpit.myst`, per the manifests on 2026-09-29). So this is no +longer a free replacement — it needs the `AGENTS.md` "new vintage → new filename" treatment (`caron.csv` lands alongside, consumers opt in, the `.npy` is swept later) or a coordinated set under repoint rules 1–3. **Decide before Q6**, since `french_rev` is a pilot. ### 8.4 `dataBHS.mat` — convert at migration, or exclude? +**Resolved 2026-08-18: converted at migration.** `dataBHS.csv` landed in +[#98](https://github.com/QuantEcon/data-lectures/pull/98) and the lecture reads it; the `.mat` stays as the +builder's input in `sources/`. The analysis that led there, as written: + 5,588 bytes; `c`, `rb`, `rs`, each (236, 1) float64; the lecture uses only `data['c']` and the read is inside a `hide-input` cell, so nothing about it is taught. Trivially a 236×3 CSV — but see §4.3 on the read. A Track C decision; the only true impossibility among static files. ### 8.5 Is `lecture-intro.zh-cn` in scope? -It carries data reads, appears in **zero** `consumers` blocks, is excluded from `SCAN_REPOS` by decision, has -no data CI, publishes on a `publish*` tag, and inherits install cells automatically via the `.md`-only sync — +It carries data reads, is excluded from `SCAN_REPOS` by decision, has no data CI, publishes on a `publish*` tag, and inherits install cells automatically via the `.md`-only sync — so it acquires whatever intro acquires without anyone deciding. It also has files with no data-lectures key -and no business having one (`country_code_cn.csv`, a translation asset). +and no business having one (`country_code_cn.csv`, a translation asset). When this was written it appeared +in no `consumers` block; the Track A and P3 repoints have since recorded it in 25 `consumers` entries (as of +2026-09-29), so its reads are now in the manifests even though no CI sees them. **Recommendation: explicit non-goal for v1, with one fixed rule instead of machinery — any sweep touching an intro file also touches zh-cn.** ### 8.6 The rename list for generic filenames -§5.5. Needs a pass before Tracks B and C migrate, and each rename needs its prose pairing found by grep. +§5.5. **Settled.** The list moved to [#87](https://github.com/QuantEcon/data-lectures/issues/87) on +2026-08-17, and the naming policy ([#113](https://github.com/QuantEcon/data-lectures/issues/113), +2026-09-07) decided it: a consumed file keeps its name (`manifest-schema.yml` naming rule 6), and `fp.dta`, a +verbatim release, keeps its upstream name (rule 5). Tracks B and C migrated under their current names. ### 8.7 Open from the original report diff --git a/PLAN.md b/PLAN.md index 2a0a190..f19adf8 100644 --- a/PLAN.md +++ b/PLAN.md @@ -1,25 +1,27 @@ # PLAN — `data-lectures` (formerly `QuantEcon/data`) -**Status:** active roadmap (last updated 2026-09-01) — **the repo is LIVE**: the first repoint merged 2026-07-17 (P1, `lingcod_msy_recovery.csv` → `msy_fishery`), so published filenames are an API from here on +**Status:** active roadmap (last updated 2026-09-29) — **the repo is LIVE**: the first repoint merged 2026-07-17 (P1, `lingcod_msy_recovery.csv` → `msy_fishery`), so published filenames are an API from here on -**Where the numbers stand (`audit.json`, 2026-08-31):** **40 of 40 static datasets migrated and repointed — the static migration completed 2026-08-18** ([#98](https://github.com/QuantEcon/data-lectures/pull/98), [#99](https://github.com/QuantEcon/data-lectures/pull/99)); 23 lectures still fetch live API data (Track E); 24 committed orphans (Track X); 0 legacy-repo references; 2 URL forms in use. +**Where the numbers stand (`audit.json`, 2026-09-28):** **40 of 40 static datasets migrated and repointed — the static migration completed 2026-08-18** ([#98](https://github.com/QuantEcon/data-lectures/pull/98), [#99](https://github.com/QuantEcon/data-lectures/pull/99)); 23 lectures still fetch live API data (Track E); 1 committed orphan, kept deliberately (`python_advanced_features/test_table.csv`, an exercise download — the Track X sweep completed 2026-09-01); 0 legacy-repo references; 2 URL forms in use. -**`CATALOG.md` and `migrated` both say 40 today, and that agreement holds only while no wave is in flight.** They count different things: `migrated` counts datasets whose *consumers* read this repo; the catalog counts datasets that *live* here. They agree only when no wave is in flight. The P3 fold is the worked example — the six `high_dim_data` files landed 2026-08-10 at `status: landed` with `consumers: []`, opening a six-file gap that closed on 2026-08-11 when PR set C repointed them and [#69](https://github.com/QuantEcon/data-lectures/pull/69) flipped the records. Expect that gap again for the duration of any wave that lands ahead of its repoints, which is the convention here. +**`CATALOG.md` counts more datasets than `migrated` today, and the gap is expected.** They count different things: `migrated` counts datasets whose *consumers* read this repo; the catalog counts datasets that *live* here. They agree only when no wave is in flight. Today's gap is the four Track E `business_cycle` snapshots, landed 2026-09-01 at `consumers: []` ahead of their first consumer, plus any new dataset landed ahead of the lecture that will read it. The P3 fold is the worked example — the six `high_dim_data` files landed 2026-08-10 at `status: landed` with `consumers: []`, opening a six-file gap that closed on 2026-08-11 when PR set C repointed them and [#69](https://github.com/QuantEcon/data-lectures/pull/69) flipped the records. Expect that gap again for the duration of any wave that lands ahead of its repoints, which is the convention here. Every figure on that line comes from `stats` in the generated `audit.json`, and every figure below that restates one is a copy that can drift — as all seven of them had by 2026-08-07, each understating progress by three repoint sets. Re-read them from `audit.json` before quoting them, and prefer citing the generated file over this document. -This repository is being shaped into the **single canonical repository for data consumed by the QuantEcon lecture series**, referenced by stable URLs and documented in the manual. +This repository is the **single canonical repository for data consumed by the QuantEcon lecture series**, referenced by stable URLs. Documenting it for lecture authors in the manual's style guide is tracked in [QuantEcon.manual#153](https://github.com/QuantEcon/QuantEcon.manual/issues/153). ## Governing threads | Thread | Role | | --- | --- | | [QuantEcon/meta#336](https://github.com/QuantEcon/meta/issues/336) | Design discussion / future QEP — the convention itself | -| [QuantEcon/data#8](https://github.com/QuantEcon/data/issues/8) | Scaffolding checklist for **this repo** (this PLAN executes it) | -| [QuantEcon/meta#337](https://github.com/QuantEcon/meta/issues/337) | Live hosting risks + `high_dim_data` shape-up + orphan sweep in lecture repos | -| [QuantEcon/meta#338](https://github.com/QuantEcon/meta/issues/338) | Pilot: migrate one dataset per hosting pattern, landing here | -| [QuantEcon/QuantEcon.manual#108](https://github.com/QuantEcon/QuantEcon.manual/pull/108) | Draft `styleguide/datasets.md` — the convention's design surface | -| [QuantEcon/data#1](https://github.com/QuantEcon/data/issues/1), [#2](https://github.com/QuantEcon/data/issues/2), [#4](https://github.com/QuantEcon/data/issues/4) | Pre-existing execution items (LFS, fold in `high_dim_data`, repoint lectures) | +| [QuantEcon/data-lectures#8](https://github.com/QuantEcon/data-lectures/issues/8) | Scaffolding checklist for **this repo** — closed 2026-08-31, delivered | +| [QuantEcon/meta#337](https://github.com/QuantEcon/meta/issues/337) | Live hosting risks + `high_dim_data` shape-up + orphan sweep in lecture repos — closed 2026-08-06 | +| [QuantEcon/meta#338](https://github.com/QuantEcon/meta/issues/338) | Pilot: migrate one dataset per hosting pattern, landing here — P1–P3 complete, P4 (the dynamic-snapshot twin) in progress | +| [QuantEcon/QuantEcon.manual#108](https://github.com/QuantEcon/QuantEcon.manual/pull/108) | Draft `styleguide/datasets.md` — the convention's original design surface, unchanged since 2026-08-10 and so predating D11; its fate is decided under manual#153 | +| [QuantEcon/QuantEcon.manual#153](https://github.com/QuantEcon/QuantEcon.manual/issues/153) | The style-guide section for lecture authors: a pointer to this repo first, then manual#108 refreshed and merged, or closed | +| [QuantEcon/qeps#41](https://github.com/QuantEcon/qeps/issues/41) | Registering the `data-{scope}` repository type in QEP-3, which does not list it today | +| [QuantEcon/data#1](https://github.com/QuantEcon/data/issues/1), [#2](https://github.com/QuantEcon/data/issues/2), [#4](https://github.com/QuantEcon/data/issues/4) | Pre-existing execution items (LFS, fold in `high_dim_data`, repoint lectures) — all closed 2026-08-31 | | [QuantEcon/workspace-lectures#14](https://github.com/QuantEcon/workspace-lectures/issues/14) | Standing cross-repo tracker for the migration — tracks A–E and X–Y, repoint rules, outstanding write-backs (it began as the pilot's session work plan) | ## Where we are @@ -70,7 +72,7 @@ Sweep by cloning and grepping, not with `gh search code` on a URL — code searc The strict audit has **no green state for a partially-repointed dataset**. `scripts/build_audit.py` fails a record marked `pending`/`landed` while any consumer already reads data-lectures, *and* fails one marked `repointed`/`final` while any consumer still does not. That is deliberate — it is what makes the tracker trustworthy — but it means a dataset with two consuming repos cannot be moved one repo at a time without the drift alarm firing in the gap. -**2 of the 17 remaining datasets have two consuming repos, and both are `lecture-python-intro` + `lecture-wasm`** — `life-expectancy-vs-gdp-per-capita.csv` and `usa-gini-nwealth-tincome-lincome.csv`. (Six of the eight this line used to name were the `high_dim_data` files, repointed 2026-08-11.) There is no other cross-series coupling left; the last one was the P2 `pandas_panel` trio, already done. (`graph.txt` is **not** in this count and is no longer a Track A item at all — `lecture-wasm` used to read intro's committed copy, which is what made it look like a one-consumer dataset; that read is gone and it is now embedded in every consuming lecture. See the Track A row below.) +**While Track A was in flight, 2 of the 17 remaining datasets had two consuming repos, both `lecture-python-intro` + `lecture-wasm`** — `life-expectancy-vs-gdp-per-capita.csv` and `usa-gini-nwealth-tincome-lincome.csv`, repointed together in wave A4 (2026-08-12); no static dataset remains to migrate. (Six of the eight this line used to name were the `high_dim_data` files, repointed 2026-08-11.) There is no other cross-series coupling left; the last one was the P2 `pandas_panel` trio, already done. (`graph.txt` is **not** in this count and is no longer a Track A item at all — `lecture-wasm` used to read intro's committed copy, which is what made it look like a one-consumer dataset; that read is gone and it is now embedded in every consuming lecture. See the Track A row below.) **"Two consuming repos" is the audit's count, not the consumer set.** `SCAN_REPOS` is the eight Python-family repos, so a dataset the dashboard shows with two consumers may have four or five in reality. Measured 2026-08-12, both remaining pairs are read by **five** reference-holders each: intro, `lecture-wasm`, `lecture-intro.zh-cn`, `QuantEcon/test-actions-lecture-intro`, and the generated `lecture-python-intro.notebooks` mirror. The last three are invisible to every audit run — see rule 1. @@ -111,8 +113,8 @@ Adopting a newer upstream vintage is a *different change* with a different risk So when a migration finds that the committed file differs from what upstream publishes today: 1. **Migrate what the lectures use**, unchanged, with the byte-compare gate as normal. -2. **Record the delta** in the dataset's manifest (`integrity.upstream`) *and* in the register at [#39](https://github.com/QuantEcon/data-lectures/issues/39) — the manifest makes it visible in the catalog from day one, the register is where it gets reasoned about. -3. **Review the register once the migration completes**, and decide each case on its merits. +2. **Record the delta** in the dataset's manifest (`integrity.upstream`) — the manifest is the register of record, and the catalog shows the delta from the day it is found. During the migration a separate register at [#39](https://github.com/QuantEcon/data-lectures/issues/39) carried it too; since #39 closed on 2026-09-07 a new delta needs no issue (AGENTS.md). +3. **Decide each case on its merits, after the migration** — the migration-era review ran on 2026-09-01 (#39). Adopting a newer vintage changes lecture output, so it is an author's call and gets a new filename. Two deltas look alike and need opposite responses. *Upstream moved* — a newer vintage exists; adopting it means a **new filename**, per "Corrections vs vintages" in `AGENTS.md`, so consumers opt in. *Our copy diverges* — upstream is unchanged but our file was modified; resolving means reconciling the edit. `mpd2020.xlsx` is the first recorded instance of the second kind, and it is instructive: the local edits are load-bearing for the consuming lecture, so the file and the lecture have to move together. @@ -217,8 +219,8 @@ The remaining work decomposes by **consuming series** rather than by hosting pat | **B — `python.myst`** | 7, **all done**, cut into two waves. **B1′**: the `ols` trio, `fp.dta` and `NEWQDATA.csv` landed in [#79](https://github.com/QuantEcon/data-lectures/pull/79), flipped in [#80](https://github.com/QuantEcon/data-lectures/pull/80). **B2′**: `hansen_singleton_1982/1983_data.csv` landed in [#82](https://github.com/QuantEcon/data-lectures/pull/82), flipped in [#83](https://github.com/QuantEcon/data-lectures/pull/83); both waves independently validated in [#84](https://github.com/QuantEcon/data-lectures/issues/84) | **three consumers, not one** — `lecture-python.zh-cn` reads by URL *and* holds byte-identical copies of all 7 plus both builders (and is outside `SCAN_REPOS`, so the audit cannot see it); `lecture-python.notebooks` lags a publish tag; `lecture-stats` carried a published-site prose link to `fp.dta` behind a daily linkcheck. B2′ adds a fourth kind: the two builders migrate too, and each lecture names them twice outside its data cell | nothing | | **C — `advanced.myst`** | 6, **all done**. **C1**: the `bbh` pair and `hansen_jagannathan_1991_data.json` landed in [#92](https://github.com/QuantEcon/data-lectures/pull/92), flipped in [#95](https://github.com/QuantEcon/data-lectures/pull/95), validated in [#96](https://github.com/QuantEcon/data-lectures/issues/96). **C2**: `fred_data.csv`, `acs_data_summary.csv` and `dataBHS.mat` (converted to `dataBHS.csv`) landed in [#98](https://github.com/QuantEcon/data-lectures/pull/98), flipped in [#99](https://github.com/QuantEcon/data-lectures/pull/99), validated in [#100](https://github.com/QuantEcon/data-lectures/issues/100) | none by URL — but six other org repos hold byte-identical copies of `acs_data_summary.csv` and `dataBHS.mat`, and `lecture-tools-techniques` publishes its own read of the latter, so acceptance was scoped to advanced's own URLs | — | | **D — `programming`** | 1, **done**: `test_pwt.csv` rode with wave C2 ([#98](https://github.com/QuantEcon/data-lectures/pull/98), [#99](https://github.com/QuantEcon/data-lectures/pull/99)); its four consuming repos are recorded in [#101](https://github.com/QuantEcon/data-lectures/pull/101) | none | — | -| **E — dynamic / live-API** | the UNRATE twin, then the 15 incidental API lectures | wasm is the forcing customer | [#14](https://github.com/QuantEcon/data-lectures/issues/14) schema decisions, [#26](https://github.com/QuantEcon/data-lectures/issues/26) fetch layer | -| **X — orphan sweep** | **done 2026-09-01** — 23 of the 24 `audit.json` orphans deleted (dp 10, programming 4, intro 4, python.myst 2, wasm 2, `continuous_time_mcs` 1), `python_advanced_features/test_table.csv` kept as an exercise download, plus 46 translation copies the audit cannot see (`lecture-python.zh-cn` 15, `lecture-python-programming.{zh-cn,fr,fa}` 9 each, `.ml` 6, `lecture-intro.zh-cn` 4); twelve PRs, ledger at [QuantEcon/workspace-lectures#57](https://github.com/QuantEcon/workspace-lectures/issues/57) | — | site clearance rides the settle policy, verified at [QuantEcon/workspace-lectures#40](https://github.com/QuantEcon/workspace-lectures/issues/40) | +| **E — dynamic / live-API** | the `business_cycle` set first (P4, reframed 2026-09-01 from a lone UNRATE twin): four dynamic snapshots, landed 2026-09-01 at `consumers: []` ([#109](https://github.com/QuantEcon/data-lectures/pull/109), [#114](https://github.com/QuantEcon/data-lectures/pull/114)); then candidates among the other live-API lectures (23 in the audit) | `lecture-wasm` is the forcing customer; `lecture-python-intro` keeps its live calls, because there the API is the lesson (decided 2026-09-01) | nothing — the schema decisions ([#14](https://github.com/QuantEcon/data-lectures/issues/14)) landed 2026-09-07 and the fetch layer ([#26](https://github.com/QuantEcon/data-lectures/issues/26), `builders/_fred.py`) 2026-09-01. Next: the `lecture-wasm` adoption ([QuantEcon/lecture-wasm#70](https://github.com/QuantEcon/lecture-wasm/issues/70)), then the flip; session plan [#118](https://github.com/QuantEcon/data-lectures/issues/118) | +| **X — orphan sweep** | **done 2026-09-01** — 23 of the 24 `audit.json` orphans deleted (dp 10, programming 4, intro 4, python.myst 2, wasm 2, `continuous_time_mcs` 1), `python_advanced_features/test_table.csv` kept as an exercise download, plus 46 translation copies the audit cannot see (`lecture-python.zh-cn` 15, `lecture-python-programming.{zh-cn,fr,fa}` 9 each, `.ml` 6, `lecture-intro.zh-cn` 4); twelve PRs, ledger at [QuantEcon/workspace-lectures#57](https://github.com/QuantEcon/workspace-lectures/issues/57). **One remainder, found 2026-09-29:** `lecture-intro.zh-cn` still commits 10 byte-identical copies of Track A datasets — the eight under `lectures/datasets/` plus the `usa-gini` and `life-expectancy` CSVs under `_static/`, which its site serves — because Track A's deletions were never mirrored there (only P5's were). Every one of its reads points at this repo, so all 10 are orphans awaiting a deletion PR | — | site clearance rides the settle policy, verified at [QuantEcon/workspace-lectures#40](https://github.com/QuantEcon/workspace-lectures/issues/40) | | **Y — consumer interface (`qeld`)** | the `qeld` package, Q1–Q7 of `PLAN-QELD-PACKAGE.md` — audit support, the package, pilots, then adoption by win; QEP graduation stays | — | nothing — re-scoped 2026-08-12 (D11): the DNS → custom domain → URL-sweep sequence this row used to carry is retired | **`graph.txt` was closed out as a non-migration (2026-08-12).** It is synthetic teaching data — `provenance: toy`, null in every real provenance field — and the shortest-path exercise teaches its format by quoting the first line, so the data has to stay visible on the page. Hosting it here would have put a toy in a registry that exists to carry provenance. Instead `lecture-wasm` stopped fetching intro's committed copy over the network and embeds it with `%%file` like every sibling ([QuantEcon/lecture-wasm#63](https://github.com/QuantEcon/lecture-wasm/pull/63)), which retired the last cross-repo read of that blob anywhere in the organisation. `graph.txt` consequently no longer appears as a scanned dataset at all. Four repos embed it via `%%file` — intro, dp, jax and wasm — and two of those (intro, dp) also commit a copy the cell overwrites before reading, so those two are shadowed orphans; jax and wasm commit none, which is the cleaner shape. The remaining committed copies (`lecture-intro.zh-cn`, the canary, `lecture-python.zh-cn`, `lecture-dp.monorepo`, `ipynb_pdf_constructor`) are read by nothing. Intro's committed copy is now deletable as Track X — but the same blob sits in **7** repos byte-identically (a further two, `QuantEcon.jl` and `QuantEcon.lectures.code`, hold a 4,692-byte variant differing by one trailing space) and is regenerated at 17 `%%file` sites, including archived `.rst` ancestors that `gh search code` cannot see, so that deletion needed its own per-repo reader sweep rather than an org-wide sweep. **Deleted 2026-09-01** from intro, dp, `lecture-intro.zh-cn` and `lecture-python.zh-cn` (Track X); the one outside-org reader, `devopseng99/project.lecture-wasm`, was recorded and accepted on QuantEcon/workspace-lectures#57. The copies in the canary, `lecture-dp.monorepo`, `ipynb_pdf_constructor`, `QuantEcon.jl` and `QuantEcon.lectures.code` are read by nothing and stay. @@ -265,9 +267,9 @@ Only one file genuinely forces LFS, and it is not a dataset: | Path | Contents | Storage | Served? | | --- | --- | --- | --- | | `lectures/` | every published dataset, including both SCF minis and the 4 `cross_section` CSVs | **plain git** | yes | -| `sources/` | upstream inputs that builders consume but no lecture reads — `SCF_plus.dta` (99.1 MiB) | **per-path LFS** | **no** | +| `sources/` | upstream inputs that builders consume but no lecture reads — `SCF_plus.dta` (99.1 MiB), and since Tracks B and C the small `NEWQDATA.MAT` and `dataBHS.mat` | **per-path LFS** | **no** | -`sources/` exists as of [#63](https://github.com/QuantEcon/data-lectures/pull/63). Its defining property is **un-refetchability, not size** — all six `committed` builders here fetch from their third-party upstream at run time and none has a committed input, so this is an exception layer rather than a general input tree. Confirmed not served, 2026-08-10: `quantecon.github.io/data-lectures/sources/SCF_plus.dta` and `.../sources/README.md` both return **404**, against **200** for any `lectures/` path. +`sources/` exists as of [#63](https://github.com/QuantEcon/data-lectures/pull/63). Its defining property is **un-refetchability, not size** — every `committed` builder here fetches from its third-party upstream at run time except the few whose input cannot be re-fetched (`dataBHS.py` and `NEWQDATA.py` read `sources/`), so this is an exception layer rather than a general input tree. Confirmed not served, 2026-08-10: `quantecon.github.io/data-lectures/sources/SCF_plus.dta` and `.../sources/README.md` both return **404**, against **200** for any `lectures/` path. - [x] Add `sources/` for builder inputs, with `sources/README.md` as the **audit trail**: a **`## ` section per committed file**, recording where it came from, when, its licence, the upstream identifier (DOI where one exists), its `sha256`, and which builder consumes it. That shape is **enforced, not conventional** — `check_sources()` splits the README on `## ` headings and reads the first 64-hex token in each filename-shaped section, so a file documented as a row in a shared table has no recorded hash and fails the required check. (This line previously said "one row per committed file", which was written before the format existed and would now send you into a red build.) A file in `sources/` is not a published dataset and gets no sidecar manifest — this README is its provenance record. Landed in [#63](https://github.com/QuantEcon/data-lectures/pull/63), which also made the README **load-bearing rather than documentary**: `check_consumed_files.py` now asserts that every file here is captured by the LFS rule and hashes to a `sha256` recorded under a `## ` heading, and fails on a README entry naming a file that is not there. It reads the pointer's `oid` rather than the object, so it verifies ~100 MiB under `lfs: false` at zero bandwidth - [x] Per-path LFS via `.gitattributes`, scoped to `sources/**` only — never a blanket rule like `high_dim_data`'s `*.csv` **and** `*.dta` (data#1). Landed in [#57](https://github.com/QuantEcon/data-lectures/pull/57), with `sources/README.md` excluded so the audit trail stays readable text @@ -284,7 +286,7 @@ Only one file genuinely forces LFS, and it is not a dataset: - [x] GitHub Pages deploy of the published tree, **`lfs: false` at checkout** (inverted by [#57](https://github.com/QuantEcon/data-lectures/pull/57): a mis-tracked file must publish as its pointer, so the mistake is visible rather than masked) — landed 2026-07-17 with the audit dashboard (`.github/workflows/audit-dashboard.yml`, [#20](https://github.com/QuantEcon/data-lectures/issues/20)): the default `quantecon.github.io/data-lectures/` site serves the dashboard at `/` and the published tree at `/lectures/`. The custom domain below stays open - [x] ~~`data.quantecon.org` DNS + custom domain~~ — **deferred indefinitely 2026-08-12, do not do this without revisiting D11** (`PLAN-QELD-PACKAGE.md` §2.2): the `qeld` package is the stable consumer interface instead of a branded host, and the direct raw URLs are standing rather than interim. The measurement stands for whenever this is revisited: as of 2026-08-10 the name is NXDOMAIN at `quantecon.org`'s own authoritative nameserver, the repo's Pages `cname` is null, and the stale A record and the AWS box it pointed at are gone — so reviving it is two actions QuantEcon controls (create the record, set the custom domain), plus D11's condition that old wheels keep working because qeld's base never sat on `quantecon.github.io`. [#37](https://github.com/QuantEcon/data-lectures/issues/37) stays open as the deferred tracker - [x] Verify `access-control-allow-origin: *` on served files (pyodide/JupyterLite, meta#143) — **verified 2026-08-06**: `quantecon.github.io/data-lectures/lectures/lingcod_msy_recovery.csv` returns `access-control-allow-origin: *`. The requirement is met on the default Pages domain today and never waited on a custom domain; re-verify only if DNS is ever revisited (D11) -- [ ] Monitor Pages soft limits (~1 GB site, 100 GB/month) +- [ ] Monitor Pages soft limits (~1 GB site, 100 GB/month) — measured 2026-09-29: `lectures/` is 113.5 MiB, about 11% of the site limit; bandwidth is not observable through the API ### Phase 5 — Automation (`.github/`) @@ -305,9 +307,9 @@ Full automation: ### Phase 6 — Metadata backfill for existing holdings -- [x] Manifest per dataset for the files now in `lectures/`: source, license, retrieval date, schema, consumers, provenance class. Schema sketched in `manifest-schema.yml` (Phase 2); backfill is per-file work gated on the license check below. **Complete 2026-09-01** — 41 datasets, 41 manifests; the last was `business_cycle_data.csv`, and the two `business_cycle` `.md` dumps left `lectures/` for `provenance/` the same day (`business_cycle_info.md` and `business_cycle_metadata.md` are prose, not datasets, so the gap is 1 file and not 3) +- [x] Manifest per dataset for the files now in `lectures/`: source, license, retrieval date, schema, consumers, provenance class. Schema sketched in `manifest-schema.yml` (Phase 2); backfill is per-file work gated on the license check below. **Complete 2026-09-01** — 41 datasets, 41 manifests at that point (the `business_cycle` set added three the same day, [#114](https://github.com/QuantEcon/data-lectures/pull/114), each landing with its manifest); the last backfill was `business_cycle_data.csv` (now `gdp_growth_annual.csv`), and the two `business_cycle` `.md` dumps left `lectures/` for `provenance/` the same day (`business_cycle_info.md` and `business_cycle_metadata.md` are prose, not datasets, so the gap is 1 file and not 3) - [x] Classify: the 8 static intro files are author-assembled or verbatim; `business_cycle_data.csv` is the one dynamic snapshot — `class: dynamic-snapshot`, `cadence: annual`, declared 2026-09-01 -- [ ] Licence check **per source**, not per file: the question is *"may this source be cached and served publicly, with attribution?"* — a cheap binary gate (`redistribution: permitted | restricted`, see AGENTS.md "Licensing and attribution"), a fast yes for public data sources. Two sources already answered: World Bank is **CC BY-4.0** (`business_cycle_metadata.md`, the model for what a manifest should capture) and RAM Legacy is **CC BY 4.0** (established against its Zenodo DOI record, P1). The remaining sources need the equivalent established by hand +- [x] Licence check **per source**, not per file — **answered for every manifest (verified 2026-09-29): each records `redistribution` as `permitted` or `restricted`, with the `verified` date it was established; the restricted and caveated files are the standing review on [#35](https://github.com/QuantEcon/data-lectures/issues/35), which the paragraph below defers.** The original requirement: the question is *"may this source be cached and served publicly, with attribution?"* — a cheap binary gate (`redistribution: permitted | restricted`, see AGENTS.md "Licensing and attribution"), a fast yes for public data sources. Two sources already answered: World Bank is **CC BY-4.0** (`business_cycle_metadata.md`, the model for what a manifest should capture) and RAM Legacy is **CC BY 4.0** (established against its Zenodo DOI record, P1). The remaining sources need the equivalent established by hand. **Licensing does not gate migration** (settled 2026-08-06, [#35](https://github.com/QuantEcon/data-lectures/issues/35)). Inherited data — anything the lecture repos already serve publicly — migrates with its licence recorded **as found**, including `redistribution: restricted` and `name: null` where that is the honest answer. Moving the same bytes to a canonical host with better provenance and an explicit licence field improves on the status quo, so the migration does not wait on review; what needs further thought is tracked in [#35](https://github.com/QuantEcon/data-lectures/issues/35) with alternatives, and resolved before this repo is ever promoted as a branded public open-data host. That promotion is the gate, not each file's move — and with the custom domain deferred indefinitely (2026-08-12, D11), no such promotion is scheduled: the #35 inventory stays open and the gate binds only if a public host is someday established after all. This generalises the exception AGENTS.md already carried for `countries.csv`, and applies to **inherited** data only — a genuinely new dataset still establishes its licence before it lands - [x] Keep-or-drop decision for the files with no consumer anywhere — **dropped 2026-07-16** in the Phase 2 restructure, rather than promoting them into the published namespace: @@ -319,11 +321,11 @@ Full automation: Verify that what this repo holds is actually the data it claims to be — against upstream sources, and against the copies lectures consume today — before any lecture is repointed here. -- [ ] **Byte-compare against the in-use copies**: each file migrated in Feb 2025 must be identical to the copy `lecture-python-intro` currently consumes (git blob hash compare). If a copy diverged, a repoint silently changes lecture output — this check is a hard prerequisite for Phase 8. Recorded **in the repoint PR** as a one-time gate, reproducible later from the manifest's `sha256` — not a manifest field (P1 decision) -- [ ] **Verbatim files**: re-fetch from the upstream source and compare (e.g. `mpd2020.xlsx` against the published Maddison Project 2020 release); record `sha256`, `status`, what it was compared `against`, and the date in the manifest's `integrity.upstream` +- [x] **Byte-compare against the in-use copies** — **met (reviewed 2026-09-29)**: recorded in the repoint PRs for the eight Feb 2025 files (QuantEcon/lecture-python-intro#823, #824, #826) and in the landing or repoint PR of 38 of the 40 repointed datasets; the other two (`epl_match_goals.csv`, `japan_earthquakes.csv`) were new, with no earlier copy to compare. The original requirement: each file migrated in Feb 2025 must be identical to the copy `lecture-python-intro` currently consumes (git blob hash compare). If a copy diverged, a repoint silently changes lecture output — this check is a hard prerequisite for Phase 8. Recorded **in the repoint PR** as a one-time gate, reproducible later from the manifest's `sha256` — not a manifest field (P1 decision) +- [x] **Verbatim files** — **met: every one has a resolved `integrity.upstream` status (2026-09-29: 4 `verified`, 7 `unverifiable`, 1 `diverged`), and no manifest of any class is `unverified`.** Re-fetch from the upstream source and compare (e.g. `mpd2020.xlsx` against the published Maddison Project 2020 release); record `sha256`, `status`, what it was compared `against`, and the date in the manifest's `integrity.upstream` - [x] **Constructed / dynamic files**: re-run the committed builder (`builders/business_cycle.py` → `business_cycle_data.csv`) and confirm values agree in the overlap window with the committed snapshot — **done 2026-09-01, and they do not agree, by design**: the World Bank revised 236 of 320 overlap cells (max 1.5 pp) and appended two years. Recorded as `diverged` / `upstream-moved` in the manifest and in the register at [#39](https://github.com/QuantEcon/data-lectures/issues/39); the finding is what set the builder's 5 pp overlap bound -- [ ] **Author-assembled files** (the French Revolution spreadsheets, `caron.npy`, `nom_balances.npy` — prose-only provenance): spot-check key values against the cited publication and record what was checked; full verification may be impossible, and the manifest should say so (`status: unverifiable` with a one-line `note` — the honest known status, per P1) -- [ ] **Unverifiable or failing files**: flag in the manifest and open an issue — do not promote a file to the canonical URL namespace with a known-bad or unknown integrity status +- [ ] **Author-assembled files** (the French Revolution spreadsheets, `caron.npy`, `nom_balances.npy` — prose-only provenance): spot-check key values against the cited publication and record what was checked; full verification may be impossible, and the manifest should say so (`status: unverifiable` with a one-line `note` — the honest known status, per P1). **Status 2026-09-29:** all seven are recorded `unverifiable`; `caron.npy` and `nom_balances.npy` record checks made on 2026-08-06, and the other five notes say a spot-check is possible but none is recorded yet +- [x] **Unverifiable or failing files**: flag in the manifest — do not promote a file to the canonical URL namespace with a known-bad or unknown integrity status. **Met, under the rule as amended 2026-09-07**: the manifest is the register of record and the catalog shows its status, so no issue is opened (AGENTS.md); no manifest is `unverified`, and every `unverifiable` or `diverged` one carries a note ### Phase 8 — Pilot deployment (meta#338) @@ -341,23 +343,23 @@ The first end-to-end deployment: one dataset per hosting pattern, each the harde Five things P3 proved that were not on its test list. A `constructed` dataset's builder must land in the **same** PR as the data, because `check_consumed_files.py` asserts the `builder:` path resolves. `builders/README.md`'s coverage table is a real coverage report and goes stale silently. The plain-git decision costs ~10 MB of packed history for 110 MB of working tree, since CSV compresses 5-22×. The **C0 → C1 → C2 ordering worked and proved less than it looks like** — the sync PR it was designed to defuse (QuantEcon/lecture-intro.zh-cn#293) touched zero data-read lines, zero `# i18n` markers and zero protected localisations, but nothing ever asked the model to rewrite those cells, so the markers remain unexercised, prompt-level protection and **the hand-diff is what protects a localisation**. And the translation sync is **`.md`-only**, so no hand-localised `_static` asset can be created, updated or repaired by it — every `data.ipynb` copy had to be repointed by hand in all four repos, filed upstream as QuantEcon/action-translation#271 - [ ] **P4 — dynamic snapshot twin**: originally `UNRATE` alone; **reframed 2026-09-01** as the `business_cycle` set, because the lecture that needs a twin is excluded from `lecture-wasm` for want of one and a partial twin buys it nothing. Done so far: `business_cycle_data.csv` (renamed `gdp_growth_annual.csv` 2026-09-07 under the naming policy, [#113](https://github.com/QuantEcon/data-lectures/issues/113)) manifested and its builder retrofitted ([#109](https://github.com/QuantEcon/data-lectures/pull/109)); the refresh-as-PR and canary workflow ([#110](https://github.com/QuantEcon/data-lectures/pull/110)); the first real refresh ([#112](https://github.com/QuantEcon/data-lectures/pull/112)); the World Bank set extended to three tables and the FRED half landed as one composite monthly file on a shared `builders/_fred.py` library ([#114](https://github.com/QuantEcon/data-lectures/pull/114)). Remaining: the `lecture-wasm` adoption ([QuantEcon/lecture-wasm#70](https://github.com/QuantEcon/lecture-wasm/issues/70) — intro keeps its live calls as the lesson), the flip with `on_refresh: rebuild`, and a canary run catching an induced failure -- [ ] Verify each migrated URL with a pyodide/JupyterLite fetch (CORS, meta#143) -- [ ] Fold every validated decision into the draft `styleguide/datasets.md` (manual#108) as it is proven +- [ ] Verify each migrated URL with a pyodide/JupyterLite fetch (CORS, meta#143) — **header evidence 2026-09-29**: every published file returns 200 from `raw.githubusercontent.com` with `access-control-allow-origin: *`, no redirect, and bytes matching its manifest's sha256; these are simple GETs with no preflight, so that settles CORS. No in-browser Pyodide run is recorded, so whether to require one before ticking is open +- [ ] Fold every validated decision into the draft `styleguide/datasets.md` (manual#108) as it is proven — tracked since 2026-09-29 in [QuantEcon.manual#153](https://github.com/QuantEcon/QuantEcon.manual/issues/153): a style-guide pointer to this repo first, then manual#108 refreshed against D11, `qeld` and the naming and schema rules and merged, or closed in favour of this repo's own docs ### Phase 9 — Adoption (broad sweep — the step that stalled in Feb 2025) - [x] Repoint the remaining consuming lectures as datasets land here (data#4) — **done 2026-08-18**: all 40 static datasets in the corpus are migrated and `repointed` (tracks A–D above; the last four landed in [#98](https://github.com/QuantEcon/data-lectures/pull/98) and flipped in [#99](https://github.com/QuantEcon/data-lectures/pull/99)). The "Repoint rules" stay binding on any future wave: repoint all consumers of a dataset together, and never delete a copy a sibling repo reads - [x] Remove lecture repos' duplicate copies as each repoint merges — **done 2026-09-01** (Track X, [QuantEcon/workspace-lectures#57](https://github.com/QuantEcon/workspace-lectures/issues/57)): 23 audit orphans plus 46 translation copies deleted across twelve repos, one PR each; the wasm mirror copies went only after wasm read data-lectures directly -- [ ] Intake rule for migrations: constructed datasets arrive **with their builders**; of the 5 known constructed-but-unscripted files, three arrived with recovered builders in Track C (`fred_data.csv` and the two `bbh` extracts — see `builders/README.md`); `hansen_jagannathan_1991_data.json` and `acs_data_summary.csv` still have none — recorded as QEP follow-ups per meta#338 -- [ ] Graduate the convention to a QEP and merge manual#108, with the remaining sweep as its rollout checklist +- [x] Intake rule for migrations: constructed datasets arrive **with their builders** — **codified in AGENTS.md** (a constructed dataset ships its builder, and `unrecovered` is for inherited files only). Of the 5 known constructed-but-unscripted files, three arrived with recovered builders in Track C (`fred_data.csv` and the two `bbh` extracts — see `builders/README.md`); `hansen_jagannathan_1991_data.json` and `acs_data_summary.csv` still have none — recorded as QEP follow-ups per meta#338. Across the corpus, 10 inherited datasets are `builder_status: unrecovered` (2026-09-29), shown in the catalog; recovering them has no issue yet +- [ ] Graduate the convention to a QEP and merge manual#108, with the remaining sweep as its rollout checklist. Two pieces now have their own issues: the style-guide page ([QuantEcon.manual#153](https://github.com/QuantEcon/QuantEcon.manual/issues/153)) and the `data-{scope}` repository type, which QEP-3 does not yet list ([QuantEcon/qeps#41](https://github.com/QuantEcon/qeps/issues/41)); whether the convention as a whole still becomes a QEP is meta#336's call -## Open decisions (owned by meta#336 / manual#108, not this repo) +## Convention decisions (owned by meta#336 and the manual, not this repo) -| Decision | Current strawman | +| Decision | Status | | --- | --- | -| Repo name | **settled 2026-07-16**: renamed `data-lectures` (Phase 1) | +| Repo name | **settled 2026-07-16**: renamed `data-lectures` (Phase 1); registering its `data-{scope}` type in QEP-3 is [QuantEcon/qeps#41](https://github.com/QuantEcon/qeps/issues/41) | | URL form | **settled 2026-08-12** (D11): lecture code reads `qeld.url('')`; the direct forms are the runtime-dependent raw URLs (repoint rule 5), standing rather than interim. `data.quantecon.org` deferred indefinitely ([#37](https://github.com/QuantEcon/data-lectures/issues/37)) | -| Layout | flat | -| Licensing review | per-source cache-and-serve-with-attribution gate (`redistribution: permitted \| restricted`), recorded in the manifest — this repo is a stability cache, not a content host | +| Layout | **settled 2026-07-16**: flat (Phase 2) | +| Licensing review | **settled**: a per-source cache-and-serve-with-attribution gate (`redistribution: permitted \| restricted`), recorded in the manifest — this repo is a stability cache, not a content host; licensing does not gate migration (2026-08-06, [#35](https://github.com/QuantEcon/data-lectures/issues/35)) | -When one of these settles, update this PLAN and `AGENTS.md` in the same PR that acts on it. +All four have settled. If one is revisited, update this PLAN and `AGENTS.md` in the same PR that acts on it. diff --git a/README.md b/README.md index 30098ee..91a3fdb 100644 --- a/README.md +++ b/README.md @@ -2,7 +2,7 @@ The canonical repository for **data consumed by the QuantEcon lecture series**, referenced by stable URLs. -> **Status:** renamed from `QuantEcon/data` (2026-07-16) and being shaped into the canonical lecture-data repo per [QuantEcon/meta#336](https://github.com/QuantEcon/meta/issues/336). See [`PLAN.md`](PLAN.md) for the roadmap and [`AGENTS.md`](AGENTS.md) for working conventions. The full data-hosting convention is drafted in [QuantEcon.manual#108](https://github.com/QuantEcon/QuantEcon.manual/pull/108). +> **Status:** live. Renamed from `QuantEcon/data` on 2026-07-16 per [QuantEcon/meta#336](https://github.com/QuantEcon/meta/issues/336); the lectures' static datasets finished migrating here on 2026-08-18, and dynamic snapshots are being adopted (`PLAN.md` Phase 8, P4). See [`PLAN.md`](PLAN.md) for the roadmap and [`AGENTS.md`](AGENTS.md) for working conventions. The convention itself is this README, `AGENTS.md` and [`manifest-schema.yml`](manifest-schema.yml). A style-guide page for lecture authors is tracked in [QuantEcon.manual#153](https://github.com/QuantEcon/QuantEcon.manual/issues/153) (team access), and the `data-{scope}` repository type in [QuantEcon/qeps#41](https://github.com/QuantEcon/qeps/issues/41). ## The routing rule @@ -33,7 +33,7 @@ The `github.com/…/raw/` form is a 302 whose response carries an **empty** `acc 4. Reference it from the lecture — `qeld.url('')` once the package ships, the runtime-correct direct URL until then. The lecture PR builds green immediately, no two-step merge. 5. Add the lecture to the dataset's `consumers` list. -See the [draft convention](https://github.com/QuantEcon/QuantEcon.manual/pull/108) for the full checklist and manifest schema. +Every manifest field is documented in [`manifest-schema.yml`](manifest-schema.yml), and its `schema` rules are checked on every PR. The classes, naming rules and URL rules are in [`AGENTS.md`](AGENTS.md). ## Layout @@ -43,8 +43,8 @@ See the [draft convention](https://github.com/QuantEcon/QuantEcon.manual/pull/10 | `builders/` | one builder per constructed or dynamic dataset, `builders/.` → `lectures/.` | no | | `sources/` | builder inputs that cannot be re-fetched (per-path LFS); `sources/README.md` is their audit trail | no | | `provenance/` | upstream metadata dumps a builder writes beside its dataset — evidence for the manifest's `source` and `license` fields, regenerated every run | no | -| `scripts/` | repo tooling: the catalog generator, the audit dashboard, and the dynamic-snapshot plumbing (`snapshots.py`) | no | -| `manifest-schema.yml` | the per-dataset manifest schema (strawman — see [`PLAN.md`](PLAN.md) Phase 2) | no | +| `scripts/` | repo tooling: the catalog generator, the audit dashboard, the dataset validator (`validate_datasets.py`) and the dynamic-snapshot plumbing (`snapshots.py`) | no | +| `manifest-schema.yml` | the per-dataset manifest schema — the authoritative field reference; its `schema` rules are enforced on every PR by `validate-datasets` | no | | `migration.yml` | the migration lifecycle tracker — which PRs landed and repointed each dataset (transitional; archivable when the migration programme completes) | rendered | The tree is flat because the filename is the interface: `lectures/` is diff --git a/lectures/acs_data_summary.csv.yml b/lectures/acs_data_summary.csv.yml index 277907c..b658dea 100644 --- a/lectures/acs_data_summary.csv.yml +++ b/lectures/acs_data_summary.csv.yml @@ -3,8 +3,10 @@ # lectures/_static/lecture_specific/match_transport/ and was read by relative # path (PLAN Phase 8, wave C2). # -# The name is on the phase-2 rename review (QuantEcon/data-lectures#87, -# decision of 2026-08-17): it migrates under its current name. The org-wide +# The name was on the phase-2 rename review (QuantEcon/data-lectures#87, +# decision of 2026-08-17) and migrated under its current name, which the +# naming policy then kept (QuantEcon/data-lectures#113, 2026-09-07, rule 6: a +# consumed file keeps its name). The org-wide # Trees sweep (2026-08-17, re-derived in QuantEcon/data-lectures#96) found # four copies of this basename org-wide and all four are THE SAME BLOB # (29a0e219…) — inherited duplicates in lecture-dp, lecture-dp.monorepo and diff --git a/lectures/ames_house_prices.csv.yml b/lectures/ames_house_prices.csv.yml index d2a1bea..4f54fcd 100644 --- a/lectures/ames_house_prices.csv.yml +++ b/lectures/ames_house_prices.csv.yml @@ -35,8 +35,8 @@ license: # data-documentation file. Redistribution is long established in practice: # CRAN ships it in the AmesHousing package and OpenML serves it as dataset # 42165. Recorded as `permitted` on that basis, with `name: null` because - # there is a genuine licence statement to point at, not because one was not - # looked for. Attribution is carried by the citation above. + # there is no licence statement to point at, not because one was not looked + # for. Attribution is carried by the citation above. retrieved: 2026-08-03 maintainer: QuantEcon diff --git a/lectures/dataBHS.csv.yml b/lectures/dataBHS.csv.yml index 1814e87..bf989fc 100644 --- a/lectures/dataBHS.csv.yml +++ b/lectures/dataBHS.csv.yml @@ -7,9 +7,10 @@ # python-advanced.quantecon.org/dataBHS.mat is 404 while the published # five_preferences.ipynb is 200 and calls loadmat('dataBHS.mat'). The # downloadable notebook could not run. Serving the data from this repo as CSV -# removes the 404 and the scipy dependency at once. The stem is preserved so -# the real rename rides the phase-2 review (QuantEcon/data-lectures#87) with -# its wave-mates. +# removes the 404 and the scipy dependency at once. The stem was preserved so +# a real rename could ride the phase-2 review (QuantEcon/data-lectures#87) with +# its wave-mates; the naming policy (QuantEcon/data-lectures#113, 2026-09-07, +# rule 6) then kept the name, so no rename is planned. # # Because the extension changes, this file is NOT a byte-identical migration # and cannot claim the usual repoint gate ("a migration moves bytes"). The diff --git a/lectures/fp.dta.yml b/lectures/fp.dta.yml index 7485e0a..042c56b 100644 --- a/lectures/fp.dta.yml +++ b/lectures/fp.dta.yml @@ -4,10 +4,12 @@ # # THE FILENAME IS TWO LETTERS AND SAYS NOTHING — deliberately kept at migration # (decision 2026-08-13; a batch rename across the five generic names in -# .dev/qeld/migration-catalog.md §5.1 is proposed separately, and any rename -# has to be paired with a prose edit because mle.md names `mle/fp.dta` in -# text). Until then the title and description below are the only thing that -# tells a reader what these bytes are. +# .dev/qeld/migration-catalog.md §5.1 was proposed separately, and any rename +# would have to be paired with a prose edit because mle.md names `mle/fp.dta` +# in text). The naming policy (QuantEcon/data-lectures#113, 2026-09-07) then +# kept it for good: a verbatim release keeps its upstream name (rule 5), and a +# consumed file keeps its name (rule 6). So the title and description below +# are the only thing that tells a reader what these bytes are. # # NOT the same dataset as forbes-billionaires.csv in this same tree: that one # is a list of PEOPLE (the Forbes person-level ranking, read by heavy_tails in diff --git a/lectures/fred_data.csv.yml b/lectures/fred_data.csv.yml index 85539f6..b4c0d4d 100644 --- a/lectures/fred_data.csv.yml +++ b/lectures/fred_data.csv.yml @@ -3,11 +3,12 @@ # lectures/_static/lecture_specific/risk_aversion_or_mistaken_beliefs/ and was # read over the repo's OWN raw URL (PLAN Phase 8, wave C2). # -# The name is on the phase-2 rename review (QuantEcon/data-lectures#87, -# decision of 2026-08-17): it migrates under its current name, and the org-wide -# Trees sweep of 2026-08-17 (re-derived in the QuantEcon/data-lectures#96 -# validation) found no other file of this name anywhere in the org, so the -# deferral introduces no ambiguity. +# The name was on the phase-2 rename review (QuantEcon/data-lectures#87, +# decision of 2026-08-17) and migrated under its current name, which the +# naming policy then kept (QuantEcon/data-lectures#113, 2026-09-07, rule 6: a +# consumed file keeps its name). The org-wide Trees sweep of 2026-08-17 +# (re-derived in the QuantEcon/data-lectures#96 validation) found no other file +# of this name anywhere in the org, so keeping it introduces no ambiguity. # # The builder below is a RECONSTRUCTION, not a recovered original: no build # script for this file has ever existed in the lecture repo. Unlike the BBH diff --git a/lectures/private_credit_to_gdp.csv.yml b/lectures/private_credit_to_gdp.csv.yml index 7fbae8d..dbeb8ba 100644 --- a/lectures/private_credit_to_gdp.csv.yml +++ b/lectures/private_credit_to_gdp.csv.yml @@ -1,8 +1,8 @@ # Manifest for private_credit_to_gdp.csv — the third of three World Bank # tables builders/business_cycle.py writes for the intro `business_cycle` -# lecture. The FILENAME IS PROVISIONAL (naming policy: QuantEcon/data-lectures#113) -# and free to change while `consumers` is empty. Landed 2026-09-01 as part of -# the P4 dynamic-snapshot pilot. +# lecture. Named under the naming policy settled 2026-09-07 +# (QuantEcon/data-lectures#113: named for the variable), and final. Landed +# 2026-09-01 as part of the P4 dynamic-snapshot pilot. filename: private_credit_to_gdp.csv title: World Bank domestic credit to the private sector (% of GDP) — United Kingdom, 1960 onward diff --git a/lectures/test_pwt.csv.yml b/lectures/test_pwt.csv.yml index 96c55ee..1d8531c 100644 --- a/lectures/test_pwt.csv.yml +++ b/lectures/test_pwt.csv.yml @@ -4,8 +4,10 @@ # OWN raw URL from two lectures and three synced translation repos (PLAN # Phase 8, Track D). # -# The name is on the phase-2 rename review (QuantEcon/data-lectures#87, -# decision of 2026-08-17): it migrates under its current name. +# The name was on the phase-2 rename review (QuantEcon/data-lectures#87, +# decision of 2026-08-17) and migrated under its current name, which the +# naming policy then kept (QuantEcon/data-lectures#113, 2026-09-07, rule 6: a +# consumed file keeps its name). # # PROVENANCE IS WRONG IN THE CONSUMING LECTURE, measurably. pandas.md's prose # says the file "is taken from the Penn World Tables" and links PWT 7.0 — but diff --git a/manifest-schema.yml b/manifest-schema.yml index faf7f39..eaa189c 100644 --- a/manifest-schema.yml +++ b/manifest-schema.yml @@ -1,10 +1,12 @@ -# Per-dataset manifest — schema sketch (PLAN Phase 2 / QuantEcon.manual#108) +# Per-dataset manifest — the authoritative, commented field reference (PLAN +# Phase 2; AGENTS.md, "Every dataset needs a class and a manifest"). # -# STATUS: strawman. This is the shape the pilot (PLAN Phase 8) tests against; -# every field here is provisional until proven by a real migration. Decisions -# validated by the pilot get folded into styleguide/datasets.md (manual#108). +# STATUS: proven by the migration. Every published dataset carries a manifest +# in this shape, and the `schema` block is executable (see COMPLETENESS below). +# The summary for lecture authors belongs in the manual's style guide +# (QuantEcon.manual#153); this file stays the field-level reference. # -# PLACEMENT (strawman): a manifest is a sidecar next to the file it describes, +# PLACEMENT: a manifest is a sidecar next to the file it describes, # named ".yml": # # lectures/mpd2020.xlsx <- the dataset @@ -30,8 +32,9 @@ # then be documented back here. Two such extensions are already in wide use and # are described where they apply: `schema.sheets` for multi-sheet workbooks, and # `schema.read_as` for the pandas read-kwargs a positional read needs. Adding -# the four `source`/`license` fields below closed the last known gap between -# this file and practice (QuantEcon/data-lectures#85). +# the four `source`/`license` fields below closed the gap known at the time +# (QuantEcon/data-lectures#85); `source.file_url` and `source.citation_policy` +# were in use before they were written down here, and now are. # --------------------------------------------------------------------------- # Identity @@ -74,22 +77,32 @@ source: name: World Bank national accounts data, and OECD National Accounts data files series: NY.GDP.MKTP.KD.ZG url: https://data.worldbank.org/indicator/NY.GDP.MKTP.KD.ZG + # The file's own address, where `url` is a landing page and the artifact has + # a stable direct link or query (a release file, an API endpoint) — the thing + # a re-fetch would download. Optional: omit it when `url` already is the file, + # and leave it null when there is no single file to point at, as here (the + # builder fetches this snapshot through the `wbgapi` client). + file_url: null # Dataset DOI where the source issues one, else null. Prefer a DOI that # resolves to the DATA deposit; an article DOI is better than nothing, but say - # which it is. In use by 17 of 44 manifests. + # which it is. doi: null # The upstream's own version identifier, quoted as the source states it — an # edition ("Maddison Project Database 2020"), a vintage stamp, or a file # header. Null where the source is genuinely unversioned; say so rather than - # inventing a version. In use by 21. + # inventing a version. version: null citation: > World Bank, World Development Indicators, series NY.GDP.MKTP.KD.ZG (GDP growth, annual %). + # The source's own conditions on citation, quoted, where they go beyond + # citing the source itself (Maddison asks for the original papers whenever + # its data is graphed or fewer than 12 countries are used). Optional. + citation_policy: null # Anything a reader must know before trusting these bytes that the fields # above cannot carry — most often that the file is NOT what its name or its # cited paper implies. Where that is the case, say it here AND at the top of - # the manifest. In use by 20. + # the manifest. note: null license: @@ -110,7 +123,7 @@ license: # a named public licence: a split answer across providers, an inherited # exposure, a term found only in prose, or a reasoned judgement about what is # actually published here. A bare `permitted` beside a `name: null` is the - # case that most needs this field. In use by 16. + # case that most needs this field. note: null retrieved: 2024-04-10 # ISO date the bytes were obtained, or