Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
10 changes: 5 additions & 5 deletions AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,7 +4,7 @@ Guidance for coding agents (and humans) making changes in this repository. Read

## What this repo is

The canonical home for **data consumed by the QuantEcon lecture series** (renamed from `QuantEcon/data` on 2026-07-16, per [meta#336](https://github.com/QuantEcon/meta/issues/336)). Its purpose is **stability**: it snapshots upstream sources — with attribution to each source carried in the manifest — so a lecture build never depends on a live API or a third-party host staying up. It is a **cache, not a content-distribution host**. The published tree is **flat** (`lectures/`, since 2026-07-16) and live on GitHub Pages; consumers fetch it over the raw GitHub forms below. **There is no pending host transition**: the `data.quantecon.org` custom domain was deferred indefinitely on 2026-08-12 in favor of the `qeld` consumer package (`PLAN-QELD-PACKAGE.md`, D11) — the stable interface lectures get is a package call, `qeld.url('<filename>')`, not a branded host ([#37](https://github.com/QuantEcon/data-lectures/issues/37), [#15](https://github.com/QuantEcon/data-lectures/issues/15)). The full convention lives in the draft manual page ([QuantEcon.manual#108](https://github.com/QuantEcon/QuantEcon.manual/pull/108)).
The canonical home for **data consumed by the QuantEcon lecture series** (renamed from `QuantEcon/data` on 2026-07-16, per [meta#336](https://github.com/QuantEcon/meta/issues/336)). Its purpose is **stability**: it snapshots upstream sources — with attribution to each source carried in the manifest — so a lecture build never depends on a live API or a third-party host staying up. It is a **cache, not a content-distribution host**. The published tree is **flat** (`lectures/`, since 2026-07-16) and live on GitHub Pages; consumers fetch it over the raw GitHub forms below. **There is no pending host transition**: the `data.quantecon.org` custom domain was deferred indefinitely on 2026-08-12 in favor of the `qeld` consumer package (`PLAN-QELD-PACKAGE.md`, D11) — the stable interface lectures get is a package call, `qeld.url('<filename>')`, not a branded host ([#37](https://github.com/QuantEcon/data-lectures/issues/37), [#15](https://github.com/QuantEcon/data-lectures/issues/15)). The convention's authoritative text is this file, `README.md` and `manifest-schema.yml`. A style-guide page for lecture authors is tracked in [QuantEcon.manual#153](https://github.com/QuantEcon/QuantEcon.manual/issues/153), which also settles the older draft ([QuantEcon.manual#108](https://github.com/QuantEcon/QuantEcon.manual/pull/108)); the `data-{scope}` repository type is proposed for QEP-3 in [QuantEcon/qeps#41](https://github.com/QuantEcon/qeps/issues/41).

## Rules

Expand Down Expand Up @@ -33,9 +33,9 @@ A constructed dataset without its committed builder is a bug. Manifest fields: `

**Capture what the source gives you; never let a missing field block a useful dataset.** Rich provenance — DOI, upstream version, exact retrieval date, licence id — is always welcome and worth recording whenever it is available, because it makes the data auditable years later at almost no ongoing cost. But effort scales with what the source actually provides: where a field is genuinely unavailable, record it as an explicit, reasoned gap (see the inherited-file states below) rather than fabricating it or refusing the file. A clean, well-documented source should produce a short manifest; only genuinely messy provenance earns a long one.

#### Two inherited-file states that look like violations but are tracked, not hidden
#### Three inherited-file states that look like violations but are tracked, not hidden

The Feb 2025 migration left files that cannot fully satisfy the rules above. The manifest records each gap **explicitly** — visible in the generated catalog — rather than burying it by misclassification. Both are provisional decisions from the P1 pilot ([meta#338](https://github.com/QuantEcon/meta/issues/338)), to be folded into [manual#108](https://github.com/QuantEcon/QuantEcon.manual/pull/108).
The Feb 2025 migration left files that cannot fully satisfy the rules above. The manifest records each gap **explicitly** — visible in the generated catalog — rather than burying it by misclassification. `retrieved: null` and `unrecovered` are provisional decisions from the P1 pilot ([meta#338](https://github.com/QuantEcon/meta/issues/338)); `committed-frozen` joined them with the builders directory ([#61](https://github.com/QuantEcon/data-lectures/pull/61)). All three are to be folded into the style-guide page ([QuantEcon.manual#153](https://github.com/QuantEcon/QuantEcon.manual/issues/153)).

- **`retrieved: null` — inherited-undated bytes.** `retrieved` is required, but may be `null` when the bytes were inherited (e.g. from a lecture repo) with **no recorded upstream-retrieval date**. Do **not** reconstruct one from git history — that records when QuantEcon acquired the file, not when it was retrieved from the source, and the false precision is worse than an honest null. A null `retrieved` must be paired with an `integrity.upstream` entry that **accounts for the gap** — never left bare. Any resolved status does that: `unverifiable` says the vintage cannot be established at all, and `diverged` or `verified` say something stronger, because a re-fetch that hashes identically pins the vintage by content rather than by date, which is what a `retrieved` date was a proxy for. What is not acceptable is `null` beside `unverified` — that is two unanswered questions, not one answered a different way. (Amended 2026-08-13, wave B1': the rule previously named `unverifiable` alone, which was the only case that had arisen. `life-expectancy-vs-gdp-per-capita.csv` and `mpd2020.xlsx` already paired `null` with `diverged`; B1' added five files where an upstream re-fetch succeeded byte-identically.)
- **`builder_status: committed-frozen` — the builder is here, and deliberately will not run.** For a dataset built from a source that must not be refreshed: a frozen vintage, or a scraper we will not re-run. The artifact is kept as the record of what produced these bytes, so it is committed verbatim and not edited — editing it is what would destroy its value as provenance. Distinct from `committed`, which asserts a runnable four-stage builder, and from `unrecovered`, which says the builder is absent.
Expand Down Expand Up @@ -120,7 +120,7 @@ Where one builder produces a **set** of files, name it for the set and let each

`scripts/` is repo tooling — the audit dashboard and the catalog generator — and produces no dataset. Keep the two apart.

**Where a builder reads its input from.** The normal case is the third-party upstream, fetched at run time: eight of the nine `committed` builders here do that, and it is the fetch stage of the contract below. A builder reads from `sources/` **only when the input cannot be re-fetched** — the upstream is gone, unlocatable, or was inherited with no recoverable source. `sources/` is that exception layer, not a general input tree, and it is emphatically not "the big-file directory": the defining property is un-refetchability, not size. What it must never be is a network read from another QuantEcon repo — that is how a retired repo becomes load-bearing again.
**Where a builder reads its input from.** The normal case is the third-party upstream, fetched at run time: every `committed` builder here does that except the few that read `sources/`, and it is the fetch stage of the contract below. A builder reads from `sources/` **only when the input cannot be re-fetched** — the upstream is gone, unlocatable, or was inherited with no recoverable source. `sources/` is that exception layer, not a general input tree, and it is emphatically not "the big-file directory": the defining property is un-refetchability, not size. What it must never be is a network read from another QuantEcon repo — that is how a retired repo becomes load-bearing again.

Builders follow four stages — **fetch → pre-process → validate → write** — and only write on validation pass. **The validate stage is shared**: `builders/_validate.py` reads the manifest's `schema` block as its spec (columns and `pattern` runs, dtype families, exact `known_nulls`, the `nulls:` placement rule, `row_count_floor`, `date_range`) and measures the overlap window against the previous vintage; a builder calls `validate(frame.reset_index(), manifest, previous)` and layers on only what the schema cannot say — value bands, a grid check, the per-series revision **bound** (a tracking snapshot is revised by its source, so the test is a tolerance plus a printed summary, never equality). The same function runs over every committed CSV on every PR (`scripts/validate_datasets.py`, `validate-datasets.yml`), so a manifest that drifts from its bytes fails the PR, not the next refresh. **A dynamic builder also exposes `check_committed()`** — its own `validate()` on the committed bytes, no network — which the same workflow runs (`--builders`) under both pandas majors; a builder-specific check that breaks on a pandas change fails the PR that introduces it (#128 was a week-old pandas-3 break the shared layer could not see). Lectures always read the last-good snapshot: an upstream outage may fail a refresh, it must never break a lecture build.

Expand Down Expand Up @@ -193,7 +193,7 @@ scripts/ # repo tooling — NOT published, produces no dataset
# a refresh, render the refresh PR body
audit_annotations.yml # curated judgment for not-yet-migrated data refs
migration.yml # migration lifecycle tracker (status + PR provenance per dataset)
manifest-schema.yml # per-dataset manifest schema (strawman)
manifest-schema.yml # per-dataset manifest schema — the authoritative field reference
requirements.txt
PLAN.md # roadmap — start here
AGENTS.md # this file
Expand Down
31 changes: 22 additions & 9 deletions PLAN-QELD-PACKAGE.md
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
# PLAN — `qeld`, the consumer-side data package

**Status:** design settled, nothing implemented · **Last updated:** 2026-08-12
**Status:** design settled, nothing implemented · **Last updated:** 2026-09-29 (facts in §5.3, §8.3–§8.6 and the Q2 row refreshed; no decision changed)
**Relationship to `PLAN.md`:** that document migrates *bytes* into this repo. This one gives *consumers* a
stable way to read them. They are independent — the migration completes with or without `qeld` — but the
call-site convention here replaces repoint rules 5–6 for any lecture that adopts it.
Expand Down Expand Up @@ -273,8 +273,11 @@ The best diff in the corpus is `lecture-python-advanced.myst/lectures/hansen_jag

### 5.3 `url()` alone covers the corpus

After the `.npy` conversion (§8.3), **`dataBHS.mat` is the only file in the endgame that cannot be read from
a URL**. That is what justifies dropping `fetch()`.
After the `.npy` conversion (§8.3), **every file in the endgame can be read from a URL**: `dataBHS.mat`, the
one exception when this was written, was converted to `dataBHS.csv` at migration
([#98](https://github.com/QuantEcon/data-lectures/pull/98), 2026-08-18; §8.4). That is what justifies
dropping `fetch()`. Until the conversion, the `.npy` pair is the only published format that cannot be passed
straight to a reader function — `np.load` needs `requests` plus `BytesIO`.

### 5.4 Two latent bugs found, worth fixing regardless

Expand Down Expand Up @@ -333,7 +336,7 @@ or every migrated read classifies `local-path` and the dashboard inverts.
| phase | work | gate |
|---|---|---|
| **Q1 — Audit first** | `build_audit.py` learns `qeld.url('X')` → pattern `qeld`, counted migrated **and terminal**. For `pattern == 'qeld'`, assert the key exists in `lectures/` and is not deprecated — otherwise the qeld path loses every assertion #55/#48/#47 added. `migration.yml`: `final` := every code read via qeld, with §4.1 carve-outs terminal on the direct form (the canonical-host arm was retired with D11) | `audit.json` `stats` and `problems` unchanged on today's repos (**not** "byte-identical" — the audit stamps `date.today()`) |
| **Q2 — Schema hygiene** | Document `read_as` (used in 6 manifests) and `sheets` (5) in `manifest-schema.yml` — both are in use and neither appears in the file `AGENTS.md` calls "the authoritative, commented field reference". Add `deprecated:` (new, used nowhere yet) since §3.3 warns on it. `shape` is already documented. Delete `then: "iloc[1:]"` from `longprices.xls.yml:70` by moving `iloc[1:]` into the lecture — a post-read transform encoded as a string to evaluate is exactly what D5 excludes | `manifest-schema.yml` covers every field any manifest uses. Needs none of #14's decisions — do not block on it |
| **Q2 — Schema hygiene** | Document `read_as` (used in 6 manifests) and `sheets` (5) in `manifest-schema.yml` — both are in use and neither appears in the file `AGENTS.md` calls "the authoritative, commented field reference". Add `deprecated:` (new, used nowhere yet) since §3.3 warns on it. `shape` is already documented. Delete `then: "iloc[1:]"` from `longprices.xls.yml:70` by moving `iloc[1:]` into the lecture — a post-read transform encoded as a string to evaluate is exactly what D5 excludes | `manifest-schema.yml` covers every field any manifest uses. Needs none of #14's decisions — do not block on it. *(2026-09-29: `read_as` and `sheets` are now described there; `deprecated:` and the `then:` cleanup remain.)* |
| **Q3 — Package** | `packages/qeld/`: `url()`, `info()`, context detection, advisory catalog. Catalog compiler shares a freshness gate with `CATALOG.md`. Format tier-1 assertion. First release to PyPI via trusted publishing | Offline suite green on every PR: catalog compiles and is fresh; unknown key warns and still returns a URL; URL form correct per detected context; suffix fidelity incl. `.csv.gz`; `info()` fields present. CPython matrix |
| **Q4 — Live leg** | Post-merge + scheduled job: fetch each served URL, compare to the manifest hash, open an issue on failure | Green on `main`; an induced failure opens an issue |
| **Q5 — Browser session** | `%pip install qeld==<v>` in a real `lecture-wasm` page (**`%pip` routes through piplite, not micropip** — a console `micropip.install` is a false pass); `pd.read_excel(qeld.url('mpd2020.xlsx'), sheet_name='Regional data', header=[0,1,2], index_col=0)`; a `.csv.gz` read; record observed Pyodide and pyodide-kernel versions | Written pass/fail. Fail ⇒ wasm keeps URLs and the plan proceeds for the CPython repos |
Expand Down Expand Up @@ -374,29 +377,39 @@ columns (`date`/`specie_value`, `date`/`nominal_balances`). Converting to CSV de
and two imports from `french_rev` in every consuming repo.

**But the window closed.** When this was analysed both files had `consumers: []`; the A3 set has since been
repointed (#49) and both now have two consumers. So this is no longer a free replacement — it needs the
repointed (#49), and both now have six consumers each (intro, wasm, `lecture-intro.zh-cn`, the actions
canary, `tom-econ370-2025` and `python-lecture-sandpit.myst`, per the manifests on 2026-09-29). So this is no
longer a free replacement — it needs the
`AGENTS.md` "new vintage → new filename" treatment (`caron.csv` lands alongside, consumers opt in, the `.npy`
is swept later) or a coordinated set under repoint rules 1–3. **Decide before Q6**, since `french_rev` is a
pilot.

### 8.4 `dataBHS.mat` — convert at migration, or exclude?

**Resolved 2026-08-18: converted at migration.** `dataBHS.csv` landed in
[#98](https://github.com/QuantEcon/data-lectures/pull/98) and the lecture reads it; the `.mat` stays as the
builder's input in `sources/`. The analysis that led there, as written:

5,588 bytes; `c`, `rb`, `rs`, each (236, 1) float64; the lecture uses only `data['c']` and the read is inside
a `hide-input` cell, so nothing about it is taught. Trivially a 236×3 CSV — but see §4.3 on the read. A Track
C decision; the only true impossibility among static files.

### 8.5 Is `lecture-intro.zh-cn` in scope?

It carries data reads, appears in **zero** `consumers` blocks, is excluded from `SCAN_REPOS` by decision, has
no data CI, publishes on a `publish*` tag, and inherits install cells automatically via the `.md`-only sync —
It carries data reads, is excluded from `SCAN_REPOS` by decision, has no data CI, publishes on a `publish*` tag, and inherits install cells automatically via the `.md`-only sync —
so it acquires whatever intro acquires without anyone deciding. It also has files with no data-lectures key
and no business having one (`country_code_cn.csv`, a translation asset).
and no business having one (`country_code_cn.csv`, a translation asset). When this was written it appeared
in no `consumers` block; the Track A and P3 repoints have since recorded it in 25 `consumers` entries (as of
2026-09-29), so its reads are now in the manifests even though no CI sees them.
**Recommendation: explicit non-goal for v1, with one fixed rule instead of machinery — any sweep touching an
intro file also touches zh-cn.**

### 8.6 The rename list for generic filenames

§5.5. Needs a pass before Tracks B and C migrate, and each rename needs its prose pairing found by grep.
§5.5. **Settled.** The list moved to [#87](https://github.com/QuantEcon/data-lectures/issues/87) on
2026-08-17, and the naming policy ([#113](https://github.com/QuantEcon/data-lectures/issues/113),
2026-09-07) decided it: a consumed file keeps its name (`manifest-schema.yml` naming rule 6), and `fp.dta`, a
verbatim release, keeps its upstream name (rule 5). Tracks B and C migrated under their current names.

### 8.7 Open from the original report

Expand Down
Loading
Loading