Repository navigation
Scaffolding: QuantEcon/data → data-lectures (canonical lecture-data repo) #8
Description
Activity
Status — 2026-07-16
Two PRs open against this checklist, both in review.
#10 — the restructure (PLAN Phase 2). Flat published tree: the data moves out of
lecture-python-intro/{static,dynamic}/intolectures/, andscripts/lifts to the repo root outside the published tree. Phase 2's open question — where non-published assets live — is answered in it:scripts/andmanifest-schema.ymlat the root, manifests as sidecars insidelectures/named<full filename>.ymlso a dataset cannot be moved without its metadata. Full filename rather than stem, because a stem-keyed sidecar collides the moment a dataset ships in two formats — exactly thefig_3.xlsx/fig_3.odscase this repo already had.Doing it now was the point. Zero lectures reference this repo — confirmed by org-wide code search rather than only the audit — so every move is free today, and a breaking change the moment the first repoint merges.
AGENTS.mdnow says that where it previously promised the freedom.Also in #10: the keep-or-drop call on the three no-consumer files, made rather than deferred.
fig_3.odswas verified to be a pure format twin offig_3.xlsx(both parse to a singleSheet1, identical 34×6 frame,DataFrame.equals→ true), and the two World Bank CSVs return zero references org-wide and are re-downloadable at will. All three dropped rather than promoted into the canonical namespace, where each would then have needed a manifest, a licence check and integrity verification. They remain in git history. Happy to keep any of them if you disagree — that is the Phase 6 decision made explicit rather than silently.business_cycle.pyis repointed atlectures/in the same PR; it wrote to../dynamic, which no longer exists, and would have failed silently at the next refresh.#12 — pilot P1 (lingcod). The first dataset in the published tree and the first manifest written against the schema. Meeting one real file changed that schema in three places, all folded back into
manifest-schema.yml:integritysplit intomigrationvsupstream(the file is a byte-perfect copy of what the lecture consumes and its provenance cannot be re-derived — both true at once);builder_status: unrecovered, so a constructed-but-builderless file stays visible in the catalog instead of being quietly relabelledverbatim; andknown_nulls, becauseF_over_Fmsyis legitimately null in its terminal assessment year and a blanket no-nulls rule would reject a good file.Licence cleared before adding, and how is itself a finding: RAM Legacy is CC BY 4.0, but ramlegacy.org states no licence at all — the authority is the Zenodo record behind its DOI. "Check the homepage" is not a sufficient procedure, so the manifest now records when a licence was established and against what.
Checklist state
-
PLAN.md,AGENTS.md, README rewrite — Phase 0, Add PLAN.md, AGENTS.md and rewrite README (Phase 0 scaffolding) #9 - Rename + description/topics — Phase 1
- Restructure to the flat published tree — Phase 2, Flatten the consumer-keyed tree into the published layout #10 in review
- Sketch the per-dataset manifest schema —
manifest-schema.yml, Flatten the consumer-keyed tree into the published layout #10/P1 pilot: add lingcod_msy_recovery.csv with its manifest #12 in review - Generate the catalog page from the manifests — Phase 2, not started
- Per-path LFS / fold in
high_dim_data— Phase 3, not started. Related: Lecture data: fix live hosting risks and bring high_dim_data into shape meta#337 risk 1 is now closed, sohigh_dim_data's branch-only SCF file is onmainand its four consumers are repointed — the fold starts from a cleaner base than it would have this morning.
Full session notes: QuantEcon/workspace-lectures#14. Convention findings: QuantEcon/QuantEcon.manual#108.
-
- added a commit that references this issue
on Jul 16, 2026 Checklist above refreshed against verified state of
main(2026-08-06) rather than intent — three boxes ticked that had quietly become true (layout decisions, catalog, and the CORS check, which I verified returnsaccess-control-allow-origin: *on the Pages domain today), two marked partial with counts, and the licensing box closed out under the policy that migration does not wait on licence review.The stale framing was mostly in Publishing and Metadata: Pages has been live since #21 but the checklist still read as if none of it existed, while the manifest backfill box implied ten files when it is really ten done of nineteen — the eight static intro files and
business_cycle_data.csvstill have no sidecar, and that is the next tranche.Two items are now explicitly blocked rather than merely unstarted: PR validation waits on #14, and the DNS work is the only thing here with external lead time. Everything else in Storage and Automation is unblocked work.
Closing. The scaffolding this issue tracked is done — flat published tree, 40 manifests, Pages live and CORS-clean, branch protection with a required check, the audit dashboard, per-path LFS scoped to
sources/, and thehigh_dim_datafold — and the static migration it existed to make possible completed on 2026-08-18 (40 of 40 datasets repointed, strict audit green on the 2026-08-31 run). Checklist refreshed againstmainone last time.The boxes still open are all Phase 5 automation (PR validation, scheduled refresh, the sources-alive canary, the
business_cycle.pyretrofit, the reusable workflow) plus the Pages soft-limit watch. Those are owned by PLAN.md Phase 5 and #14, and are the next build front as Track E / pilot P4 — a scaffolding issue is no longer the right place to hold them. The orphan sweep is Track X at QuantEcon/workspace-lectures#57.
Part of QuantEcon/meta#336 (design thread) — this issue documents the scaffolding and initial maintenance work to take this repo from its current state to the canonical
data-lecturesrepository described in the draft convention (QuantEcon/QuantEcon.manual#108). The pilot (QuantEcon/meta#338) lands its migrations here, so the early items below are its prerequisites.Checklist refreshed 2026-09-01 (previously 2026-08-06) — boxes reflect verified state of
main, not intent. Closed 2026-09-01: the scaffolding is done and the static migration it was the prerequisite for completed 2026-08-18; the boxes still open below are Phase 5 automation, tracked in PLAN.md and #14 rather than here.PLAN.mdremains the roadmap; this issue is the scaffolding subset of it.Current state (audit, 2026-07-15 — superseded, kept for the record)
The repo holds 10 files for lecture-python-intro under a consumer-keyed layout (
lecture-python-intro/static/,dynamic/,scripts/), has one manual refresh script (business_cycle.py), no.github/directory at all (no CI, no automation, no scheduled refresh), no LFS, no per-dataset metadata, and is referenced by zero lectures — the sweep in #4 never happened.Where it stands 2026-08-06: flat published tree, 10 datasets with manifests, all 10
repointed,.github/with three workflows, Pages live and CORS-clean, branch protection with a required check. The audit dashboard is green against all 8 lecture repos: 41 static files, 35 orphans, 22 live-API lectures, 0 legacy refs.Identity
QuantEcon/data→data-lecturesonce the name settles in Proposal: a single canonical QuantEcon datasets repository, with a documented convention meta#336 (GitHub redirects preserve existing URLs, so this is non-breaking and can go first) — done 2026-07-16, old raw URLs verified to 200 through the redirectREADME.md: purpose, the routing rule, how to add a dataset, link to the manual page — done in Add PLAN.md, AGENTS.md and rewrite README (Phase 0 scaffolding) #9 (2026-07-16), together withPLAN.md(the phased roadmap for this issue) andAGENTS.mdLayout
lecture-python-intro/...) to the flat published tree in the draft convention — no folder implies ownership by a series — done in Flatten the consumer-keyed tree into the published layout #10 (2026-07-16)scripts/, per-datasetmanifest.yml, tests) relative to the published tree — settled in Flatten the consumer-keyed tree into the published layout #10:scripts/andmanifest-schema.ymlat the root, manifests as sidecars insidelectures/named<full filename>.ymlscripts/build_catalog.py→CATALOG.md, done in Add generated dataset catalog (CATALOG.md + build_catalog.py) #19One layout question from the flatten is still open: whether
business_cycle's two.mdprovenance dumps belong in the published tree at all — #13. Both are live public URLs today, so this got more expensive since it was raised.Storage
.gitattributes— large binaries only, small teaching files plain git (Setup Git LFS for large file support #1; avoid high_dim_data's blanket*.csvrule) — done 2026-08-10 (LFS: scope it to sources/**, and stop fetching it in CI #57), scoped tosources/**only; the published tree is plain git (Storage policy for large data files: plain git in lectures/, LFS only in sources/ — and the ladder above 100 MiB #58). Original note: note the upstream rule is*.csvand*.dta. Worth re-testing the premise: both SCF minis (31.3 MiB and 72.4 MiB) fit plain git under GitHub's 100 MiB limit, so LFS may only be needed forSCF_plus.dta(99.12 MiB), which has no lecture consumerhigh_dim_datacontent (Migratehigh-dim-datato this repo #2; coordinates with Lecture data: fix live hosting risks and bring high_dim_data into shape meta#337 for the consuming-lecture repoints and the branch-only SCF file) — done 2026-08-11 (Fold in the six high_dim_data datasets (PR B1) #62, Land SCF_plus.dta in sources/, and gate sources/ on its recorded hashes (PR B2) #63, 28 reads repointed across four repos, flipped in Flip the six high_dim_data datasets to repointed #69). Original note: the branch-only file is resolved upstream, so the fold starts from a clean basePublishing
lfs: trueat checkout so LFS objects publish as bytes rather than pointer files — done 2026-07-17 in Add the generated data-audit dashboard (gh-pages): audit + migration tracker #21; dashboard at/, published tree at/lectures/— deferred indefinitely 2026-08-12 (D11, PLAN-QELD-PACKAGE.md §2.2, recorded in Record D11: invest in qeld, defer data.quantecon.org indefinitely #76): thedata.quantecon.orgDNS + custom domainqeldpackage is the stable consumer interface instead of a branded host, so this item no longer gates anything — Phase 4 cutover: swap interim raw URLs → data.quantecon.org/lectures #15 is superseded by qeld adoption and closed, and Track Y (deferred): data.quantecon.org DNS + custom domain — retired in favor of qeld (D11); reopenable, NXDOMAIN as of 2026-08-10 #37 stays open as the deferred tracker. The external blocker also dissolved on its own: the stale A record was deleted and the name is NXDOMAIN (measured 2026-08-10), so reviving this is two actions QuantEcon controls whenever D11 is revisitedaccess-control-allow-origin: *on the served files (pyodide/JupyterLite requirement, [Plan/Idea] Supporting WASM based introductory series meta#143) — verified 2026-08-06 on the default Pages domain:quantecon.github.io/data-lectures/lectures/lingcod_msy_recovery.csvreturnsaccess-control-allow-origin: *. The requirement is satisfied today and does not wait on the custom domain; re-verify once DNS movesAutomation (
.github/)Go-live guardrails, added ahead of the first repoint:
main— PRs required, force-push and deletion blocked (protect-mainruleset, 2026-07-17).github/workflows/consumed-file-check.yml(2026-07-17).github/workflows/audit-dashboard.yml(Build a generated data-audit dashboard (gh-pages) across the 8 Python-family repos #20, closed 2026-08-06), with a drift-alarm inbox added in Give the drift alarm an inbox — assign a failure issue to mmcky #29Remaining:
known_nullsexact-vs-ceiling, dtype vocabulary). The dtype vocabulary has already drifted across the ten manifests, so this is no longer cost-free to deferscripts/business_cycle.pyto the four-stage contract — it still has no validate stageThe fetch-layer question these builders depend on is #26 (
pandas_datareaderis maintained again).Metadata backfill for existing holdings
manifest-schema.yml, revised by P1 (P1 pilot: add lingcod_msy_recovery.csv with its manifest #12) and exercised by nine more datasets sincempd2020.xlsx,longprices.xls,chapter_3.xlsx,assignat.xlsx,dette.xlsx,fig_3.xlsx,caron.npy,nom_balances.npy) still have none, and neither doesbusiness_cycle_data.csv. This is the next tranche of workfig_3.ods), recoverable from historydata.quantecon.orgis promoted as a public open-data host. See the inventory issue and QuantEcon/workspace-lectures#20Adoption (the step that stalled in Feb 2025)
repointedinmigration.yml. The remaining 31 are waved inPLAN.mdTwo sequencing rules learned since and now recorded in
PLAN.md: a dataset a sibling repo reads (everylecture-wasmcase) must have that sibling repointed before the owning repo's copy is deleted, or the sibling 404s; and because the strict audit has no green state for a partially-repointed dataset, all consumers of one dataset must be repointed together. Sixteen of the thirty-one remaining datasets are multi-consumer, every one of themlecture-python-intro+lecture-wasm.