Skip to content

Scaffolding: QuantEcon/data → data-lectures (canonical lecture-data repo) #8

Description

@mmcky

Part of QuantEcon/meta#336 (design thread) — this issue documents the scaffolding and initial maintenance work to take this repo from its current state to the canonical data-lectures repository described in the draft convention (QuantEcon/QuantEcon.manual#108). The pilot (QuantEcon/meta#338) lands its migrations here, so the early items below are its prerequisites.

Checklist refreshed 2026-09-01 (previously 2026-08-06) — boxes reflect verified state of main, not intent. Closed 2026-09-01: the scaffolding is done and the static migration it was the prerequisite for completed 2026-08-18; the boxes still open below are Phase 5 automation, tracked in PLAN.md and #14 rather than here. PLAN.md remains the roadmap; this issue is the scaffolding subset of it.

Current state (audit, 2026-07-15 — superseded, kept for the record)

The repo holds 10 files for lecture-python-intro under a consumer-keyed layout (lecture-python-intro/static/, dynamic/, scripts/), has one manual refresh script (business_cycle.py), no .github/ directory at all (no CI, no automation, no scheduled refresh), no LFS, no per-dataset metadata, and is referenced by zero lectures — the sweep in #4 never happened.

Where it stands 2026-08-06: flat published tree, 10 datasets with manifests, all 10 repointed, .github/ with three workflows, Pages live and CORS-clean, branch protection with a required check. The audit dashboard is green against all 8 lecture repos: 41 static files, 35 orphans, 22 live-API lectures, 0 legacy refs.

Identity

Layout

One layout question from the flatten is still open: whether business_cycle's two .md provenance dumps belong in the published tree at all — #13. Both are live public URLs today, so this got more expensive since it was raised.

Storage

Publishing

Automation (.github/)

Go-live guardrails, added ahead of the first repoint:

Remaining:

The fetch-layer question these builders depend on is #26 (pandas_datareader is maintained again).

Metadata backfill for existing holdings

  • Per-dataset manifest schema — manifest-schema.yml, revised by P1 (P1 pilot: add lingcod_msy_recovery.csv with its manifest #12) and exercised by nine more datasets since
  • [~] Per-dataset manifest for the existing files — 10 of 19 done. The ten migrated datasets all have sidecars; the 8 static intro files in the published tree (mpd2020.xlsx, longprices.xls, chapter_3.xlsx, assignat.xlsx, dette.xlsx, fig_3.xlsx, caron.npy, nom_balances.npy) still have none, and neither does business_cycle_data.csv. This is the next tranche of work
  • Keep-or-drop for the three no-consumer files — dropped in Flatten the consumer-keyed tree into the published layout #10 (two World Bank CSVs, fig_3.ods), recoverable from history
  • License check per file before the repo is promoted as the canonical public home — policy settled 2026-08-06: licensing does not gate migration. Data already served publicly by the lectures migrates with its status recorded in the manifest; anything needing further thought is tracked for review before data.quantecon.org is promoted as a public open-data host. See the inventory issue and QuantEcon/workspace-lectures#20

Adoption (the step that stalled in Feb 2025)

  • [~] Repoint the consuming lectures as datasets land here (Add data and scripts #4's unticked box) — 10 of 41 done, all repointed in migration.yml. The remaining 31 are waved in PLAN.md
  • Remove lecture repos' duplicate copies as each repoint merges — 24 orphans across 6 repos as of the 2026-08-31 audit (down from 35); every repoint has landed, so this is now Track X, tracked at QuantEcon/workspace-lectures#57

Two sequencing rules learned since and now recorded in PLAN.md: a dataset a sibling repo reads (every lecture-wasm case) must have that sibling repointed before the owning repo's copy is deleted, or the sibling 404s; and because the strict audit has no green state for a partially-repointed dataset, all consumers of one dataset must be repointed together. Sixteen of the thirty-one remaining datasets are multi-consumer, every one of them lecture-python-intro + lecture-wasm.

Activity

  1. mmcky commented on Jul 16, 2026

    @mmcky
    ContributorAuthor

    Status — 2026-07-16

    Two PRs open against this checklist, both in review.

    #10 — the restructure (PLAN Phase 2). Flat published tree: the data moves out of lecture-python-intro/{static,dynamic}/ into lectures/, and scripts/ lifts to the repo root outside the published tree. Phase 2's open question — where non-published assets live — is answered in it: scripts/ and manifest-schema.yml at the root, manifests as sidecars inside lectures/ named <full filename>.yml so a dataset cannot be moved without its metadata. Full filename rather than stem, because a stem-keyed sidecar collides the moment a dataset ships in two formats — exactly the fig_3.xlsx / fig_3.ods case this repo already had.

    Doing it now was the point. Zero lectures reference this repo — confirmed by org-wide code search rather than only the audit — so every move is free today, and a breaking change the moment the first repoint merges. AGENTS.md now says that where it previously promised the freedom.

    Also in #10: the keep-or-drop call on the three no-consumer files, made rather than deferred. fig_3.ods was verified to be a pure format twin of fig_3.xlsx (both parse to a single Sheet1, identical 34×6 frame, DataFrame.equals → true), and the two World Bank CSVs return zero references org-wide and are re-downloadable at will. All three dropped rather than promoted into the canonical namespace, where each would then have needed a manifest, a licence check and integrity verification. They remain in git history. Happy to keep any of them if you disagree — that is the Phase 6 decision made explicit rather than silently.

    business_cycle.py is repointed at lectures/ in the same PR; it wrote to ../dynamic, which no longer exists, and would have failed silently at the next refresh.

    #12 — pilot P1 (lingcod). The first dataset in the published tree and the first manifest written against the schema. Meeting one real file changed that schema in three places, all folded back into manifest-schema.yml: integrity split into migration vs upstream (the file is a byte-perfect copy of what the lecture consumes and its provenance cannot be re-derived — both true at once); builder_status: unrecovered, so a constructed-but-builderless file stays visible in the catalog instead of being quietly relabelled verbatim; and known_nulls, because F_over_Fmsy is legitimately null in its terminal assessment year and a blanket no-nulls rule would reject a good file.

    Licence cleared before adding, and how is itself a finding: RAM Legacy is CC BY 4.0, but ramlegacy.org states no licence at all — the authority is the Zenodo record behind its DOI. "Check the homepage" is not a sufficient procedure, so the manifest now records when a licence was established and against what.

    Checklist state

    Full session notes: QuantEcon/workspace-lectures#14. Convention findings: QuantEcon/QuantEcon.manual#108.

  2. mmcky commented on Aug 6, 2026

    @mmcky
    ContributorAuthor

    Checklist above refreshed against verified state of main (2026-08-06) rather than intent — three boxes ticked that had quietly become true (layout decisions, catalog, and the CORS check, which I verified returns access-control-allow-origin: * on the Pages domain today), two marked partial with counts, and the licensing box closed out under the policy that migration does not wait on licence review.

    The stale framing was mostly in Publishing and Metadata: Pages has been live since #21 but the checklist still read as if none of it existed, while the manifest backfill box implied ten files when it is really ten done of nineteen — the eight static intro files and business_cycle_data.csv still have no sidecar, and that is the next tranche.

    Two items are now explicitly blocked rather than merely unstarted: PR validation waits on #14, and the DNS work is the only thing here with external lead time. Everything else in Storage and Automation is unblocked work.

  3. mmcky commented on Aug 31, 2026

    @mmcky
    ContributorAuthor

    Closing. The scaffolding this issue tracked is done — flat published tree, 40 manifests, Pages live and CORS-clean, branch protection with a required check, the audit dashboard, per-path LFS scoped to sources/, and the high_dim_data fold — and the static migration it existed to make possible completed on 2026-08-18 (40 of 40 datasets repointed, strict audit green on the 2026-08-31 run). Checklist refreshed against main one last time.

    The boxes still open are all Phase 5 automation (PR validation, scheduled refresh, the sources-alive canary, the business_cycle.py retrofit, the reusable workflow) plus the Pages soft-limit watch. Those are owned by PLAN.md Phase 5 and #14, and are the next build front as Track E / pilot P4 — a scaffolding issue is no longer the right place to hold them. The orphan sweep is Track X at QuantEcon/workspace-lectures#57.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions