Repository navigation
Builder architecture: four-stage template, traceback-clean failures, notebook policy (Phase 5) #14
Description
Activity
Schema implications of this architecture
Writing the template surfaced a consequence worth stating outright: the moment a builder's
validate()reads the sidecar manifest and enforces it, the manifest'sschemablock stops being documentation and becomes an executable contract. Several fields are not yet precise enough to execute. None of this needs action now — the schema is a strawman and says so — but these are exactly what the next pilot and Phase 5 should test the schema against, so they are recorded here rather than lost.1. Column patterns — a real gap our own flagship file already hits. The template checks
list(df.columns) != expected, but the schema example forbusiness_cycle_data.csvlistsYR1960with the comment "one column per year" — a single exemplar standing in for ~64 columns. A machine validator cannot read that comment; it would fail the file. The schema needs a way to express repeated columns before validation can be real, e.g.:columns: - {name: economy, dtype: string, description: ISO3 country code} - {name: Country, dtype: string, description: Country name} - {pattern: 'YR\d{4}', dtype: float64, description: Annual GDP growth (%), one column per year}
2.
known_nullssemantics need pinning: exact count or ceiling? The template enforces exact (n != known), the stronger invariant and right for frozen files like lingcod (F_over_Fmsy: 1, always). But for a dynamic snapshot, null counts can legitimately drift between refreshes — a new terminal assessment year appears and the null moves or multiplies. Likely answer: exact for static extracts, ceiling for dynamic snapshots — or a per-column form like{F_over_Fmsy: {max: 3}}. Needs an explicit decision, not a silent default.3. dtype vocabulary is currently inconsistent. The lingcod manifest says
int64/float64(pandas names); the schema sketch saysstring/float(loose names). A human reader doesn't care;validate()must compare against something canonical. Recommendation: standardise on pandas dtype strings, since that is what the checking code actually sees.4. Float tolerance for the overlap window. "Values unchanged vs last-good" cannot be
!=on floats (the template already flags this —np.iscloseneeded). If tolerance ever needs to vary by dataset it becomes a schema field (overlap_tolerance); until then a builder-side default suffices. Awareness item, not a field yet.5.
builder_statusfor notebook-origin datasets — already an open sub-question in the issue body; noted here only so the schema-facing list is complete in one place.Items 1–3 are decisions the schema has to make before Phase 5's validation can run; 4–5 are cheap to defer. The P2 pilot is the natural place to test 1–3 against a second real file.
- added 2 commits that reference this issue
on Jul 16, 2026 These three are now blockers, not deferrals — and item 3 has already drifted
Recording where items 1–3 stand against
maintoday (2026-08-06), because the framing has changed since this was written. When this issue was opened the schema was a strawman with one real file behind it. There are now ten manifests, and the next substantive piece of work (the P4 dynamic snapshot) cannot be written until these are settled —validate()reads the sidecar as its spec, so an imprecise sidecar is an unwritable builder. Treating them as "decide when convenient" means P4 stalls on them mid-implementation.Concrete recommendations below, with what the ten manifests actually show.
1. Column patterns — confirmed, and the flagship file is worse than described
business_cycle_data.csvhas 66 columns:economy,Country, thenYR1960…YR2023. Enumerating them in the manifest is both unreadable and wrong — it hard-codes an end year into a file whose whole purpose is to grow annually. Apatternentry is the only sane form.Recommendation: a
columnsentry carries eithername(exact) orpattern(a regex), never both, and the validator walks the entries in order, letting apatternentry consume one or more consecutive columns. That keeps column order checkable — which matters, since positional reads exist in the wild — while making the repeated block expressible.For
business_cycle_data.csvthat is three entries instead of sixty-six, and it stays correct when 2024 lands. Apatternentry should also requiremin_count(here 1) so an upstream change that silently drops every year column still fails.2.
known_nulls— exact vs ceiling splits cleanly on classAll ten manifests use exact counts today, and every one of them is a static file where that is right:
lingcodF_over_Fmsy: 1,realwagevalue: 68,countriesCapital: 248,employfour columns at1080each. Seven declareknown_nulls: {}with a prose note explaining why zero is meaningful.Nothing has yet needed a ceiling — because the repo has no dynamic snapshot under validation.
UNRATEwill be the first, and a monthly refresh legitimately changes null counts when a terminal observation is revised.Recommendation: keep the plain integer meaning exact, and add
{max: N}as the ceiling form, with the default chosen by class rather than per file — exact forverbatimandconstructed, ceiling fordynamic snapshot. Class already exists in every manifest, so this needs no new field and no per-file judgment. A dynamic dataset that genuinely wants an exact count can still write the integer form and get the stronger check.3. dtype vocabulary — this has already drifted, and pandas 3 makes it urgent
This is the one that stopped being theoretical. Across the ten manifests: 21
string, 19int64, 12float64, 7object— and the split is chronological, not principled:Manifests Text dtype used countries.csv,employ.csv,realwage.csv(P2, July)stringames_house_prices.csv,epl_match_goals.csv,japan_earthquakes.csv,us_adult_heights.csv(P5, August)objectNobody decided to change it; the convention flipped between pilots. Two names for one concept in a field that is about to become executable is exactly the drift this issue predicted.
The pandas 3 angle makes the choice for us. The lecture repos are on
anaconda=2026.07, i.e. pandas 3.x, where the default dtype for text read from CSV is the new string dtype rather thanobject. A manifest assertingobjectwill therefore be wrong against a defaultread_csvunder the pin the lectures already use — the sevenobjectentries are legacy the day they were written, not a live alternative.Recommendation: canonicalise on the pandas dtype string as
str(df[col].dtype)produces it, withstringfor text, and havevalidate()normaliseobjectandstringto one another when comparing so the check does not become version-brittle across the pandas 2/3 boundary. Then fix the sevenobjectentries in a single sweep — it is a comment-level change to four manifests with no data impact, cheapest done now while it is four files rather than forty.Suggested order
Land 1–3 in
manifest-schema.ymlas one PR before any P4 builder code is written, then retrofit the fourobjectmanifests in the same PR so the repo is internally consistent from that point. Items 4 (float tolerance) and 5 (notebook-originbuilder_status) stay deferred as this issue already argues — 4 has a sensible builder-side default, and 5 has no instance yet.Part of #8. Gates the P4 work in
PLAN.mdPhase 5 / Phase 8, and thebusiness_cycle_data.csvmanifest discussed in #13.Correction to item 3: the pandas 3 name is
str, notstringMy earlier comment recommended canonicalising text columns on
string. That was wrong, and writing the eight manifests in #38 surfaced it — I now have a pandas 3 environment matching the lectures'anaconda=2026.07pin, and measured it rather than reasoning about it.Under pandas 3.0.5,
str(df[col].dtype)for a text column read byread_csvisstr. All three names exist and are distinct:Construction str(dtype)read_csvtext column (pandas 3 default)strpd.Series([...], dtype='object')objectpd.Series([...], dtype='string')stringSo there are now three candidate spellings in play, not two, and the manifests currently use two of them — 21
stringfrom the P2 pilot and 7objectfrom P5 — neither of which is what a default read actually produces under the pin the lectures run on. Avalidate()comparingstr(df[col].dtype)against either would fail on every text column.Revised recommendation. Canonicalise on
str, since that is what the checking code will actually see. Havevalidate()normalisestr,stringandobjectto one another for text comparison, so the check survives the pandas 2/3 boundary and does not become a version tripwire — this matters becauselecture-datascience.mystis still onanaconda=2024.06, where the same read yieldsobject. Then sweep the existing manifests once.The new manifests in #38 use
strfor text columns (mpd2020.xlsx'scountrycodeandcountry), measured against the file rather than assumed, so they are already on the proposed canonical form.Two schema cases #38 also surfaces
These are the first multi-sheet workbooks to get manifests, and they do not fit the flat-table sketch at all.
Sheets, not just columns. A workbook's contract is which sheets and cell ranges the lectures read.
dette.xlsxhas 27 sheets, of which the lecture reads five ranges across two. The manifests use asheets:list where each entry recordsread_as(the exact pandas call), the resulting shape, dtypes and null counts — and names the unread sheets so nobody strips them as dead weight when they are in fact the authors' working provenance.Positional reads are a fragility class the schema has no vocabulary for. Every French Revolution read is
header=Nonewith explicitusecols/skiprows/nrows— for instancepd.read_excel(dette_url, sheet_name='Militspe', usecols='M:X', skiprows=7, nrows=102, header=None). A row or column inserted anywhere above or left of that range silently shifts what the lecture plots, with no error. Column-name validation cannot detect it, because there are no column names. The manifests flag it withpositional_reads: true, but it is worth deciding whether the schema should express the invariant properly — a checksum over the read block would catch what a column check cannot.That last one strengthens the case for settling items 1–3 before Phase 5 validation is written: the schema currently has no way to state the invariant that actually protects these four files.
Deferred item 5 stops being deferred at step 3 — decide it before the fold PR, not after
Step 3 of QuantEcon/workspace-lectures#23 (fold in
high_dim_data) migrates the builders as well as the data, and both builders are notebook-form:cross_section/webscrape_forbes.ipynb(5,861 B, 8 cells, no stored outputs) andSCF_plus/generating_mini.md— the latter a jupytext MyST twin (format_name: myst, 0.13, jupytext 1.14.1), not the.py:percenttwin this issue's notebook policy names. So the "decide when the first such dataset lands" sub-question lands now, inside the largest remaining migration set, and it is the first thing in this repo that its answer has to fit.Today the repo has exactly one builder shape and nothing else: seven plain
.pyfiles underscripts/, zero.ipynbanywhere in the tree, and zero mentions of jupytext, papermill or nbconvert inAGENTS.md,scripts/README.mdormanifest-schema.yml(requirements.txtis two lines). Step 3 introduces a second and a third form in one PR.The enum has no cell for it, and nothing checks.
manifest-schema.yml:178allowscommitted | unrecovered | not-applicable.build_audit.py:489only records the value andbuild_catalog.py:111only formats it — no check enforces the vocabulary, so whatever step 3 writes lands silently. Both usable values are false here:committedasserts a CI-runnable builder of the shape defined above, andunrecoveredsays the builder was lost when it is arriving in the same PR.generating_mini.mdmakescommitteduntenable on its own terms: bothto_csvcalls are commented out upstream (# df1.to_csv('SCF_plus_mini.csv', index=None),# df2.to_csv('SCF_plus_mini_no_weights.csv', index=None)), and two cells are bare display expressions (df1,df2). Fetch and transform are complete; write and validate do not exist. That is the same retrofit shape asbusiness_cycle.py— item 2 of this issue's "When Phase 5 is actioned" — which is worth pairing, becausebusiness_cycle_data.csvhas no sidecar manifest (18 manifests for 21 files inlectures/), so itsvalidate()has noschemablock to read either. Two Phase 5 retrofits, one missing spec between them.Two things argue for "provenance, not runnable builder" rather than a port:
- Both are one-off constructions by this issue's own test — the Forbes lists are a frozen 2020 vintage, and the SCF minis derive from a static deposit
.dta. Neither refreshes, so the port-to-.pybranch does not apply. webscrape_forbes.ipynbfetches Forbes' undocumented internal API with a spoofed browser user-agent and hardcoded GDPR consent cookies (notice_gdpr_prefs), and writes with bare relativeto_csv('forbes-global2000.csv')into the CWD. This should never become a scheduled builder. Its cell 2 also doesSOURCES_DIR = Path('./sources')andmkdir()s it without ever writing there — dead code that wouldmkdirinto the verysources/tree step 3 creates for LFS-scoped builder inputs.
Also worth knowing for the same PR:
cross_section/README.mdnames the notebook as builder for the two Forbes files only.cities_us.csvandcities_brazil.csvcarry a worldpopulationreview source URL and no builder at all, so step 3 lands two genuinelyunrecovereddatasets beside the two notebook-origin ones — four manifests, three different builder states.Ask. Either settle item 5 here first — the fourth status value, and where the artifact lives (I would keep both out of
scripts/, whichscripts/README.mddefines as the builders directory; its table is stale anyway, listing onlybusiness_cycle.pywhile six committed builders exist) — or land them under an explicitly provisional path with a note on this issue naming it provisional. Without one of those, this issue becomes post-hoc ratification of whatever the fold PR happened to do, in the set that is hardest to revisit.Blocks step 3 of QuantEcon/workspace-lectures#23. Part of #8.
- Both are one-off constructions by this issue's own test — the Forbes lists are a frozen 2020 vintage, and the SCF minis derive from a static deposit
Decision: builders get their own directory, and the enum tells the truth about runnability
Settled the notebook question this issue has been holding, and the answer changed shape once the repo was measured rather than reasoned about. Implemented in #60.
What the repo actually does today
All six
committedbuilders fetch from their third-party upstream at run time —jse.amstat.org,earthquake.usgs.gov,wwwn.cdc.gov,stat.go.jp,openfootball. Not one has a committed input. So the four-stage contract's fetch stage is the established pattern, and there is no precedent here for a builder that reads a local file.That matters because it reframes
sources/. It is not "where builder inputs live" as a general rule — it is the exception layer for an input that cannot be re-fetched.SCF_plus.dtaqualifies for a specific, measured reason: the upstream deposit could not be located at all (Crossref registerslicense: nulland no data relation; DataCite and Harvard Dataverse return zero; openICPSR and the JPE supplement path are 403 behind Cloudflare). If a real deposit URL existed,generating_miniwould fetch from it andsources/would not be needed.So
sources/is defined by un-refetchability, not by size — and emphatically not "the LFS directory", even though the.gitattributesin #57 makes everything under it LFS-tracked.The decision
Builders live in
builders/, one per published dataset, whatever their form. An earlier draft of this proposed putting the notebook-form builders insources/; that conflates input with code and I have dropped it.sources/holds inputs.builders/holds builders.scripts/holds repo tooling and produces no dataset.builders/<stem>.<ext>buildslectures/<stem>.<ext2>. Six of the seven existing builders already matched their output stem exactly, so this makes a latent convention enforceable rather than inventing one. Where a builder produces a set of files it is named for the set and several manifests point at the same path —business_cycle.pywrites three, and bothhigh_dim_databuilders write two, so the very first files this convention meets break naive 1:1.And the enum gains
committed-frozen: the builder is here and deliberately will not run, for a dataset built from a source that must not be refreshed.That is the substantive answer to this issue's question. Neither existing value was true of the two artifacts:
why it is wrong for these committedasserts a runnable four-stage builder. generating_mini.mdhas both itsto_csvwrites commented out and no validate stage;webscrape_forbes.ipynbtargets a frozen 2020 vintage through an undocumented endpoint with a spoofed user-agent and hardcoded GDPR consent cookiesunrecoveredsays the builder was lost — while it arrives in the same PR not-applicableis for verbatim files Since nothing validated the enum —
build_audit.pyrecords the value,build_catalog.pyformats it with astr()fallback — whatever the fold happened to write would have become the precedent silently. #60 adds three assertions: the status must be known, acommitted*status must name a builder, and a named builder must exist on disk.Two consequences for the fold
Do not edit
generating_mini.md. Undercommitted-frozenthe artifact is kept as the record of what produced these bytes, so editing it destroys the property that makes it worth keeping.PLAN.mdand QuantEcon/workspace-lectures#23 both previously instructed uncommenting its twoto_csvwrites; both are now reversed, and #59 carries the PLAN half.The one edit that looked required either way — repointing its input URL off the retired repo — is not required under this reading. The rule it would satisfy ("a builder must read its input from
sources/, never over the network from another QuantEcon repo") binds builders that run; a never-executed record has no live dependency to re-introduce. Record the substitution as prose insources/README.mdinstead: the input is now committed atsources/SCF_plus.dta, and the URL in the file is historical.The two READMEs should not migrate as files.
cross_section/README.md's source tables belong in the manifests'source.*fields andSCF_plus/README.md's variable dictionary belongs inschema.columns[].description— that is what the manifest schema is for, and migrating them as loose files creates a second, unstructured provenance record beside the structured one.cross_section/README.mdadditionally carries a "view raw → copy the url path" how-to, which in an LFS repo yields amedia.githubusercontent.comURL — that instruction is the origin of the 17 media-host reads the fold exists to undo.One thing this exposes that is not yet decided
sources/will have no validation at all — no manifest, no CI check, no hash verification — while every file inlectures/is validated as it migrates. For a file with no consumer that is arguably fine, exceptSCF_plus.dtais the root of the provenance chain for two published datasets: if its bytes drift, both minis become unreproducible and nothing notices.The machinery already exists. #56 rekeyed the consumed-file check onto "hash whenever a hash is recorded", and the same idea extends to the
sha256values insources/README.mdin about ten lines. Worth doing with PR B rather than after it, otherwisesources/is the one tree in the repo where we assert provenance and check nothing.- added 8 commits that reference this issue
on Aug 10, 2026 Status, 2026-09-01: the "when Phase 5 is actioned" list above is two-thirds done.
Step 1 — the template — done in #110, as
builders/_template.pyrather thanscripts/_template_builder.py(builders got their own directory in #60), with the contract section inbuilders/README.mdandAGENTS.md"Builders". Two details differ from the sketch in the body: the overlap window is bounded and reported, not asserted equal — the first dynamic snapshot's source (World Bank WDI) revised 236 of 320 overlap cells between vintages, so equality would fail every refresh by design — andValidationErrormaps to exit code 2 against 1 for a fetch failure, which is how the canary issue tells the two failure classes apart.Step 2 — the
business_cycle.pyretrofit — done in #109, with the two metadata dumps moved toprovenance/(#13).Step 3 — the notebook-origin
builder_status— was settled on 2026-08-10 (comment above, #60):committed-frozen.Also landed in #110: the refresh-as-PR and sources-alive canary (
.github/workflows/refresh-snapshots.yml,scripts/snapshots.py), and the manifest write-back this issue deferred to Phase 5 (snapshots.py stamp).What stays open here is the PR-validation half — a manifest-driven
validate()reading the sidecar'sschemablock as its spec, and the three schema decisions from the 2026-08-06 comments (columnpattern, now in use by one manifest;known_nullsexact-vs-ceiling; the dtype vocabulary, which the pandas-3strcorrection settles). Those are the next Phase 5 slice; the harness extraction (scripts/_builder.py) waits for the second dynamic builder, as the body says.- added a commit that references this issue
on Sep 1, 2026 - added sub-issues
on Sep 2, 2026 - added a commit that references this issue
on Sep 2, 2026 Closed 2026-09-07: every item this tracker carried has landed
The
data-lectures-automationproject is done. What the body and its comments asked for, and where each landed:Asked for Landed A copy-able four-stage builder template with traceback-clean failures and exit codes a canary can read builders/_template.py(#110) — exit 2 for aValidationError, 1 for a fetch failurebusiness_cycle.pyretrofitted to the contract#109; extended to a three-table set in #114 The refresh-as-PR job and the weekly sources-alive canary .github/workflows/refresh-snapshots.yml+scripts/snapshots.py(#110); first real refresh #112; runner User-Agent fix #117The notebook-origin builder_statusquestioncommitted-frozen(#60, decided 2026-08-10)The three schema decisions the first comment flagged — column pattern,known_nullsexact-vs-ceiling, the dtype vocabulary — plus the naming policy#120, #121, #122, #113, decided 2026-09-07 and written into every manifest and manifest-schema.ymlby #124A manifest-driven validate()reading the sidecar'sschemablock, shared by the builders and a PR checkbuilders/_validate.py,scripts/validate_datasets.py,validate-datasets.yml(#119, landed in #126) — both dynamic builders now delegate to it; 44 manifests conformance-checked and 31 CSVs byte-validated on every PRTwo contract bugs the validator found on the way in, both fixed in #126: the FRED refresh would have failed its first monthly run on a publication lag (now
recent: 1, placement-only), andcountries.csvhad an undeclared null from Namibia's ISO codeNAbeing parsed as missing.What is not here, deliberately. The two Phase 5 boxes still open in
PLAN.mdare not this tracker's: consumer fan-out is block D of the P4 work plan (#118) and becomes real at thelecture-wasmflip, and packaging the refresh job as a reusable workflow waits until it has run for a while. Reading xlsx/dta/npy/json ranges the way each lecture does, so those 13 manifests get byte validation too, is a follow-on if it earns one; today they are conformance-checked only, and the workflow log says so.The design notes above stay as written; they are the reasoning this repo's builders follow.
Design notes from a review discussion on PR #12, recorded here to be actioned in PLAN Phase 5 (automation): the supporting research for the builder architecture and a copy-able template. This issue is the tracker for the
data-lectures-automationproject; the open work is its sub-issues.Where we stand (verified 2026-09-07)
The architecture below has landed, and the schema it depends on is now an executable contract rather than a strawman. Landed: the copy-able template
builders/_template.py;builders/business_cycle.pyretrofitted to the four-stage contract on 2026-09-01 (itsvalidate()bounds a revised upstream's overlap window rather than asserting equality);.github/workflows/refresh-snapshots.ymlwith both the refresh-as-PR job and the weekly sources-alive canary;builder_status: committed-frozenfor builders that ran once and will not run again; and, on 2026-09-07, the three schema decisions this tracker had been carrying since its first comment — #120 columnpattern(ordered, contiguous, exhaustive, capture group carries the date), #121 nulls (integers exact, dynamic snapshots declare placement undernulls:), #122 dtypes (pandas-3 names,strfor text, compared by family) — written intomanifest-schema.ymland every manifest by #124, together with the naming policy (#113;business_cycle_data.csvbecamegdp_growth_annual.csvwhile nothing read it). The one open item is #119: the shared manifest-drivenvalidate()that reads the sidecar'sschemablock as its spec, used unchanged by every dynamic builder and by a PR-validation workflow. Its spec is now fully recorded; nothing reads it yet, and the two builders keep their hand-written constants until it lands. Landing #119 closes this tracker. The design notes that follow are kept as written; they are the reasoning, not the work list.Why builders are plain
.pyscripts, not notebooksThe question that prompted this: we could run a notebook to generate a dataset — but extracting tracebacks and failure states from a notebook run is difficult. That difficulty is structural, not incidental, and it settles the architecture:
ValidationError(the data is wrong — needs a human) is not an infrastructure failure (network/credentials — retry). A script makes this a cleanexceptboundary; a notebook blurs it.schemablock is the validation spec. The builder'svalidate()stage reads the dataset's sidecar<filename>.ymland enforces itscolumns,row_count_floor,known_nullsanddate_range. The builder and the PR-CI check then enforce the same contract from one source of truth — no drift between what the manifest claims and what the builder guarantees.os.replace()into place. A failed refresh or a crash mid-write leaves the previously-committed snapshot untouched — which is theAGENTS.mdpromise that an upstream outage may fail a refresh but must never break a lecture build.Notebook-origin datasets
Common for QuantEcon: a dataset's construction starts life in a lecture notebook. The rule:
.pybuilder — or pair the notebook with jupytext, whose.py:percenttwin is the runnable, diffable, traceback-clean artifact that becomes the committedbuilder. The notebook stays in the lecture repo as pedagogy.builderfield, with abuilder_statusthat signals it is not a CI-runnable builder (open sub-question: a distinct status value such asnotebook, vs reusingnot-applicablewith a note — decide when the first such dataset lands).Template
The four-stage contract (
AGENTS.md: fetch → pre-process → validate → write), withpreprocesskept pure (no I/O) so it is unit-testable in isolation:Deliberately deferred
atomic_write,ValidationError, and manifest-drivenvalidate()are obviously shared machinery, but the repo has exactly one dynamic builder today (business_cycle.py). Keep builders self-contained; extractscripts/_builder.pywhen the second one lands.retrieved/integrityon a successful refresh belongs to the refresh-as-PR wiring, i.e. Phase 5 proper, not the template.When Phase 5 is actioned
scripts/_template_builder.pyand add a "Builder architecture" section toscripts/README.md.business_cycle.pyto the four-stage contract — it currently writes three files with no validate stage (AGENTS.mdalready flags this).builder_statussub-question above when the first such dataset lands.Part of #8
See QuantEcon/meta#338