Skip to content

Builder architecture: four-stage template, traceback-clean failures, notebook policy (Phase 5) #14

Description

@mmcky

Design notes from a review discussion on PR #12, recorded here to be actioned in PLAN Phase 5 (automation): the supporting research for the builder architecture and a copy-able template. This issue is the tracker for the data-lectures-automation project; the open work is its sub-issues.

Where we stand (verified 2026-09-07)

The architecture below has landed, and the schema it depends on is now an executable contract rather than a strawman. Landed: the copy-able template builders/_template.py; builders/business_cycle.py retrofitted to the four-stage contract on 2026-09-01 (its validate() bounds a revised upstream's overlap window rather than asserting equality); .github/workflows/refresh-snapshots.yml with both the refresh-as-PR job and the weekly sources-alive canary; builder_status: committed-frozen for builders that ran once and will not run again; and, on 2026-09-07, the three schema decisions this tracker had been carrying since its first comment — #120 column pattern (ordered, contiguous, exhaustive, capture group carries the date), #121 nulls (integers exact, dynamic snapshots declare placement under nulls:), #122 dtypes (pandas-3 names, str for text, compared by family) — written into manifest-schema.yml and every manifest by #124, together with the naming policy (#113; business_cycle_data.csv became gdp_growth_annual.csv while nothing read it). The one open item is #119: the shared manifest-driven validate() that reads the sidecar's schema block as its spec, used unchanged by every dynamic builder and by a PR-validation workflow. Its spec is now fully recorded; nothing reads it yet, and the two builders keep their hand-written constants until it lands. Landing #119 closes this tracker. The design notes that follow are kept as written; they are the reasoning, not the work list.

Why builders are plain .py scripts, not notebooks

The question that prompted this: we could run a notebook to generate a dataset — but extracting tracebacks and failure states from a notebook run is difficult. That difficulty is structural, not incidental, and it settles the architecture:

  1. Errors must propagate as exceptions with a full traceback and a non-zero exit. A plain script gives CI exactly that for free. Papermill/nbconvert can execute a notebook, but a failure surfaces as a truncated cell-execution error embedded in JSON output — the natural stack trace is lost, and the exit-code contract is murkier.
  2. CI needs to distinguish two failure classes. A ValidationError (the data is wrong — needs a human) is not an infrastructure failure (network/credentials — retry). A script makes this a clean except boundary; a notebook blurs it.
  3. The manifest's schema block is the validation spec. The builder's validate() stage reads the dataset's sidecar <filename>.yml and enforces its columns, row_count_floor, known_nulls and date_range. The builder and the PR-CI check then enforce the same contract from one source of truth — no drift between what the manifest claims and what the builder guarantees.
  4. Atomic write + last-good guarantee. Validate the in-memory frame before writing; write to a temp file in the same directory, then os.replace() into place. A failed refresh or a crash mid-write leaves the previously-committed snapshot untouched — which is the AGENTS.md promise that an upstream outage may fail a refresh but must never break a lecture build.

Notebook-origin datasets

Common for QuantEcon: a dataset's construction starts life in a lecture notebook. The rule:

  • If it needs to refresh (dynamic snapshot): port the notebook logic to a plain .py builder — or pair the notebook with jupytext, whose .py:percent twin is the runnable, diffable, traceback-clean artifact that becomes the committed builder. The notebook stays in the lecture repo as pedagogy.
  • If it was a one-off construction that will not refresh: record the notebook as provenance in the manifest's builder field, with a builder_status that signals it is not a CI-runnable builder (open sub-question: a distinct status value such as notebook, vs reusing not-applicable with a note — decide when the first such dataset lands).
  • Papermill-executing a notebook as the builder is the fallback only when the notebook itself must be the deliverable, accepting the worse failure ergonomics.

Template

The four-stage contract (AGENTS.md: fetch → pre-process → validate → write), with preprocess kept pure (no I/O) so it is unit-testable in isolation:

"""Builder for <output>.csv — <one-line description>.

Contract (AGENTS.md / PLAN Phase 5): fetch -> pre-process -> validate -> write.
Writes ONLY on validation pass, atomically, so a failed refresh leaves the
last-good snapshot untouched and never breaks a lecture build.

Run:  python scripts/<name>.py
Exit: 0 on success; non-zero WITH a traceback on any failure (CI-readable).
"""
from __future__ import annotations
import logging, os, tempfile
import pandas as pd, yaml
# import wbgapi as wb   # or the relevant upstream client

HERE = os.path.dirname(os.path.abspath(__file__))
LECTURES = os.path.join(os.path.dirname(HERE), "lectures")
OUTPUT = "example.csv"              # the published filename
MANIFEST = OUTPUT + ".yml"          # its sidecar — the validation spec

log = logging.getLogger(OUTPUT)


class ValidationError(Exception):
    """Data failed an invariant — distinct from an infrastructure failure, so CI
    can tell 'the data is wrong' (human) from 'the network was down' (retry)."""


def fetch():
    """Pull raw data from upstream. Network/credential failures surface here."""
    raise NotImplementedError


def preprocess(raw) -> pd.DataFrame:
    """Pure raw -> published frame. No I/O, so it's unit-testable in isolation."""
    raise NotImplementedError


def validate(df: pd.DataFrame, spec: dict, previous: pd.DataFrame | None) -> None:
    """Enforce the manifest's `schema` block — the single source of truth."""
    schema = spec["schema"]
    expected = [c["name"] for c in schema["columns"]]
    if list(df.columns) != expected:
        raise ValidationError(f"columns {list(df.columns)} != manifest {expected}")
    if len(df) < schema["row_count_floor"]:
        raise ValidationError(f"{len(df)} rows < floor {schema['row_count_floor']}")

    known = schema.get("known_nulls") or {}
    for col in df.columns:
        n = int(df[col].isna().sum())
        if n != known.get(col, 0):
            raise ValidationError(f"{col}: {n} nulls, manifest allows {known.get(col, 0)}")

    # Overlap window: history must not change vs the last-good vintage.
    if previous is not None:
        common = previous.index.intersection(df.index)
        a, b = df.loc[common], previous.loc[common]
        # NOTE: float columns need a tolerance — use np.isclose, not `!=`.
        changed = ((a != b) & a.notna() & b.notna()).to_numpy().any()
        if changed:
            raise ValidationError("values changed in the overlap window vs last-good")


def write(df: pd.DataFrame, path: str) -> None:
    """Atomic write: temp file in the same dir, then os.replace() (atomic on POSIX).
    A crash mid-write can never leave a partial file where a lecture would read it."""
    fd, tmp = tempfile.mkstemp(dir=os.path.dirname(path), suffix=".tmp")
    os.close(fd)
    try:
        df.to_csv(tmp)
        os.replace(tmp, path)
    finally:
        if os.path.exists(tmp):
            os.remove(tmp)


def run() -> None:
    out = os.path.join(LECTURES, OUTPUT)
    spec = yaml.safe_load(open(os.path.join(LECTURES, MANIFEST)))
    previous = pd.read_csv(out, index_col=0) if os.path.exists(out) else None

    log.info("fetch");        raw = fetch()
    log.info("pre-process");  df = preprocess(raw)
    log.info("validate");     validate(df, spec, previous)   # raises -> no write
    log.info("write");        write(df, out)
    log.info("ok: %d rows", len(df))


if __name__ == "__main__":
    logging.basicConfig(level=logging.INFO, format="%(name)s %(levelname)s %(message)s")
    run()   # uncaught exceptions -> full traceback + non-zero exit (CI-friendly)

Deliberately deferred

  • A shared harness. atomic_write, ValidationError, and manifest-driven validate() are obviously shared machinery, but the repo has exactly one dynamic builder today (business_cycle.py). Keep builders self-contained; extract scripts/_builder.py when the second one lands.
  • Manifest write-back — stamping retrieved / integrity on a successful refresh belongs to the refresh-as-PR wiring, i.e. Phase 5 proper, not the template.

When Phase 5 is actioned

  1. Commit the template as scripts/_template_builder.py and add a "Builder architecture" section to scripts/README.md.
  2. Retrofit business_cycle.py to the four-stage contract — it currently writes three files with no validate stage (AGENTS.md already flags this).
  3. Decide the notebook-origin builder_status sub-question above when the first such dataset lands.

Part of #8
See QuantEcon/meta#338

Activity

  1. mmcky commented on Jul 16, 2026

    @mmcky
    ContributorAuthor

    Schema implications of this architecture

    Writing the template surfaced a consequence worth stating outright: the moment a builder's validate() reads the sidecar manifest and enforces it, the manifest's schema block stops being documentation and becomes an executable contract. Several fields are not yet precise enough to execute. None of this needs action now — the schema is a strawman and says so — but these are exactly what the next pilot and Phase 5 should test the schema against, so they are recorded here rather than lost.

    1. Column patterns — a real gap our own flagship file already hits. The template checks list(df.columns) != expected, but the schema example for business_cycle_data.csv lists YR1960 with the comment "one column per year" — a single exemplar standing in for ~64 columns. A machine validator cannot read that comment; it would fail the file. The schema needs a way to express repeated columns before validation can be real, e.g.:

    columns:
      - {name: economy, dtype: string, description: ISO3 country code}
      - {name: Country, dtype: string, description: Country name}
      - {pattern: 'YR\d{4}', dtype: float64, description: Annual GDP growth (%), one column per year}

    2. known_nulls semantics need pinning: exact count or ceiling? The template enforces exact (n != known), the stronger invariant and right for frozen files like lingcod (F_over_Fmsy: 1, always). But for a dynamic snapshot, null counts can legitimately drift between refreshes — a new terminal assessment year appears and the null moves or multiplies. Likely answer: exact for static extracts, ceiling for dynamic snapshots — or a per-column form like {F_over_Fmsy: {max: 3}}. Needs an explicit decision, not a silent default.

    3. dtype vocabulary is currently inconsistent. The lingcod manifest says int64 / float64 (pandas names); the schema sketch says string / float (loose names). A human reader doesn't care; validate() must compare against something canonical. Recommendation: standardise on pandas dtype strings, since that is what the checking code actually sees.

    4. Float tolerance for the overlap window. "Values unchanged vs last-good" cannot be != on floats (the template already flags this — np.isclose needed). If tolerance ever needs to vary by dataset it becomes a schema field (overlap_tolerance); until then a builder-side default suffices. Awareness item, not a field yet.

    5. builder_status for notebook-origin datasets — already an open sub-question in the issue body; noted here only so the schema-facing list is complete in one place.

    Items 1–3 are decisions the schema has to make before Phase 5's validation can run; 4–5 are cheap to defer. The P2 pilot is the natural place to test 1–3 against a second real file.

  2. mmcky commented on Aug 6, 2026

    @mmcky
    ContributorAuthor

    These three are now blockers, not deferrals — and item 3 has already drifted

    Recording where items 1–3 stand against main today (2026-08-06), because the framing has changed since this was written. When this issue was opened the schema was a strawman with one real file behind it. There are now ten manifests, and the next substantive piece of work (the P4 dynamic snapshot) cannot be written until these are settled — validate() reads the sidecar as its spec, so an imprecise sidecar is an unwritable builder. Treating them as "decide when convenient" means P4 stalls on them mid-implementation.

    Concrete recommendations below, with what the ten manifests actually show.

    1. Column patterns — confirmed, and the flagship file is worse than described

    business_cycle_data.csv has 66 columns: economy, Country, then YR1960 … YR2023. Enumerating them in the manifest is both unreadable and wrong — it hard-codes an end year into a file whose whole purpose is to grow annually. A pattern entry is the only sane form.

    Recommendation: a columns entry carries either name (exact) or pattern (a regex), never both, and the validator walks the entries in order, letting a pattern entry consume one or more consecutive columns. That keeps column order checkable — which matters, since positional reads exist in the wild — while making the repeated block expressible.

    For business_cycle_data.csv that is three entries instead of sixty-six, and it stays correct when 2024 lands. A pattern entry should also require min_count (here 1) so an upstream change that silently drops every year column still fails.

    2. known_nulls — exact vs ceiling splits cleanly on class

    All ten manifests use exact counts today, and every one of them is a static file where that is right: lingcod F_over_Fmsy: 1, realwage value: 68, countries Capital: 248, employ four columns at 1080 each. Seven declare known_nulls: {} with a prose note explaining why zero is meaningful.

    Nothing has yet needed a ceiling — because the repo has no dynamic snapshot under validation. UNRATE will be the first, and a monthly refresh legitimately changes null counts when a terminal observation is revised.

    Recommendation: keep the plain integer meaning exact, and add {max: N} as the ceiling form, with the default chosen by class rather than per file — exact for verbatim and constructed, ceiling for dynamic snapshot. Class already exists in every manifest, so this needs no new field and no per-file judgment. A dynamic dataset that genuinely wants an exact count can still write the integer form and get the stronger check.

    3. dtype vocabulary — this has already drifted, and pandas 3 makes it urgent

    This is the one that stopped being theoretical. Across the ten manifests: 21 string, 19 int64, 12 float64, 7 object — and the split is chronological, not principled:

    Manifests Text dtype used
    countries.csv, employ.csv, realwage.csv (P2, July) string
    ames_house_prices.csv, epl_match_goals.csv, japan_earthquakes.csv, us_adult_heights.csv (P5, August) object

    Nobody decided to change it; the convention flipped between pilots. Two names for one concept in a field that is about to become executable is exactly the drift this issue predicted.

    The pandas 3 angle makes the choice for us. The lecture repos are on anaconda=2026.07, i.e. pandas 3.x, where the default dtype for text read from CSV is the new string dtype rather than object. A manifest asserting object will therefore be wrong against a default read_csv under the pin the lectures already use — the seven object entries are legacy the day they were written, not a live alternative.

    Recommendation: canonicalise on the pandas dtype string as str(df[col].dtype) produces it, with string for text, and have validate() normalise object and string to one another when comparing so the check does not become version-brittle across the pandas 2/3 boundary. Then fix the seven object entries in a single sweep — it is a comment-level change to four manifests with no data impact, cheapest done now while it is four files rather than forty.

    Suggested order

    Land 1–3 in manifest-schema.yml as one PR before any P4 builder code is written, then retrofit the four object manifests in the same PR so the repo is internally consistent from that point. Items 4 (float tolerance) and 5 (notebook-origin builder_status) stay deferred as this issue already argues — 4 has a sensible builder-side default, and 5 has no instance yet.

    Part of #8. Gates the P4 work in PLAN.md Phase 5 / Phase 8, and the business_cycle_data.csv manifest discussed in #13.

  3. mmcky commented on Aug 6, 2026

    @mmcky
    ContributorAuthor

    Correction to item 3: the pandas 3 name is str, not string

    My earlier comment recommended canonicalising text columns on string. That was wrong, and writing the eight manifests in #38 surfaced it — I now have a pandas 3 environment matching the lectures' anaconda=2026.07 pin, and measured it rather than reasoning about it.

    Under pandas 3.0.5, str(df[col].dtype) for a text column read by read_csv is str. All three names exist and are distinct:

    Construction str(dtype)
    read_csv text column (pandas 3 default) str
    pd.Series([...], dtype='object') object
    pd.Series([...], dtype='string') string

    So there are now three candidate spellings in play, not two, and the manifests currently use two of them — 21 string from the P2 pilot and 7 object from P5 — neither of which is what a default read actually produces under the pin the lectures run on. A validate() comparing str(df[col].dtype) against either would fail on every text column.

    Revised recommendation. Canonicalise on str, since that is what the checking code will actually see. Have validate() normalise str, string and object to one another for text comparison, so the check survives the pandas 2/3 boundary and does not become a version tripwire — this matters because lecture-datascience.myst is still on anaconda=2024.06, where the same read yields object. Then sweep the existing manifests once.

    The new manifests in #38 use str for text columns (mpd2020.xlsx's countrycode and country), measured against the file rather than assumed, so they are already on the proposed canonical form.

    Two schema cases #38 also surfaces

    These are the first multi-sheet workbooks to get manifests, and they do not fit the flat-table sketch at all.

    Sheets, not just columns. A workbook's contract is which sheets and cell ranges the lectures read. dette.xlsx has 27 sheets, of which the lecture reads five ranges across two. The manifests use a sheets: list where each entry records read_as (the exact pandas call), the resulting shape, dtypes and null counts — and names the unread sheets so nobody strips them as dead weight when they are in fact the authors' working provenance.

    Positional reads are a fragility class the schema has no vocabulary for. Every French Revolution read is header=None with explicit usecols/skiprows/nrows — for instance pd.read_excel(dette_url, sheet_name='Militspe', usecols='M:X', skiprows=7, nrows=102, header=None). A row or column inserted anywhere above or left of that range silently shifts what the lecture plots, with no error. Column-name validation cannot detect it, because there are no column names. The manifests flag it with positional_reads: true, but it is worth deciding whether the schema should express the invariant properly — a checksum over the read block would catch what a column check cannot.

    That last one strengthens the case for settling items 1–3 before Phase 5 validation is written: the schema currently has no way to state the invariant that actually protects these four files.

  4. mmcky commented on Aug 7, 2026

    @mmcky
    ContributorAuthor

    Deferred item 5 stops being deferred at step 3 — decide it before the fold PR, not after

    Step 3 of QuantEcon/workspace-lectures#23 (fold in high_dim_data) migrates the builders as well as the data, and both builders are notebook-form: cross_section/webscrape_forbes.ipynb (5,861 B, 8 cells, no stored outputs) and SCF_plus/generating_mini.md — the latter a jupytext MyST twin (format_name: myst, 0.13, jupytext 1.14.1), not the .py:percent twin this issue's notebook policy names. So the "decide when the first such dataset lands" sub-question lands now, inside the largest remaining migration set, and it is the first thing in this repo that its answer has to fit.

    Today the repo has exactly one builder shape and nothing else: seven plain .py files under scripts/, zero .ipynb anywhere in the tree, and zero mentions of jupytext, papermill or nbconvert in AGENTS.md, scripts/README.md or manifest-schema.yml (requirements.txt is two lines). Step 3 introduces a second and a third form in one PR.

    The enum has no cell for it, and nothing checks. manifest-schema.yml:178 allows committed | unrecovered | not-applicable. build_audit.py:489 only records the value and build_catalog.py:111 only formats it — no check enforces the vocabulary, so whatever step 3 writes lands silently. Both usable values are false here: committed asserts a CI-runnable builder of the shape defined above, and unrecovered says the builder was lost when it is arriving in the same PR.

    generating_mini.md makes committed untenable on its own terms: both to_csv calls are commented out upstream (# df1.to_csv('SCF_plus_mini.csv', index=None), # df2.to_csv('SCF_plus_mini_no_weights.csv', index=None)), and two cells are bare display expressions (df1, df2). Fetch and transform are complete; write and validate do not exist. That is the same retrofit shape as business_cycle.py — item 2 of this issue's "When Phase 5 is actioned" — which is worth pairing, because business_cycle_data.csv has no sidecar manifest (18 manifests for 21 files in lectures/), so its validate() has no schema block to read either. Two Phase 5 retrofits, one missing spec between them.

    Two things argue for "provenance, not runnable builder" rather than a port:

    • Both are one-off constructions by this issue's own test — the Forbes lists are a frozen 2020 vintage, and the SCF minis derive from a static deposit .dta. Neither refreshes, so the port-to-.py branch does not apply.
    • webscrape_forbes.ipynb fetches Forbes' undocumented internal API with a spoofed browser user-agent and hardcoded GDPR consent cookies (notice_gdpr_prefs), and writes with bare relative to_csv('forbes-global2000.csv') into the CWD. This should never become a scheduled builder. Its cell 2 also does SOURCES_DIR = Path('./sources') and mkdir()s it without ever writing there — dead code that would mkdir into the very sources/ tree step 3 creates for LFS-scoped builder inputs.

    Also worth knowing for the same PR: cross_section/README.md names the notebook as builder for the two Forbes files only. cities_us.csv and cities_brazil.csv carry a worldpopulationreview source URL and no builder at all, so step 3 lands two genuinely unrecovered datasets beside the two notebook-origin ones — four manifests, three different builder states.

    Ask. Either settle item 5 here first — the fourth status value, and where the artifact lives (I would keep both out of scripts/, which scripts/README.md defines as the builders directory; its table is stale anyway, listing only business_cycle.py while six committed builders exist) — or land them under an explicitly provisional path with a note on this issue naming it provisional. Without one of those, this issue becomes post-hoc ratification of whatever the fold PR happened to do, in the set that is hardest to revisit.

    Blocks step 3 of QuantEcon/workspace-lectures#23. Part of #8.

  5. mmcky commented on Aug 10, 2026

    @mmcky
    ContributorAuthor

    Decision: builders get their own directory, and the enum tells the truth about runnability

    Settled the notebook question this issue has been holding, and the answer changed shape once the repo was measured rather than reasoned about. Implemented in #60.

    What the repo actually does today

    All six committed builders fetch from their third-party upstream at run time — jse.amstat.org, earthquake.usgs.gov, wwwn.cdc.gov, stat.go.jp, openfootball. Not one has a committed input. So the four-stage contract's fetch stage is the established pattern, and there is no precedent here for a builder that reads a local file.

    That matters because it reframes sources/. It is not "where builder inputs live" as a general rule — it is the exception layer for an input that cannot be re-fetched. SCF_plus.dta qualifies for a specific, measured reason: the upstream deposit could not be located at all (Crossref registers license: null and no data relation; DataCite and Harvard Dataverse return zero; openICPSR and the JPE supplement path are 403 behind Cloudflare). If a real deposit URL existed, generating_mini would fetch from it and sources/ would not be needed.

    So sources/ is defined by un-refetchability, not by size — and emphatically not "the LFS directory", even though the .gitattributes in #57 makes everything under it LFS-tracked.

    The decision

    Builders live in builders/, one per published dataset, whatever their form. An earlier draft of this proposed putting the notebook-form builders in sources/; that conflates input with code and I have dropped it. sources/ holds inputs. builders/ holds builders. scripts/ holds repo tooling and produces no dataset.

    builders/<stem>.<ext> builds lectures/<stem>.<ext2>. Six of the seven existing builders already matched their output stem exactly, so this makes a latent convention enforceable rather than inventing one. Where a builder produces a set of files it is named for the set and several manifests point at the same path — business_cycle.py writes three, and both high_dim_data builders write two, so the very first files this convention meets break naive 1:1.

    And the enum gains committed-frozen: the builder is here and deliberately will not run, for a dataset built from a source that must not be refreshed.

    That is the substantive answer to this issue's question. Neither existing value was true of the two artifacts:

    why it is wrong for these
    committed asserts a runnable four-stage builder. generating_mini.md has both its to_csv writes commented out and no validate stage; webscrape_forbes.ipynb targets a frozen 2020 vintage through an undocumented endpoint with a spoofed user-agent and hardcoded GDPR consent cookies
    unrecovered says the builder was lost — while it arrives in the same PR
    not-applicable is for verbatim files

    Since nothing validated the enum — build_audit.py records the value, build_catalog.py formats it with a str() fallback — whatever the fold happened to write would have become the precedent silently. #60 adds three assertions: the status must be known, a committed* status must name a builder, and a named builder must exist on disk.

    Two consequences for the fold

    Do not edit generating_mini.md. Under committed-frozen the artifact is kept as the record of what produced these bytes, so editing it destroys the property that makes it worth keeping. PLAN.md and QuantEcon/workspace-lectures#23 both previously instructed uncommenting its two to_csv writes; both are now reversed, and #59 carries the PLAN half.

    The one edit that looked required either way — repointing its input URL off the retired repo — is not required under this reading. The rule it would satisfy ("a builder must read its input from sources/, never over the network from another QuantEcon repo") binds builders that run; a never-executed record has no live dependency to re-introduce. Record the substitution as prose in sources/README.md instead: the input is now committed at sources/SCF_plus.dta, and the URL in the file is historical.

    The two READMEs should not migrate as files. cross_section/README.md's source tables belong in the manifests' source.* fields and SCF_plus/README.md's variable dictionary belongs in schema.columns[].description — that is what the manifest schema is for, and migrating them as loose files creates a second, unstructured provenance record beside the structured one. cross_section/README.md additionally carries a "view raw → copy the url path" how-to, which in an LFS repo yields a media.githubusercontent.com URL — that instruction is the origin of the 17 media-host reads the fold exists to undo.

    One thing this exposes that is not yet decided

    sources/ will have no validation at all — no manifest, no CI check, no hash verification — while every file in lectures/ is validated as it migrates. For a file with no consumer that is arguably fine, except SCF_plus.dta is the root of the provenance chain for two published datasets: if its bytes drift, both minis become unreproducible and nothing notices.

    The machinery already exists. #56 rekeyed the consumed-file check onto "hash whenever a hash is recorded", and the same idea extends to the sha256 values in sources/README.md in about ten lines. Worth doing with PR B rather than after it, otherwise sources/ is the one tree in the repo where we assert provenance and check nothing.

  6. mmcky commented on Sep 1, 2026

    @mmcky
    ContributorAuthor

    Status, 2026-09-01: the "when Phase 5 is actioned" list above is two-thirds done.

    Step 1 — the template — done in #110, as builders/_template.py rather than scripts/_template_builder.py (builders got their own directory in #60), with the contract section in builders/README.md and AGENTS.md "Builders". Two details differ from the sketch in the body: the overlap window is bounded and reported, not asserted equal — the first dynamic snapshot's source (World Bank WDI) revised 236 of 320 overlap cells between vintages, so equality would fail every refresh by design — and ValidationError maps to exit code 2 against 1 for a fetch failure, which is how the canary issue tells the two failure classes apart.

    Step 2 — the business_cycle.py retrofit — done in #109, with the two metadata dumps moved to provenance/ (#13).

    Step 3 — the notebook-origin builder_status — was settled on 2026-08-10 (comment above, #60): committed-frozen.

    Also landed in #110: the refresh-as-PR and sources-alive canary (.github/workflows/refresh-snapshots.yml, scripts/snapshots.py), and the manifest write-back this issue deferred to Phase 5 (snapshots.py stamp).

    What stays open here is the PR-validation half — a manifest-driven validate() reading the sidecar's schema block as its spec, and the three schema decisions from the 2026-08-06 comments (column pattern, now in use by one manifest; known_nulls exact-vs-ceiling; the dtype vocabulary, which the pandas-3 str correction settles). Those are the next Phase 5 slice; the harness extraction (scripts/_builder.py) waits for the second dynamic builder, as the body says.

  7. added theissue type on Sep 2, 2026
  8. added a commit that references this issue on Sep 2, 2026
  9. mmcky commented on Sep 7, 2026

    @mmcky
    ContributorAuthor

    Closed 2026-09-07: every item this tracker carried has landed

    The data-lectures-automation project is done. What the body and its comments asked for, and where each landed:

    Asked for Landed
    A copy-able four-stage builder template with traceback-clean failures and exit codes a canary can read builders/_template.py (#110) — exit 2 for a ValidationError, 1 for a fetch failure
    business_cycle.py retrofitted to the contract #109; extended to a three-table set in #114
    The refresh-as-PR job and the weekly sources-alive canary .github/workflows/refresh-snapshots.yml + scripts/snapshots.py (#110); first real refresh #112; runner User-Agent fix #117
    The notebook-origin builder_status question committed-frozen (#60, decided 2026-08-10)
    The three schema decisions the first comment flagged — column pattern, known_nulls exact-vs-ceiling, the dtype vocabulary — plus the naming policy #120, #121, #122, #113, decided 2026-09-07 and written into every manifest and manifest-schema.yml by #124
    A manifest-driven validate() reading the sidecar's schema block, shared by the builders and a PR check builders/_validate.py, scripts/validate_datasets.py, validate-datasets.yml (#119, landed in #126) — both dynamic builders now delegate to it; 44 manifests conformance-checked and 31 CSVs byte-validated on every PR

    Two contract bugs the validator found on the way in, both fixed in #126: the FRED refresh would have failed its first monthly run on a publication lag (now recent: 1, placement-only), and countries.csv had an undeclared null from Namibia's ISO code NA being parsed as missing.

    What is not here, deliberately. The two Phase 5 boxes still open in PLAN.md are not this tracker's: consumer fan-out is block D of the P4 work plan (#118) and becomes real at the lecture-wasm flip, and packaging the refresh job as a reusable workflow waits until it has run for a while. Reading xlsx/dta/npy/json ranges the way each lecture does, so those 13 manifests get byte validation too, is a follow-on if it earns one; today they are conformance-checked only, and the workflow log says so.

    The design notes above stay as written; they are the reasoning this repo's builders follow.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions