Skip to content

Migrate high-dim-data to this repo #2

Description

@mmcky

(Low Priority)

We should migrate https://github.com/QuantEcon/high_dim_data to be hosted in this repo in a folder high_dimension perhaps.

Activity

  1. mmcky commented on Jul 16, 2026

    @mmcky
    ContributorAuthor

    Status — 2026-07-16

    Still open, and still PLAN Phase 3 — but high_dim_data is in materially better shape for the fold than it was, so noting what changed.

    Its branch-only file is gone. SCF_plus_mini_no_weights.csv existed only on the update_scf_noweights branch, and lecture-python-intro/lectures/mle.md had been reading it from there since 2023. That branch is now merged (QuantEcon/high_dim_data#6, open since 2023-04-28) and deleted, and all four consuming repos are repointed at main — QuantEcon/meta#337 risk 1, now closed. So a fold no longer has to reason about a branch-only file with live consumers.

    The .gitattributes problem is unchanged and is the real work here. high_dim_data carries a blanket *.csv filter=lfs rule, which is why mle.md needs an LFS-aware URL at all. This repo deliberately will not inherit that rule (AGENTS.md), so the fold has to decide per-path LFS for the SCF pair — which is exactly #1.

    Sequencing constraint worth restating (from PLAN.md): enabling LFS on a path breaks every raw.githubusercontent.com URL for it — those URLs return pointer text, not data, so consumers fail with a confusing parse error rather than a 404. Verified live against SCF_plus_mini.csv on high_dim_data main today. Do not LFS-track an existing file until its consumers use a form that survives it (github.com/{org}/{repo}/raw/{ref}/… interim, or Pages final).

    The cross_section/ set (Forbes ×2, cities ×2) plus the SCF pair are pilot P3 of QuantEcon/meta#338 — that pilot is where this fold gets executed and tested, rather than as a separate migration.

  2. mmcky commented on Aug 7, 2026

    @mmcky
    ContributorAuthor

    Status — 2026-08-07: this is the execution home for the fold, and its shape is settled

    Everything the 2026-07-16 comment left open now has an answer. The fold is step 3 of QuantEcon/workspace-lectures#23 and P3 of QuantEcon/meta#338 — that step carries the ordered checklist; this issue is where the work lands.

    Shape

    Note the layout answer first, since the issue body proposes one: no high_dimension folder. The published namespace is flat (settled in Phase 2), so cross_section/ and SCF_plus/ are flattened away and the six datasets land directly in lectures/.

    What Destination Storage Served
    forbes-global2000.csv, forbes-billionaires.csv, cities_us.csv, cities_brazil.csv, SCF_plus_mini.csv, SCF_plus_mini_no_weights.csv lectures/ (flat) plain git yes
    SCF_plus.dta sources/ per-path LFS no
    generating_mini.md, webscrape_forbes.ipynb builder placement open, see below plain git n/a

    The per-path LFS question from the 2026-07-16 comment is answered, and the answer is "not for the SCF pair". At 31.3 MiB and 72.4 MiB both sit comfortably under the blob limit, so lectures/ stays 100% plain git and no consumer can ever meet the raw-vs-media trap. #1 narrows to exactly one path: SCF_plus.dta is 103,934,093 B — 923,507 B (0.88%) under the hard 104,857,600 B limit — so it must stay LFS-tracked permanently, and an upstream vintage 1% larger could not be pushed as plain git at all.

    It deletes nothing, and it is the most reversible set in the programme

    Neither lecture-python-intro nor lecture-wasm holds a copy of any of the six — every read is a cross-repo URL. And archiving high_dim_data preserves serving on both hosts: measured against four archived public LFS repos, media.githubusercontent.com still returns 200 with real bytes and access-control-allow-origin: *. So repoint rule 3's phase 2 and the translation "third beat" from QuantEcon/workspace-lectures#25 are both inapplicable to this set. Deletion bites at step 4, not here.

    The host inversion — and a correction to the count, worth making before #53 merges

    All 21 reads change; 17 of them change host, not 12. The 12 in PLAN.md's rule-6 line and in #50 is the intro + wasm subset, and predates widening the count to three repos — lecture-intro.zh-cn adds five more media-host reads. Verified by grep today:

    Repo on media.githubusercontent.com on github.com/*/raw/ total
    lecture-python-intro 5 — heavy_tails.md:827/854/855/879, _static/…/inequality/data.ipynb:37 2 — mle.md:93, inequality.md:249 7
    lecture-wasm 7 — heavy_tails.md:827/854/855/879, mle.md:95, inequality.md:250, data.ipynb:37 0 7
    lecture-intro.zh-cn 5 — heavy_tails.md:810/837/838/862, data.ipynb:37 2 — mle.md:105, inequality.md:256 7

    The four github.com/*/raw/ reads survive with only an org/repo swap, because that form is a smart redirect that routes per path by LFS status. The 17 media reads do not: the media endpoint also routes per path, not per repo (measured against high_dim_data's own untracked README.md), and after the fold none of these paths is LFS-tracked anywhere. Targets are intro → github.com/QuantEcon/data-lectures/raw/main/lectures/<file>, wasm → raw.githubusercontent.com/… per rule 5, zh-cn by hand.

    The move is also a bandwidth win, not just a chore: the media endpoint serves uncompressed — SCF_plus_mini.csv is 32,853,734 B on the wire even with Accept-Encoding: gzip — while both target hosts compress (realwage.csv here: 121,589 B identity, 14,020 B gzipped from raw, 13,907 B from Pages). For the SCF minis that is roughly a 5x saving on every reader's fetch.

    Gates, in order

    1. Audit: assert on lfs_media and on ref/path resolvability — strict exits 0 on a repoint that 404s every read #54 first. The strict audit exits 0 on a fold left entirely on media.githubusercontent.com: lfs_media is computed at build_audit.py:196 and asserted nowhere — it only feeds a dashboard badge at render_audit.py:841, and branch_pinned has the same shape. Until that lands, rule 6's acceptance grep is the only net.
    2. The lfs: false workflow PR, before the object lands. .github/workflows/audit-dashboard.yml:43 and .github/workflows/consumed-file-check.yml:22 both check out with lfs: true today, and the latter runs on every PR. LFS bandwidth is an org-wide quota shared with high_dim_data, and a 403 there takes out the live media-host reads in all three consuming repos simultaneously — a lecture outage, not a CI failure.
    3. Three hand edits, three repos. build_audit.py:215-216/:222-224 scans lectures/**.md and excludes /_static/, so all three data.ipynb:37 builder reads are invisible to the audit; and the translation sync is .md-only (sync-orchestrator.ts:269), so zh-cn's byte-identical twin can never receive a sync PR. Nothing mechanical catches a miss on these three.
    4. sha256 will not be checked on arrival. check_consumed_files.py:62-63 skips the byte check whenever consumers is empty, and that check is this repo's only required status check — so the "manifest lands ahead of the repoint with empty consumers" convention lands all six files with sha256 unverified. Do the byte-compare explicitly in the fold PR (Phase 7's first bullet) rather than relying on CI.

    Three open decisions

    1. SCF+ provenance — nothing in the fold supplies it (→ #35). SCF_plus/README.md is 4,974 B of variable dictionary: zero URLs, zero dates, no licence, no DOI, and it does not contain the word "source". Only cross_section/README.md carries source tables. So the SCF pair needs the Kuhn–Schularick–Steins (JPE 2020) deposit record, or an honest record — retrieved: null, license.name: null, integrity.upstream.status: unverifiable. Do not stamp retrieved from a commit date; that records when a blob was pushed, not when the data was obtained.

    2. Builder placement (→ #14). Both builders are notebook-origin one-offs, which is precisely #14's deferred builder_status sub-question — "decide when the first such dataset lands", and it lands here. Two facts for that decision: generating_mini.md's two to_csv calls are commented out upstream, so the committed builder fetches and transforms correctly and then writes nothing; and its input URL reads SCF_plus.dta over the network from the repo being retired, so it must repoint at sources/ before archiving.

    3. The webscrape ToS row (→ #35, as its own row). webscrape_forbes.ipynb targets Forbes' undocumented internal API with a spoofed browser user-agent and hardcoded GDPR consent cookies, and the endpoint is still live. That is a republication question about the method, distinct from the data licence #35 already anticipates for the two Forbes CSVs, and it should not be folded into that row.

    One thing worth taking now

    https://quantecon.github.io/data-lectures/lectures/<file> is live today — 200, access-control-allow-origin: *, gzipped, no redirect, correct text/csv — and is immune to both rule 5 and rule 6 regardless of any future sources/ LFS decision. It does not wait on the DNS work in #37. Its one catch: the Pages deploy is gated behind the strict audit (deploy: needs: build), where raw.githubusercontent.com updates on push with no such dependency.

  3. mmcky commented on Aug 31, 2026

    @mmcky
    ContributorAuthor

    Closing — complete. The six high_dim_data datasets landed here as plain git in #62, sources/SCF_plus.dta in #63, all 28 consuming reads across four repos were repointed on 2026-08-11, and the tracker flipped in #69. high_dim_data is archived, not deleted, and still serves on both hosts. The full record is PLAN.md Phase 3 and repoint rule 6.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions