Repository navigation
Migrate high-dim-data to this repo #2
Description
Activity
Status — 2026-07-16
Still open, and still PLAN Phase 3 — but
high_dim_datais in materially better shape for the fold than it was, so noting what changed.Its branch-only file is gone.
SCF_plus_mini_no_weights.csvexisted only on theupdate_scf_noweightsbranch, andlecture-python-intro/lectures/mle.mdhad been reading it from there since 2023. That branch is now merged (QuantEcon/high_dim_data#6, open since 2023-04-28) and deleted, and all four consuming repos are repointed atmain— QuantEcon/meta#337 risk 1, now closed. So a fold no longer has to reason about a branch-only file with live consumers.The
.gitattributesproblem is unchanged and is the real work here.high_dim_datacarries a blanket*.csv filter=lfsrule, which is whymle.mdneeds an LFS-aware URL at all. This repo deliberately will not inherit that rule (AGENTS.md), so the fold has to decide per-path LFS for the SCF pair — which is exactly #1.Sequencing constraint worth restating (from
PLAN.md): enabling LFS on a path breaks everyraw.githubusercontent.comURL for it — those URLs return pointer text, not data, so consumers fail with a confusing parse error rather than a 404. Verified live againstSCF_plus_mini.csvonhigh_dim_datamaintoday. Do not LFS-track an existing file until its consumers use a form that survives it (github.com/{org}/{repo}/raw/{ref}/…interim, or Pages final).The
cross_section/set (Forbes ×2, cities ×2) plus the SCF pair are pilot P3 of QuantEcon/meta#338 — that pilot is where this fold gets executed and tested, rather than as a separate migration.Status — 2026-08-07: this is the execution home for the fold, and its shape is settled
Everything the 2026-07-16 comment left open now has an answer. The fold is step 3 of QuantEcon/workspace-lectures#23 and P3 of QuantEcon/meta#338 — that step carries the ordered checklist; this issue is where the work lands.
Shape
Note the layout answer first, since the issue body proposes one: no
high_dimensionfolder. The published namespace is flat (settled in Phase 2), socross_section/andSCF_plus/are flattened away and the six datasets land directly inlectures/.What Destination Storage Served forbes-global2000.csv,forbes-billionaires.csv,cities_us.csv,cities_brazil.csv,SCF_plus_mini.csv,SCF_plus_mini_no_weights.csvlectures/(flat)plain git yes SCF_plus.dtasources/per-path LFS no generating_mini.md,webscrape_forbes.ipynbbuilder placement open, see below plain git n/a The per-path LFS question from the 2026-07-16 comment is answered, and the answer is "not for the SCF pair". At 31.3 MiB and 72.4 MiB both sit comfortably under the blob limit, so
lectures/stays 100% plain git and no consumer can ever meet the raw-vs-media trap. #1 narrows to exactly one path:SCF_plus.dtais 103,934,093 B — 923,507 B (0.88%) under the hard 104,857,600 B limit — so it must stay LFS-tracked permanently, and an upstream vintage 1% larger could not be pushed as plain git at all.It deletes nothing, and it is the most reversible set in the programme
Neither
lecture-python-intronorlecture-wasmholds a copy of any of the six — every read is a cross-repo URL. And archivinghigh_dim_datapreserves serving on both hosts: measured against four archived public LFS repos,media.githubusercontent.comstill returns 200 with real bytes andaccess-control-allow-origin: *. So repoint rule 3's phase 2 and the translation "third beat" from QuantEcon/workspace-lectures#25 are both inapplicable to this set. Deletion bites at step 4, not here.The host inversion — and a correction to the count, worth making before #53 merges
All 21 reads change; 17 of them change host, not 12. The 12 in
PLAN.md's rule-6 line and in #50 is the intro + wasm subset, and predates widening the count to three repos —lecture-intro.zh-cnadds five more media-host reads. Verified by grep today:Repo on media.githubusercontent.comon github.com/*/raw/total lecture-python-intro5 — heavy_tails.md:827/854/855/879,_static/…/inequality/data.ipynb:372 — mle.md:93,inequality.md:2497 lecture-wasm7 — heavy_tails.md:827/854/855/879,mle.md:95,inequality.md:250,data.ipynb:370 7 lecture-intro.zh-cn5 — heavy_tails.md:810/837/838/862,data.ipynb:372 — mle.md:105,inequality.md:2567 The four
github.com/*/raw/reads survive with only an org/repo swap, because that form is a smart redirect that routes per path by LFS status. The 17 media reads do not: the media endpoint also routes per path, not per repo (measured againsthigh_dim_data's own untrackedREADME.md), and after the fold none of these paths is LFS-tracked anywhere. Targets are intro →github.com/QuantEcon/data-lectures/raw/main/lectures/<file>, wasm →raw.githubusercontent.com/…per rule 5, zh-cn by hand.The move is also a bandwidth win, not just a chore: the media endpoint serves uncompressed —
SCF_plus_mini.csvis 32,853,734 B on the wire even withAccept-Encoding: gzip— while both target hosts compress (realwage.csvhere: 121,589 B identity, 14,020 B gzipped fromraw, 13,907 B from Pages). For the SCF minis that is roughly a 5x saving on every reader's fetch.Gates, in order
- Audit: assert on lfs_media and on ref/path resolvability — strict exits 0 on a repoint that 404s every read #54 first. The strict audit exits 0 on a fold left entirely on
media.githubusercontent.com:lfs_mediais computed atbuild_audit.py:196and asserted nowhere — it only feeds a dashboard badge atrender_audit.py:841, andbranch_pinnedhas the same shape. Until that lands, rule 6's acceptance grep is the only net. - The
lfs: falseworkflow PR, before the object lands..github/workflows/audit-dashboard.yml:43and.github/workflows/consumed-file-check.yml:22both check out withlfs: truetoday, and the latter runs on every PR. LFS bandwidth is an org-wide quota shared withhigh_dim_data, and a 403 there takes out the live media-host reads in all three consuming repos simultaneously — a lecture outage, not a CI failure. - Three hand edits, three repos.
build_audit.py:215-216/:222-224scanslectures/**.mdand excludes/_static/, so all threedata.ipynb:37builder reads are invisible to the audit; and the translation sync is.md-only (sync-orchestrator.ts:269), so zh-cn's byte-identical twin can never receive a sync PR. Nothing mechanical catches a miss on these three. sha256will not be checked on arrival.check_consumed_files.py:62-63skips the byte check wheneverconsumersis empty, and that check is this repo's only required status check — so the "manifest lands ahead of the repoint with empty consumers" convention lands all six files withsha256unverified. Do the byte-compare explicitly in the fold PR (Phase 7's first bullet) rather than relying on CI.
Three open decisions
1. SCF+ provenance — nothing in the fold supplies it (→ #35).
SCF_plus/README.mdis 4,974 B of variable dictionary: zero URLs, zero dates, no licence, no DOI, and it does not contain the word "source". Onlycross_section/README.mdcarries source tables. So the SCF pair needs the Kuhn–Schularick–Steins (JPE 2020) deposit record, or an honest record —retrieved: null,license.name: null,integrity.upstream.status: unverifiable. Do not stampretrievedfrom a commit date; that records when a blob was pushed, not when the data was obtained.2. Builder placement (→ #14). Both builders are notebook-origin one-offs, which is precisely #14's deferred
builder_statussub-question — "decide when the first such dataset lands", and it lands here. Two facts for that decision:generating_mini.md's twoto_csvcalls are commented out upstream, so the committed builder fetches and transforms correctly and then writes nothing; and its input URL readsSCF_plus.dtaover the network from the repo being retired, so it must repoint atsources/before archiving.3. The webscrape ToS row (→ #35, as its own row).
webscrape_forbes.ipynbtargets Forbes' undocumented internal API with a spoofed browser user-agent and hardcoded GDPR consent cookies, and the endpoint is still live. That is a republication question about the method, distinct from the data licence #35 already anticipates for the two Forbes CSVs, and it should not be folded into that row.One thing worth taking now
https://quantecon.github.io/data-lectures/lectures/<file>is live today — 200,access-control-allow-origin: *, gzipped, no redirect, correcttext/csv— and is immune to both rule 5 and rule 6 regardless of any futuresources/LFS decision. It does not wait on the DNS work in #37. Its one catch: the Pages deploy is gated behind the strict audit (deploy: needs: build), whereraw.githubusercontent.comupdates on push with no such dependency.- Audit: assert on lfs_media and on ref/path resolvability — strict exits 0 on a repoint that 404s every read #54 first. The strict audit exits 0 on a fold left entirely on
Closing — complete. The six
high_dim_datadatasets landed here as plain git in #62,sources/SCF_plus.dtain #63, all 28 consuming reads across four repos were repointed on 2026-08-11, and the tracker flipped in #69.high_dim_datais archived, not deleted, and still serves on both hosts. The full record is PLAN.md Phase 3 and repoint rule 6.
(Low Priority)
We should migrate https://github.com/QuantEcon/high_dim_data to be hosted in this repo in a folder
high_dimensionperhaps.