Repository navigation
Setup Git LFS for large file support #1
Description
Activity
Status — 2026-07-16
Still open and still needed, but the shape of it is now clearer, and there is a hard constraint to record before anyone acts.
Per-path, not blanket.
high_dim_datais the cautionary example: it carries*.csv filter=lfs, which tracks every CSV regardless of size.AGENTS.mdhere says this repo will not inherit that — LFS is opt-in, large binaries only. Nothing in the current holdings needs it: the largest file in the published tree ismpd2020.xlsxat 1.7 MB, comfortably plain-git. The first real candidates arrive with pilot P3 (QuantEcon/meta#338): the SCF pair fromhigh_dim_data(SCF_plus.dta~99 MB,SCF_plus_mini.csv~31 MB,SCF_plus_mini_no_weights.csv~72 MB).The constraint — enabling LFS on an existing path silently breaks its consumers. An LFS-tracked file served over
raw.githubusercontent.comreturns pointer text, not data. It does not 404; the consumer gets three lines ofversion https://git-lfs.github.com/spec/v1…and fails as a parse error somewhere confusing. Verified live today againstSCF_plus_mini.csvonhigh_dim_datamain:URL form LFS-tracked file plain-git file raw.githubusercontent.com/…❌ pointer text ✅ github.com/{org}/{repo}/raw/{ref}/…✅ ✅ media.githubusercontent.com/media/…✅ ❌ 404 So: do not LFS-track an existing file until every consumer uses a form that survives it. The
/raw/form is the only one that works either way, which is why it is the interim canonical form (QuantEcon/QuantEcon.manual#108).One more, for whoever wires up Pages (PLAN Phase 4): the deploy workflow must check out with
lfs: true, or it publishes pointer files.Right order: adopt the
/raw/URL form for consumers → then enable per-path LFS → then Pages withlfs: true. Doing it in any other order breaks something quietly.Phase 3's measurements have largely dissolved this issue
The sizes in the 2026-07-16 comment were approximations, and measured exactly they change the answer. Both SCF minis land under the 100 MiB (104,857,600 B) hard blob limit —
SCF_plus_mini.csvis 32,853,734 B andSCF_plus_mini_no_weights.csvis 75,902,999 B — so both go in as plain git and the published tree stays 100% plain git. That means no consumer of a published dataset ever meets the raw-vs-media trap this issue exists to guard against, and every read keeps raw's gzip:SCF_plus_mini.csvtransfers 6,038,868 B gzipped fromraw.githubusercontent.comagainst 32,853,734 B frommedia(which does not gzip) — a 5.44x wire saving per read that LFS-tracking would forfeit.What is left is exactly one file.
SCF_plus.dtais 103,934,093 B — 923,507 B, or 0.88%, under the hard limit. It cannot ever be plain git, so it needs per-path LFS and must stay LFS-tracked permanently; there is no future version of the repo where it graduates off. Scope the.gitattributesrule tosources/and to that path, never a blanket*.dta/*.csv— high_dim_data's blanket rule is precisely how a 31 MB CSV that never needed LFS ended up serving pointer text to live lectures.Two operational facts that are not recorded anywhere in this thread.
Both workflows that check this repo out say
lfs: truetoday —consumed-file-check.yml:22andaudit-dashboard.yml:43— and both must flip tofalsebefore the LFS object lands. Neither needs it: the Pages tree is assembled fromlectures/+site/+audit.json, sosources/is never published, and the consumed-file hash only ever coverslectures/. Left attrue, every run pulls ~104 MB against the org-wide LFS bandwidth quota, which is shared with high_dim_data — and if that quota trips, LFS serving returns 403 and takes out the live lecture reads still pointed at high_dim_data. CI would become the thing that breaks published lectures. Flipping tofalsealso strengthensconsumed-file-check.yml's stated intent: if a consumed file is ever LFS-tracked by accident, hashing the pointer fails loudly, which is exactly the guard that comment wants, at zero bandwidth.LFS-on-raw returns HTTP 200 with pointer text, so any status-code check is a false green. It is also worse than "a parse error somewhere confusing":
pd.read_csvon pointer text raises nothing at all — it returns a 2x1 frame whose sole column name isversion https://git-lfs.github.com/spec/v1.pd.read_statadoes raise. So a pointer regression on a CSV is silent all the way into a published figure, and any guard has to be a byte/content check, neverraise_for_status. The strict audit does not currently assert on this —lfs_mediais computed and asserted nowhere, tracked in #54.Recommendation. This issue is now one
.gitattributesline scoped tosources/, and nothing else. Either narrow the title to that scope, or close it when the line lands with the two workflow flips, and let #54 own the pointer-regression guard.- added a commit that references this issue
on Aug 9, 2026 LFS is set up, in use, and now CI-enforced — recommending this can close
Done, though not in the shape this issue implies. #58 is the standing record of the reasoning and explicitly supersedes this thread; what follows is just the confirmation that the work exists.
.gitattributesscopes LFS tosources/**only, never a blanket*.csvrule (LFS: scope it to sources/**, and stop fetching it in CI #57).sources/README.mdis excluded so the audit trail stays readable text.- The published tree stays 100% plain git. That is the substantive answer to the question here: LFS is deliberately not enabled for
lectures/, because an LFS-tracked path returns HTTP 200 with ~130 bytes of pointer text fromraw.githubusercontent.com— a reader gets a parse error rather than a 404, andpd.read_csvraises nothing at all. Keeping the served tree plain-git removes that hazard rather than managing it. - First real use landed 2026-08-10:
sources/SCF_plus.dta, 103,934,093 B, in Land SCF_plus.dta in sources/, and gate sources/ on its recorded hashes (PR B2) #63. The index holds a 134-byte pointer and the object serves real bytes frommain. - Both workflow checkouts are
lfs: false, which is the assertion that nothing published is an LFS object rather than a bandwidth saving. - CI now enforces it:
check_consumed_files.pyfails if asources/file is not captured by the LFS rule, or does not hash to thesha256recorded insources/README.md. It reads the pointer'soid, so it verifies ~100 MiB at zero LFS bandwidth.
One measurement worth leaving here, since this issue is where someone will look for it: the org's net LFS charge for all of 2026 is $0.04 on $1.04 gross, and only four repos org-wide use LFS. The mechanism still deserves respect — anonymous downloads bill the repository owner with no open-source exemption, forks count against the parent, and a $0 budget blocks downloads rather than billing them — but it is not currently a constraint.
Close whenever you like; #58 carries everything durable.
Closing, as the 2026-08-10 comment recommended. LFS is set up, scoped to
sources/**only, and CI-enforced; the published tree is deliberately plain git. The reasoning and the measurements live in #58 and PLAN.md Phase 3, which supersede this thread.
@mmcky to setup
git-lfsfor large file support on this repository