Skip to content

Setup Git LFS for large file support #1

Description

@mmcky

@mmcky to setup git-lfs for large file support on this repository

Activity

  1. mmcky commented on Jul 16, 2026

    @mmcky
    ContributorAuthor

    Status — 2026-07-16

    Still open and still needed, but the shape of it is now clearer, and there is a hard constraint to record before anyone acts.

    Per-path, not blanket. high_dim_data is the cautionary example: it carries *.csv filter=lfs, which tracks every CSV regardless of size. AGENTS.md here says this repo will not inherit that — LFS is opt-in, large binaries only. Nothing in the current holdings needs it: the largest file in the published tree is mpd2020.xlsx at 1.7 MB, comfortably plain-git. The first real candidates arrive with pilot P3 (QuantEcon/meta#338): the SCF pair from high_dim_data (SCF_plus.dta ~99 MB, SCF_plus_mini.csv ~31 MB, SCF_plus_mini_no_weights.csv ~72 MB).

    The constraint — enabling LFS on an existing path silently breaks its consumers. An LFS-tracked file served over raw.githubusercontent.com returns pointer text, not data. It does not 404; the consumer gets three lines of version https://git-lfs.github.com/spec/v1… and fails as a parse error somewhere confusing. Verified live today against SCF_plus_mini.csv on high_dim_data main:

    URL form LFS-tracked file plain-git file
    raw.githubusercontent.com/… ❌ pointer text ✅
    github.com/{org}/{repo}/raw/{ref}/… ✅ ✅
    media.githubusercontent.com/media/… ✅ ❌ 404

    So: do not LFS-track an existing file until every consumer uses a form that survives it. The /raw/ form is the only one that works either way, which is why it is the interim canonical form (QuantEcon/QuantEcon.manual#108).

    One more, for whoever wires up Pages (PLAN Phase 4): the deploy workflow must check out with lfs: true, or it publishes pointer files.

    Right order: adopt the /raw/ URL form for consumers → then enable per-path LFS → then Pages with lfs: true. Doing it in any other order breaks something quietly.

  2. mmcky commented on Aug 7, 2026

    @mmcky
    ContributorAuthor

    Phase 3's measurements have largely dissolved this issue

    The sizes in the 2026-07-16 comment were approximations, and measured exactly they change the answer. Both SCF minis land under the 100 MiB (104,857,600 B) hard blob limit — SCF_plus_mini.csv is 32,853,734 B and SCF_plus_mini_no_weights.csv is 75,902,999 B — so both go in as plain git and the published tree stays 100% plain git. That means no consumer of a published dataset ever meets the raw-vs-media trap this issue exists to guard against, and every read keeps raw's gzip: SCF_plus_mini.csv transfers 6,038,868 B gzipped from raw.githubusercontent.com against 32,853,734 B from media (which does not gzip) — a 5.44x wire saving per read that LFS-tracking would forfeit.

    What is left is exactly one file. SCF_plus.dta is 103,934,093 B — 923,507 B, or 0.88%, under the hard limit. It cannot ever be plain git, so it needs per-path LFS and must stay LFS-tracked permanently; there is no future version of the repo where it graduates off. Scope the .gitattributes rule to sources/ and to that path, never a blanket *.dta / *.csv — high_dim_data's blanket rule is precisely how a 31 MB CSV that never needed LFS ended up serving pointer text to live lectures.

    Two operational facts that are not recorded anywhere in this thread.

    Both workflows that check this repo out say lfs: true today — consumed-file-check.yml:22 and audit-dashboard.yml:43 — and both must flip to false before the LFS object lands. Neither needs it: the Pages tree is assembled from lectures/ + site/ + audit.json, so sources/ is never published, and the consumed-file hash only ever covers lectures/. Left at true, every run pulls ~104 MB against the org-wide LFS bandwidth quota, which is shared with high_dim_data — and if that quota trips, LFS serving returns 403 and takes out the live lecture reads still pointed at high_dim_data. CI would become the thing that breaks published lectures. Flipping to false also strengthens consumed-file-check.yml's stated intent: if a consumed file is ever LFS-tracked by accident, hashing the pointer fails loudly, which is exactly the guard that comment wants, at zero bandwidth.

    LFS-on-raw returns HTTP 200 with pointer text, so any status-code check is a false green. It is also worse than "a parse error somewhere confusing": pd.read_csv on pointer text raises nothing at all — it returns a 2x1 frame whose sole column name is version https://git-lfs.github.com/spec/v1. pd.read_stata does raise. So a pointer regression on a CSV is silent all the way into a published figure, and any guard has to be a byte/content check, never raise_for_status. The strict audit does not currently assert on this — lfs_media is computed and asserted nowhere, tracked in #54.

    Recommendation. This issue is now one .gitattributes line scoped to sources/, and nothing else. Either narrow the title to that scope, or close it when the line lands with the two workflow flips, and let #54 own the pointer-regression guard.

  3. mmcky commented on Aug 10, 2026

    @mmcky
    ContributorAuthor

    LFS is set up, in use, and now CI-enforced — recommending this can close

    Done, though not in the shape this issue implies. #58 is the standing record of the reasoning and explicitly supersedes this thread; what follows is just the confirmation that the work exists.

    • .gitattributes scopes LFS to sources/** only, never a blanket *.csv rule (LFS: scope it to sources/**, and stop fetching it in CI #57). sources/README.md is excluded so the audit trail stays readable text.
    • The published tree stays 100% plain git. That is the substantive answer to the question here: LFS is deliberately not enabled for lectures/, because an LFS-tracked path returns HTTP 200 with ~130 bytes of pointer text from raw.githubusercontent.com — a reader gets a parse error rather than a 404, and pd.read_csv raises nothing at all. Keeping the served tree plain-git removes that hazard rather than managing it.
    • First real use landed 2026-08-10: sources/SCF_plus.dta, 103,934,093 B, in Land SCF_plus.dta in sources/, and gate sources/ on its recorded hashes (PR B2) #63. The index holds a 134-byte pointer and the object serves real bytes from main.
    • Both workflow checkouts are lfs: false, which is the assertion that nothing published is an LFS object rather than a bandwidth saving.
    • CI now enforces it: check_consumed_files.py fails if a sources/ file is not captured by the LFS rule, or does not hash to the sha256 recorded in sources/README.md. It reads the pointer's oid, so it verifies ~100 MiB at zero LFS bandwidth.

    One measurement worth leaving here, since this issue is where someone will look for it: the org's net LFS charge for all of 2026 is $0.04 on $1.04 gross, and only four repos org-wide use LFS. The mechanism still deserves respect — anonymous downloads bill the repository owner with no open-source exemption, forks count against the parent, and a $0 budget blocks downloads rather than billing them — but it is not currently a constraint.

    Close whenever you like; #58 carries everything durable.

  4. mmcky commented on Aug 31, 2026

    @mmcky
    ContributorAuthor

    Closing, as the 2026-08-10 comment recommended. LFS is set up, scoped to sources/** only, and CI-enforced; the published tree is deliberately plain git. The reasoning and the measurements live in #58 and PLAN.md Phase 3, which supersede this thread.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions