Repository navigation
fix(benchmarks): plug three eq_bench_public data-integrity gaps - #22
TechNickAI wants to merge 1 commit into
Conversation
- merge_model: preserve sources.eq_bench_public alongside artificial_analysis and eq_bench so weekly --refresh doesn't silently drop the provenance flag while keeping the public scores - apply_public_eq: clear stale PUBLIC_EQ_FIELDS + source flag when public_eq_block() returns None, preventing outdated data from being published after a mapped leaderboard row disappears or is renamed - --discover: auto-apply public EQ scores for newly discovered models; discovery is already gated on EQBENCH_PUBLIC_MAP so every found model has a verified mapping — requiring a separate --eq-public pass was an unnecessary footgun Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: be8607a5cd
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| # --discover gates on EQ-Bench map entries, so every discovered model | ||
| # already has a verified mapping — fetch their public scores automatically. | ||
| print(f"Fetching public EQ-Bench scores for {len(discovered)} discovered model(s)...") | ||
| updated, missing = apply_public_eq(data, only_ids=set(discovered)) |
There was a problem hiding this comment.
Include dry-run discoveries before applying public EQ
When --discover is combined with --dry-run, the fetch loop only prints each transformed model and never merges it into data, so this call filters over a dataset containing none of the IDs in discovered. It therefore performs both EQ-Bench requests but always reports zero updates and shows no public fields for the new models, preventing the dry run from validating the enrichment this branch adds. Merge the temporary models into an in-memory copy for this step or skip/clearly report the enrichment during dry runs.
Useful? React with 👍 / 👎.
|
Review from the benchmark steward pass (2026-07-28). One blocking hazard, one now-redundant hunk, one good catch worth keeping.
|
Summary
Follow-up to #21. Three data-integrity / workflow defects flagged by Cursor and Codex bots:
merge_modeldropssources.eq_bench_publicon refresh — weekly--refreshrebuildssourcesfromtransform_model()and preservedartificial_analysis+eq_benchbut noteq_bench_public. Public scores survived; their provenance flag silently vanished. Addedeq_bench_publicto the preservation loop.apply_public_eqleaves stale fields when upstream entry disappears — if a mapped leaderboard row is removed or renamed,public_eq_block()returnsNoneand the oldpublic_*fields + source flag were left untouched, publishing outdated data. Now clears allPUBLIC_EQ_FIELDSand removes the source flag when no valid block is returned.--discoverskipped public EQ fetch — discovery is already gated onEQBENCH_PUBLIC_MAP(hand-verified), so every discovered model has a known mapping. Running--discoverwithout--eq-publicleft new rows with empty public scores. Now auto-appliesapply_public_eqscoped to the discovered IDs.Declined: Codex comment 3653485341 ("render public EQ metric in the UI") — the product deliberately keeps public 17-trait data separate from the displayed local 22-trait v3 metric. The column shows
v3_scoreonly; this is by design.Test plan
uv run --no-project python -m unittest discover -s model-benchmarks/tests)--refreshrun on a model withsources.eq_bench_public: true— verify flag survives--discoveron a clean dataset — verify public EQ scores are populated for discovered models without needing--eq-public🤖 Generated with Claude Code