Repository navigation
Public EQ-Bench as its own metric (13→35 models) + Python 3.13 pin - #21
Conversation
…wn metric Adds the "ADD, don't just refresh" half of the update loop, plus a public EQ-Bench v3 source, and grows the dataset from 13 to 26 models. The central finding: eqbench.com's public `rubric_0_100` is NOT the same metric as the site's `v3_score`. `v3_score` comes from Nick's own local v3 runs over 22 traits; the public leaderboard scores 17 traits and lands 5-9 points higher on every model present in both: openai/gpt-5.4 local 73.20 public 82.4 +9.20 anthropic/claude-sonnet-4.6 local 71.70 public 80.0 +8.30 anthropic/claude-opus-4.6 local 71.85 public 79.2 +7.35 google/gemini-3.1-pro-... local 68.95 public 74.3 +5.35 google/gemma-4-31b-it local 66.10 public 70.8 +4.70 The public Elo differs too (sonnet-4.6: stored 1876.8 vs public 1714.1). So public values land in their own namespaced fields (public_rubric_0_100, public_elo_norm, public_traits_17, public_source_key) and never overwrite v3_score, v3_traits, or elo. Writing them into v3_score would have silently inflated 13 new rows against the 13 existing ones. Mapping is hand-verified, never fuzzy. Substring matching actively mis-assigns: gpt-5.4-mini→gpt-5.4, glm-5-turbo→GLM-5, grok-4.20→grok-4. Unmapped models get an empty cell, which is honest. 13 models added, all with real public scores and no local v3_score invented: claude-opus-4.8 (83.3, top of the public leaderboard), claude-fable-5 (82.6), gpt-5.5 (82.4), glm-5.2, deepseek-v4-pro/flash, kimi-k2.6/k2.5, glm-5.1, glm-4.7, glm-4.7-flash, gemma-4-26b-a4b, qwen3.5-397b. Verified: 13 pre-existing models diffed field-by-field before/after — zero values changed, zero fields lost. Local v3_score count stays 13. Every public score re-checked against the upstream leaderboard row. Deliberately NOT added (confirmed absent from EQ-Bench, so no real score exists): claude-opus-5, claude-opus-5-fast, grok-4.5, kimi-k3, gpt-5.6-*, gemini-3.6-flash. New flags: --discover (add untracked models with verified EQ entries), --eq-public (fetch public leaderboard fields). Tests 11 → 19.
…edecessor Charter non-negotiable: "never present a successor model's score as a predecessor's." qwen/qwen3.6-plus:free carried EQ-Bench data (elo 1417.4, v3_score 60.45, and a full 22-trait breakdown) copied verbatim from Qwen3.5-397B-A17B. It was footnoted, not fixed — a note explaining inherited data does not make it that model's score, and the number still sorted and rendered as if it were real. The previous commit added Qwen3.5-397B-A17B as its own tracked row, so the borrowed copy is now both a violation and redundant. Removed elo, v3_score, v3_traits, and the inline note. Qwen3.6 Plus now shows an empty EQ cell, which is the honest state: it has not been benchmarked on EQ-Bench. PinchBench (88.6/84.0) is retained — per the original note it was run on Qwen3.6 Plus directly, so it is genuinely this model's own data. Verified: field-by-field diff shows exactly 5 changes, all on qwen3.6-plus, 0 changes to the other 25 models. Four regression tests added (19 -> 23) reading the real data file: no inherited EQ fields on qwen3.6, predecessor tracked separately, PinchBench survives, and no two rows may share a public_source_key (two rows citing one EQ-Bench entry means one of them inherited it). Found by the pass's independent review panel, which correctly flagged that the steward had reported this as a "pre-existing integrity issue" instead of enforcing the charter on it.
Extends EQBENCH_PUBLIC_MAP with 9 hand-verified mappings, each confirmed both present on eqbench.com and live on OpenRouter (matched by name + canonical_slug, never by substring). Dataset 26 -> 35 models. Added: claude-opus-4.7 (82.6), gpt-5.2 (80.4), gpt-5.1 (78.1), claude-opus-4.5 (76.8), glm-5 (75.0), kimi-k2 (75.0), sonnet-4.5 (72.2), gpt-5.3-chat (70.7), hermes-4-405b (65.2). None carries a v3_score: those are Nick's local runs and remain empty rather than invented. Verified: field-by-field diff of the pre-existing 26 shows 0 changes and 0 fields lost; all 56 public values re-read from upstream by an independent parser match exactly; no public rubric leaked into v3_score. Resolves the glm-5-turbo provenance question: it is NOT a copy of GLM-5. Only 4 of 11 rescaled traits agree within rounding and the stored elo (1631.9) differs from GLM-5's public elo_norm (1526.0). GLM-5 is a distinct OpenRouter model and now has its own row. Tests 23 -> 26: public map is injective (no two models can claim one upstream row), and glm-5-turbo can never be mapped to GLM-5.
Apple's python3 is 3.9 and chokes on str | None in fetch-model.py. Pin 3.13 in .python-version, document the uv venv workflow, and switch all AGENTS.md commands to uv run --no-project python. Ignore .venv/ and __pycache__/.
There was a problem hiding this comment.
Cursor Bugbot has reviewed your changes using default effort and found 2 potential issues.
❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.
Reviewed by Cursor Bugbot for commit 68bc4e5. Configure here.
| found.append(model_id) | ||
| return sorted(found) | ||
|
|
||
|
|
There was a problem hiding this comment.
Refresh drops public source flag
Medium Severity
merge_model restores sources.artificial_analysis and sources.eq_bench after OpenRouter refresh, but not sources.eq_bench_public, which apply_public_eq sets. Each --refresh (including weekly CI) replaces sources from transform_model, so eq_bench_public: true is removed from models even when benchmarks.eq_bench public fields are preserved.
Reviewed by Cursor Bugbot for commit 68bc4e5. Configure here.
| print( | ||
| f" Public EQ applied to {updated} model(s); " | ||
| f"{missing} have no hand-verified public entry (left empty)." | ||
| ) |
There was a problem hiding this comment.
Discover skips public EQ fetch
Medium Severity
--discover adds rows because they have hand-verified EQ-Bench map entries, but public scores are only fetched when --eq-public is also passed. Running --discover alone merges OpenRouter metadata without calling apply_public_eq, leaving new models without public_rubric_0_100 despite discovery being EQ-gated.
Reviewed by Cursor Bugbot for commit 68bc4e5. Configure here.
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 68bc4e5d57
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| "artificial_analysis": {}, | ||
| "eq_bench": { | ||
| "public_source_key": "claude-fable-5", | ||
| "public_rubric_0_100": 82.6, |
There was a problem hiding this comment.
Render the newly stored public EQ metric
For public-only rows such as this one, the site still sorts and renders exclusively from v3_score in js/app.js (getSortValue() at lines 133–134, eqHtml() at lines 201–203, and mobile cards at line 336). Consequently, 22 newly added models display an empty EQ cell and sort below locally tested models despite having public_rubric_0_100 values; the public rubric needs its own presentation and sorting path rather than remaining unused raw data.
AGENTS.md reference: model-benchmarks/AGENTS.md:L132-L136
Useful? React with 👍 / 👎.
There was a problem hiding this comment.
👎 WONTFIX — the EQ column deliberately shows only v3_score (22-trait local run). Public data lives in separate public_* fields as raw provenance storage; the product keeps these two metrics distinct. Displaying public_rubric_0_100 in the same column would conflate two different benchmarks.
| if args.eq_public: | ||
| print("Fetching public EQ-Bench v3 leaderboard...") | ||
| updated, missing = apply_public_eq(data) |
There was a problem hiding this comment.
Populate public EQ data during discovery
When an operator follows the new help text and runs --discover by itself, the discovered IDs are fetched and persisted with the empty eq_bench block produced by transform_model(), because this guard prevents apply_public_eq() from running unless the separate --eq-public flag was also supplied. Models selected specifically for verified EQ evidence therefore get added without that evidence; discovery should apply the public fields automatically or explicitly require both flags.
AGENTS.md reference: model-benchmarks/AGENTS.md:L3-L5
Useful? React with 👍 / 👎.
| block = eqbench_public.public_eq_block(m["id"], leaderboard, traits) | ||
| if block is None: | ||
| missing += 1 | ||
| continue |
There was a problem hiding this comment.
Clear public fields when an upstream entry disappears
If a mapped leaderboard row is removed, renamed, or becomes malformed, public_eq_block() returns None and this branch leaves every previously stored public_* field and the true source flag untouched. Thus a later --eq-public run reports the model as missing while continuing to publish its stale score; fields owned by this module should be cleared when no valid block is returned, and omitted cells should likewise remove their prior values.
Useful? React with 👍 / 👎.
| f" + {m['id']} public rubric=" | ||
| f"{block.get('public_rubric_0_100')} elo={block.get('public_elo_norm')}" | ||
| ) | ||
| m.setdefault("sources", {})["eq_bench_public"] = True |
There was a problem hiding this comment.
Preserve the public EQ provenance flag on refresh
After any ordinary --refresh—including the weekly workflow—the existing eq_bench object is retained, but merge_model() rebuilds sources and only preserves artificial_analysis and eq_bench at lines 531–536. The newly introduced eq_bench_public flag set here is therefore dropped while its public scores remain, leaving persisted data with missing provenance; include this source key in the merge preservation logic.
Useful? React with 👍 / 👎.


What this does
Two things: grows the benchmark dataset from 13 → 35 models with real, hand-verified EQ-Bench data, and pins the toolchain to Python 3.13 so the pipeline stops silently breaking on Apple's system 3.9.
The central finding
eqbench.com's public
rubric_0_100is not the same metric as ourv3_score.v3_scorecomes from Nick's own local v3 runs over 22 traits; the public leaderboard scores 17 traits and lands 5-9 points higher on every model present in both:So public values live in their own namespaced fields (
public_rubric_0_100,public_elo_norm,public_traits_17,public_source_key) and never overwritev3_score,v3_traits, orelo. Writing them intov3_scorewould have silently inflated the new rows against the existing ones.Mapping is hand-verified, never fuzzy — substring matching actively mis-assigns (
gpt-5.4-mini→gpt-5.4,glm-5-turbo→GLM-5,grok-4.20→grok-4). Unmapped models get an empty cell, which is the honest state.Charter enforcement: qwen3.6-plus
qwen/qwen3.6-plus:freecarried EQ data (elo 1417.4, v3_score 60.45, full 22-trait breakdown) copied verbatim from its predecessor Qwen3.5-397B-A17B. It was footnoted, not fixed — a note explaining inherited data doesn't make it that model's score, and the number still sorted and rendered as real. Removed; the model now shows an empty EQ cell. PinchBench (88.6/84.0) is retained, as that was run on Qwen3.6 directly.This is why local
v3_scorecount reads 12, not 13.Python 3.13 pin
fetch-model.py:113uses PEP 604str | Nonewithoutfrom __future__ import annotations. Under Apple's/usr/bin/python3(3.9) the suite collects 1 test and fails; under 3.13 it collects 26 and passes. Pinned via.python-version+uv venv, AGENTS.md commands switched touv run --no-project python, and both CI workflows moved 3.12 → 3.13 to match.Verification
v3_scoreglm-5-turbocan never map to GLM-5Deliberately not added (confirmed absent from EQ-Bench, so no real score exists): claude-opus-5, claude-opus-5-fast, grok-4.5, kimi-k3, gpt-5.6-*, gemini-3.6-flash.
Known issue, not addressed here
--refreshmay destroyartificial_analysisscores whenAA_API_KEYis absent. AGENTS.md claims keyless refresh is safe; that claim is unverified. The weekly CI job runs--refresh --no-aa, so this is worth settling before relying on it.