Skip to content

Public EQ-Bench as its own metric (13→35 models) + Python 3.13 pin - #21

Merged
TechNickAI merged 6 commits into
mainfrom
benchmark/eq-public-discovery
Jul 27, 2026
Merged

TechNickAI merged 6 commits into
mainfrom
benchmark/eq-public-discovery

Conversation

@TechNickAI

Copy link
Copy Markdown
Owner

What this does

Two things: grows the benchmark dataset from 13 → 35 models with real, hand-verified EQ-Bench data, and pins the toolchain to Python 3.13 so the pipeline stops silently breaking on Apple's system 3.9.

The central finding

eqbench.com's public rubric_0_100 is not the same metric as our v3_score. v3_score comes from Nick's own local v3 runs over 22 traits; the public leaderboard scores 17 traits and lands 5-9 points higher on every model present in both:

Model local public Δ
openai/gpt-5.4 73.20 82.4 +9.20
anthropic/claude-sonnet-4.6 71.70 80.0 +8.30
anthropic/claude-opus-4.6 71.85 79.2 +7.35
google/gemini-3.1-pro 68.95 74.3 +5.35
google/gemma-4-31b-it 66.10 70.8 +4.70

So public values live in their own namespaced fields (public_rubric_0_100, public_elo_norm, public_traits_17, public_source_key) and never overwrite v3_score, v3_traits, or elo. Writing them into v3_score would have silently inflated the new rows against the existing ones.

Mapping is hand-verified, never fuzzy — substring matching actively mis-assigns (gpt-5.4-mini→gpt-5.4, glm-5-turbo→GLM-5, grok-4.20→grok-4). Unmapped models get an empty cell, which is the honest state.

Charter enforcement: qwen3.6-plus

qwen/qwen3.6-plus:free carried EQ data (elo 1417.4, v3_score 60.45, full 22-trait breakdown) copied verbatim from its predecessor Qwen3.5-397B-A17B. It was footnoted, not fixed — a note explaining inherited data doesn't make it that model's score, and the number still sorted and rendered as real. Removed; the model now shows an empty EQ cell. PinchBench (88.6/84.0) is retained, as that was run on Qwen3.6 directly.

This is why local v3_score count reads 12, not 13.

Python 3.13 pin

fetch-model.py:113 uses PEP 604 str | None without from __future__ import annotations. Under Apple's /usr/bin/python3 (3.9) the suite collects 1 test and fails; under 3.13 it collects 26 and passes. Pinned via .python-version + uv venv, AGENTS.md commands switched to uv run --no-project python, and both CI workflows moved 3.12 → 3.13 to match.

Verification

  • 26/26 unittest OK, 14/14 pre-commit hooks Passed, clean tree
  • Pre-existing models diffed field-by-field before/after: zero values changed, zero fields lost
  • All public scores re-read from upstream by an independent parser — exact match
  • No public rubric leaked into v3_score
  • Tests 11 → 26, including: public map is injective (no two rows may claim one upstream entry), and glm-5-turbo can never map to GLM-5

Deliberately not added (confirmed absent from EQ-Bench, so no real score exists): claude-opus-5, claude-opus-5-fast, grok-4.5, kimi-k3, gpt-5.6-*, gemini-3.6-flash.

Known issue, not addressed here

--refresh may destroy artificial_analysis scores when AA_API_KEY is absent. AGENTS.md claims keyless refresh is safe; that claim is unverified. The weekly CI job runs --refresh --no-aa, so this is worth settling before relying on it.

Nick Sullivan added 6 commits July 25, 2026 23:28
…wn metric

Adds the "ADD, don't just refresh" half of the update loop, plus a public
EQ-Bench v3 source, and grows the dataset from 13 to 26 models.

The central finding: eqbench.com's public `rubric_0_100` is NOT the same
metric as the site's `v3_score`. `v3_score` comes from Nick's own local v3
runs over 22 traits; the public leaderboard scores 17 traits and lands 5-9
points higher on every model present in both:

  openai/gpt-5.4               local 73.20  public 82.4  +9.20
  anthropic/claude-sonnet-4.6  local 71.70  public 80.0  +8.30
  anthropic/claude-opus-4.6    local 71.85  public 79.2  +7.35
  google/gemini-3.1-pro-...    local 68.95  public 74.3  +5.35
  google/gemma-4-31b-it        local 66.10  public 70.8  +4.70

The public Elo differs too (sonnet-4.6: stored 1876.8 vs public 1714.1).
So public values land in their own namespaced fields (public_rubric_0_100,
public_elo_norm, public_traits_17, public_source_key) and never overwrite
v3_score, v3_traits, or elo. Writing them into v3_score would have silently
inflated 13 new rows against the 13 existing ones.

Mapping is hand-verified, never fuzzy. Substring matching actively
mis-assigns: gpt-5.4-mini→gpt-5.4, glm-5-turbo→GLM-5, grok-4.20→grok-4.
Unmapped models get an empty cell, which is honest.

13 models added, all with real public scores and no local v3_score invented:
claude-opus-4.8 (83.3, top of the public leaderboard), claude-fable-5 (82.6),
gpt-5.5 (82.4), glm-5.2, deepseek-v4-pro/flash, kimi-k2.6/k2.5, glm-5.1,
glm-4.7, glm-4.7-flash, gemma-4-26b-a4b, qwen3.5-397b.

Verified: 13 pre-existing models diffed field-by-field before/after — zero
values changed, zero fields lost. Local v3_score count stays 13. Every
public score re-checked against the upstream leaderboard row.

Deliberately NOT added (confirmed absent from EQ-Bench, so no real score
exists): claude-opus-5, claude-opus-5-fast, grok-4.5, kimi-k3, gpt-5.6-*,
gemini-3.6-flash.

New flags: --discover (add untracked models with verified EQ entries),
--eq-public (fetch public leaderboard fields). Tests 11 → 19.
…edecessor

Charter non-negotiable: "never present a successor model's score as a
predecessor's." qwen/qwen3.6-plus:free carried EQ-Bench data (elo 1417.4,
v3_score 60.45, and a full 22-trait breakdown) copied verbatim from
Qwen3.5-397B-A17B. It was footnoted, not fixed — a note explaining inherited
data does not make it that model's score, and the number still sorted and
rendered as if it were real.

The previous commit added Qwen3.5-397B-A17B as its own tracked row, so the
borrowed copy is now both a violation and redundant. Removed elo, v3_score,
v3_traits, and the inline note. Qwen3.6 Plus now shows an empty EQ cell, which
is the honest state: it has not been benchmarked on EQ-Bench.

PinchBench (88.6/84.0) is retained — per the original note it was run on
Qwen3.6 Plus directly, so it is genuinely this model's own data.

Verified: field-by-field diff shows exactly 5 changes, all on qwen3.6-plus,
0 changes to the other 25 models.

Four regression tests added (19 -> 23) reading the real data file: no
inherited EQ fields on qwen3.6, predecessor tracked separately, PinchBench
survives, and no two rows may share a public_source_key (two rows citing one
EQ-Bench entry means one of them inherited it).

Found by the pass's independent review panel, which correctly flagged that
the steward had reported this as a "pre-existing integrity issue" instead of
enforcing the charter on it.
Extends EQBENCH_PUBLIC_MAP with 9 hand-verified mappings, each confirmed
both present on eqbench.com and live on OpenRouter (matched by name +
canonical_slug, never by substring). Dataset 26 -> 35 models.

Added: claude-opus-4.7 (82.6), gpt-5.2 (80.4), gpt-5.1 (78.1),
claude-opus-4.5 (76.8), glm-5 (75.0), kimi-k2 (75.0), sonnet-4.5 (72.2),
gpt-5.3-chat (70.7), hermes-4-405b (65.2). None carries a v3_score:
those are Nick's local runs and remain empty rather than invented.

Verified: field-by-field diff of the pre-existing 26 shows 0 changes and
0 fields lost; all 56 public values re-read from upstream by an
independent parser match exactly; no public rubric leaked into v3_score.

Resolves the glm-5-turbo provenance question: it is NOT a copy of GLM-5.
Only 4 of 11 rescaled traits agree within rounding and the stored elo
(1631.9) differs from GLM-5's public elo_norm (1526.0). GLM-5 is a
distinct OpenRouter model and now has its own row.

Tests 23 -> 26: public map is injective (no two models can claim one
upstream row), and glm-5-turbo can never be mapped to GLM-5.
Apple's python3 is 3.9 and chokes on str | None in fetch-model.py.
Pin 3.13 in .python-version, document the uv venv workflow, and
switch all AGENTS.md commands to uv run --no-project python.
Ignore .venv/ and __pycache__/.

@cursor cursor Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes using default effort and found 2 potential issues.

Fix All in Cursor

❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.

Reviewed by Cursor Bugbot for commit 68bc4e5. Configure here.

found.append(model_id)
return sorted(found)


Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Refresh drops public source flag

Medium Severity

merge_model restores sources.artificial_analysis and sources.eq_bench after OpenRouter refresh, but not sources.eq_bench_public, which apply_public_eq sets. Each --refresh (including weekly CI) replaces sources from transform_model, so eq_bench_public: true is removed from models even when benchmarks.eq_bench public fields are preserved.

Fix in Cursor Fix in Web

Reviewed by Cursor Bugbot for commit 68bc4e5. Configure here.

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in PR #22 (be8607a). Added eq_bench_public to the source-key preservation loop in merge_model.

print(
f" Public EQ applied to {updated} model(s); "
f"{missing} have no hand-verified public entry (left empty)."
)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Discover skips public EQ fetch

Medium Severity

--discover adds rows because they have hand-verified EQ-Bench map entries, but public scores are only fetched when --eq-public is also passed. Running --discover alone merges OpenRouter metadata without calling apply_public_eq, leaving new models without public_rubric_0_100 despite discovery being EQ-gated.

Fix in Cursor Fix in Web

Reviewed by Cursor Bugbot for commit 68bc4e5. Configure here.

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in PR #22 (be8607a). --discover now auto-calls apply_public_eq scoped to the discovered IDs — discovery is gated on EQBENCH_PUBLIC_MAP so the mapping is already verified.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 68bc4e5d57

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

"artificial_analysis": {},
"eq_bench": {
"public_source_key": "claude-fable-5",
"public_rubric_0_100": 82.6,

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Render the newly stored public EQ metric

For public-only rows such as this one, the site still sorts and renders exclusively from v3_score in js/app.js (getSortValue() at lines 133–134, eqHtml() at lines 201–203, and mobile cards at line 336). Consequently, 22 newly added models display an empty EQ cell and sort below locally tested models despite having public_rubric_0_100 values; the public rubric needs its own presentation and sorting path rather than remaining unused raw data.

AGENTS.md reference: model-benchmarks/AGENTS.md:L132-L136

Useful? React with 👍 / 👎.

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

👎 WONTFIX — the EQ column deliberately shows only v3_score (22-trait local run). Public data lives in separate public_* fields as raw provenance storage; the product keeps these two metrics distinct. Displaying public_rubric_0_100 in the same column would conflate two different benchmarks.

Comment on lines +784 to +786
if args.eq_public:
print("Fetching public EQ-Bench v3 leaderboard...")
updated, missing = apply_public_eq(data)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Populate public EQ data during discovery

When an operator follows the new help text and runs --discover by itself, the discovered IDs are fetched and persisted with the empty eq_bench block produced by transform_model(), because this guard prevents apply_public_eq() from running unless the separate --eq-public flag was also supplied. Models selected specifically for verified EQ evidence therefore get added without that evidence; discovery should apply the public fields automatically or explicitly require both flags.

AGENTS.md reference: model-benchmarks/AGENTS.md:L3-L5

Useful? React with 👍 / 👎.

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in PR #22 (be8607a). Matches the Cursor diagnosis — --discover now runs public EQ automatically for newly found models.

Comment on lines +452 to +455
block = eqbench_public.public_eq_block(m["id"], leaderboard, traits)
if block is None:
missing += 1
continue

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Clear public fields when an upstream entry disappears

If a mapped leaderboard row is removed, renamed, or becomes malformed, public_eq_block() returns None and this branch leaves every previously stored public_* field and the true source flag untouched. Thus a later --eq-public run reports the model as missing while continuing to publish its stale score; fields owned by this module should be cleared when no valid block is returned, and omitted cells should likewise remove their prior values.

Useful? React with 👍 / 👎.

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in PR #22 (be8607a). apply_public_eq now clears all PUBLIC_EQ_FIELDS and removes the source flag when public_eq_block() returns None.

f" + {m['id']} public rubric="
f"{block.get('public_rubric_0_100')} elo={block.get('public_elo_norm')}"
)
m.setdefault("sources", {})["eq_bench_public"] = True

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Preserve the public EQ provenance flag on refresh

After any ordinary --refresh—including the weekly workflow—the existing eq_bench object is retained, but merge_model() rebuilds sources and only preserves artificial_analysis and eq_bench at lines 531–536. The newly introduced eq_bench_public flag set here is therefore dropped while its public scores remain, leaving persisted data with missing provenance; include this source key in the merge preservation logic.

Useful? React with 👍 / 👎.

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in PR #22 (be8607a) alongside the matching Cursor flag — same one-line change, same root cause.

@TechNickAI
TechNickAI merged commit 68bc4e5 into main Jul 27, 2026
9 of 11 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant