Skip to content

feat(benchmarks): render public EQ-Bench score as its own column (28 real scores currently invisible) - #23

Merged
TechNickAI merged 3 commits into
mainfrom
benchmark/eq-public-display
Jul 29, 2026
Merged

TechNickAI merged 3 commits into
mainfrom
benchmark/eq-public-display

Conversation

@TechNickAI

Copy link
Copy Markdown
Owner

Decision needed from Nick. Display-only, zero data changes. This has been waiting four steward passes; opening it as a PR so it is one click either way.

The problem

model-data.json holds 35 models. 28 carry a real, upstream-verified public EQ-Bench v3 score. The page renders only 12, because the EQ column reads a single field, v3_score (Nick's locally-run 22-trait metric). The other 23 cells show an em-dash while verified data sits behind them.

What this does

Splits one ambiguous EQ column into two clearly-labelled ones.

Column Reads Rubric Renders
EQ (local) v3_score 22 traits, our own run (~$6/model) 12 of 35
EQ (public) public_rubric_0_100 17 traits, eqbench.com 28 of 35

Both sortable, blanks sort last. Mobile cards get a matching tile.

What it deliberately does NOT do

  • No merging, averaging, or converting between the two metrics. Different rubrics, and the offset is not even consistent in sign (Grok 4.20: 68.55 local vs 55.8 public, −12.75). Each column shows only its own source.
  • No fallback. A model with no local score stays blank in the local column forever, never the public number wearing a local label.
  • No data changes. Diff is js/app.js + index.html (plus the proposal doc and two screenshots).

Both tooltips state in plain words that the rubrics differ and are not directly comparable; the public one names eqbench.com. Public trait values render as x/20 because upstream traits are 0-20, not 0-100.

Evidence verified

Headless Chromium against a local server, reading the rendered DOM:

  • Before: 35 rows, 23 EQ cells empty, 12 rendered.
  • After: 28 public + 12 local scores rendered.
  • Every rendered number cross-checked cell-by-cell against model-data.json: 0 mismatches across all 35 rows and both columns.
  • Sorting by EQ (public) verified descending, blanks last.

Re-verified 2026-07-28: branch merges cleanly into main, and all 28 public scores still match live upstream (28/28, 0 mismatches).

The ask

Merge, or say no and I will close it and stop carrying it forward. Either answer unblocks the benchmark half; silence is the only outcome that keeps 28 real scores invisible.

TechNickAI added a commit that referenced this pull request Jul 29, 2026
The pre-commit gate has been red on `main` for six consecutive runs since
2026-07-25. The failures were not caused by the PRs they appeared on: a clean
`main` reproduces them exactly. PR #23 sat blocked for five review passes on a
check that had nothing to do with it.

Root cause is two formatters disagreeing with two generators, permanently:

- `scripts/fetch-model.py:559` writes model-data.json via
  `json.dump(..., indent=2)`, expanding arrays one element per line. Prettier
  collapses short arrays. Each pipeline run re-breaks the hook.
- `generate_llms_txt()` re-emits upstream OpenRouter descriptions verbatim,
  which carry smart quotes and trailing whitespace. The trailing-whitespace and
  fix-smartquotes hooks rewrite them on every run.

These files are machine-written and machine-read. Their formatting is decided by
the generator, so a style hook there cannot catch a human mistake -- it only
reports the generator being itself.

Excludes the three generated data files from Prettier via a new .prettierignore,
and excludes llms.txt from the two text-fixup hooks.

Verified, including the counterfactuals -- an exclusion that silences a check is
worse than the red it replaces:

- `pre-commit run --all-files` now exits 0 and rewrites nothing.
- Deliberately malformatting js/app.js still FAILS Prettier, so real code is
  still gated.
- Appending garbage to model-data.json still FAILS the pipeline test, so the
  ignored files are still checked for content, just not for style.

The security hook (forbid-bidi-controls) is deliberately left applying to
everything, including llms.txt.

Co-authored-by: Nick Sullivan <nick@technick.ai>
@TechNickAI
TechNickAI merged commit 3b47403 into main Jul 29, 2026
2 of 3 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant