Repository navigation
feat(benchmarks): render public EQ-Bench score as its own column (28 real scores currently invisible) - #23
Merged
Merged
Conversation
TechNickAI
added a commit
that referenced
this pull request
Jul 29, 2026
The pre-commit gate has been red on `main` for six consecutive runs since 2026-07-25. The failures were not caused by the PRs they appeared on: a clean `main` reproduces them exactly. PR #23 sat blocked for five review passes on a check that had nothing to do with it. Root cause is two formatters disagreeing with two generators, permanently: - `scripts/fetch-model.py:559` writes model-data.json via `json.dump(..., indent=2)`, expanding arrays one element per line. Prettier collapses short arrays. Each pipeline run re-breaks the hook. - `generate_llms_txt()` re-emits upstream OpenRouter descriptions verbatim, which carry smart quotes and trailing whitespace. The trailing-whitespace and fix-smartquotes hooks rewrite them on every run. These files are machine-written and machine-read. Their formatting is decided by the generator, so a style hook there cannot catch a human mistake -- it only reports the generator being itself. Excludes the three generated data files from Prettier via a new .prettierignore, and excludes llms.txt from the two text-fixup hooks. Verified, including the counterfactuals -- an exclusion that silences a check is worse than the red it replaces: - `pre-commit run --all-files` now exits 0 and rewrites nothing. - Deliberately malformatting js/app.js still FAILS Prettier, so real code is still gated. - Appending garbage to model-data.json still FAILS the pipeline test, so the ignored files are still checked for content, just not for style. The security hook (forbid-bidi-controls) is deliberately left applying to everything, including llms.txt. Co-authored-by: Nick Sullivan <nick@technick.ai>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Decision needed from Nick. Display-only, zero data changes. This has been waiting four steward passes; opening it as a PR so it is one click either way.
The problem
model-data.jsonholds 35 models. 28 carry a real, upstream-verified public EQ-Bench v3 score. The page renders only 12, because the EQ column reads a single field,v3_score(Nick's locally-run 22-trait metric). The other 23 cells show an em-dash while verified data sits behind them.What this does
Splits one ambiguous EQ column into two clearly-labelled ones.
v3_scorepublic_rubric_0_100Both sortable, blanks sort last. Mobile cards get a matching tile.
What it deliberately does NOT do
js/app.js+index.html(plus the proposal doc and two screenshots).Both tooltips state in plain words that the rubrics differ and are not directly comparable; the public one names eqbench.com. Public trait values render as
x/20because upstream traits are 0-20, not 0-100.Evidence verified
Headless Chromium against a local server, reading the rendered DOM:
model-data.json: 0 mismatches across all 35 rows and both columns.Re-verified 2026-07-28: branch merges cleanly into
main, and all 28 public scores still match live upstream (28/28, 0 mismatches).The ask
Merge, or say no and I will close it and stop carrying it forward. Either answer unblocks the benchmark half; silence is the only outcome that keeps 28 real scores invisible.