A contamination-proof LLM benchmark that can't be benchmaxxed. Every run, an LLM "teacher" invents brand-new coding tasks, verifies them by actually executing a reference solution, and then grades any model on OpenRouter against hidden tests. There is no fixed dataset to leak, memorize or overfit.
Public LLM benchmarks (HumanEval, MBPP, LiveCodeBench, SWE-bench…) are static. Once published, their problems and solutions end up in training data, and labs tune their models to top the leaderboard. This is benchmaxxing: scores go up while real-world ability barely moves. Data contamination and benchmark overfitting make it hard to tell whether a model reasons or just remembers.
unbenchmaxxed takes the opposite approach:
- No static dataset. Tasks are generated on the fly from random recipes, so each suite is new.
- Nothing to leak. Generated tasks stay in
.private/(git-ignored). Only aggregate scores are published. - Verifiable by construction. Expected outputs come from running code in a sandbox, not from an LLM's opinion.
- Any model, one key. Evaluate any model available on OpenRouter, including reasoning models at different effort levels.
git clone https://github.com/CodeSquar/unbenchmaxxed.git
cd unbenchmaxxed
npm install
cp .env.example .env # add your OPENROUTER_API_KEY
npm test # offline self-test (no API calls)
npm run bench -- run --models openai/gpt-5.6-terra,deepseek/deepseek-v4-flash@high --n 12You'll see a live progress board while tasks are generated, followed by per-model results and a Markdown report in reports/.
flowchart LR
A[Random recipe<br/>domain × concept × twist × difficulty] --> B[Teacher LLM]
B --> C[Spec + reference solution<br/>+ test generator]
C --> D[Sandbox: run generator<br/>→ test inputs]
D --> E[Sandbox: run reference<br/>→ expected outputs]
E --> F{Validation<br/>determinism · variety<br/>reviewer LLM}
F -- rejected --> B
F -- ok --> G[Evaluated models<br/>see spec + 3 examples]
G --> H[Graded on hidden tests]
-
Random recipe. An invented domain × algorithmic concept × twist × output type × difficulty (
seeds.ts). That makes tens of thousands of combinations before the teacher adds its own variation. No task is written by hand. -
Teacher. The teacher LLM writes a specification, a reference solution, and a
generateTests(rng)function. It never types test data by hand. A seeded RNG keeps the tests reproducible and allows large stress inputs. -
Asymmetric verification. Expected outputs come from executing the reference solution in a sandbox. The teacher only has to write code that works; it doesn't need to be smarter than the models it evaluates.
-
Validation. Each task must pass these checks:
- enough valid tests;
- a deterministic reference;
- varied outputs, so a constant answer can't pass;
- an ambiguity audit by a reviewer from a different model family.
Tasks that fail are discarded and regenerated.
-
Evaluation. Each model sees the spec plus 3 public examples and is graded on the hidden tests. The time limit is 10× the reference runtime, so inefficient solutions fail.
-
Report. The report includes:
- mean score with a 95% bootstrap confidence interval;
- fully-solved rate;
- breakdown by difficulty;
- latency and output tokens.
Models that also acted as the teacher are flagged
⚠️ (authorship bias).
| Command | What it does |
|---|---|
npm run bench -- run --models a,b --n 12 |
Generate a fresh suite and evaluate the models |
npm run bench -- generate --n 30 |
Only generate a suite, to evaluate several models on the same tasks later |
npm run bench -- eval --suite <id> --models a,b |
Evaluate a saved suite |
npm run bench -- categories |
List task categories |
Append @effort to any evaluated model to set its reasoning effort. The same model can be listed several times to compare effort levels side by side:
npm run bench -- eval --suite <id> --models deepseek/deepseek-v4-flash@low,deepseek/deepseek-v4-flash@high,moonshotai/kimi-k3Valid values are none | minimal | low | medium | high | xhigh. If a model doesn't support reasoning (checked against the OpenRouter catalog), the effort is ignored with a warning.
| Flag | Default | Description |
|---|---|---|
--n |
12 |
Number of tasks |
--difficulty |
mixed |
easy, medium, hard or mixed |
--teacher |
$TEACHER_MODEL |
Model that generates tasks |
--reviewer |
$REVIEWER_MODEL |
Model that audits tasks (--no-review to skip) |
--teacher-effort / --reviewer-effort |
$TEACHER_EFFORT / $REVIEWER_EFFORT |
Reasoning effort for teacher/reviewer |
--concurrency |
4 |
Parallel API calls |
--seed |
random | Recipe seed (reproducible task recipes) |
--timeout |
300 |
Seconds per model answer |
| Variable | Description |
|---|---|
OPENROUTER_API_KEY |
Your OpenRouter key |
TEACHER_MODEL |
Default teacher model |
REVIEWER_MODEL |
Default reviewer, ideally from a different family than the teacher |
TEACHER_EFFORT, REVIEWER_EFFORT |
Optional reasoning effort |
reports/<suite>.mdcontains aggregates only and is safe to publish..private/holds tasks, model answers and grading details. Never publish it: leaked tasks become training data.
The framework is category-agnostic. To add one:
- Create
src/categories/<id>/index.tsimplementingCategory(src/core/types.ts):generate(ctx)creates a task, keeping private grading data inhidden;validate(task, ctx)rejects broken, trivial or ambiguous tasks;grade(task, response)returns{ score: 0..1, passed }.
- Register it in
src/categories/index.ts.
Prefer tasks that are verifiable by construction: fictional worlds with planted facts, constraints checkable by code, simulated environments. Use an LLM judge only for what code can't verify, and then use pairwise comparison with judges from several model families.
What is benchmaxxing? Optimizing a model, intentionally or through contaminated training data, to score high on a specific public benchmark without a matching gain in general ability. Static benchmarks make it inevitable over time.
Can't labs just train on this repo? They can train on the generator, but not on the tasks: every suite is new and private. Learning to solve freshly invented, verified problems in general is the skill being measured.
Doesn't the teacher need to be smarter than the evaluated models? No. Expected outputs come from executing the teacher's reference solution, not from the teacher's claims. A mid-tier teacher can produce hard tasks as long as its reference code runs correctly. Tasks where the spec and reference disagree are caught by the reviewer and discarded.
Why JavaScript? It runs in-process with Node's permission model and needs no extra toolchain. Other languages can be added as new categories.
Which teacher should I use?
Strong instruction-followers produce the most valid tasks per dollar. Weaker or heavily quantized models work but get more tasks rejected. Avoid evaluating the teacher itself, or read its score with the
- Sandbox. Node's Permission Model blocks disk writes, child processes and workers, but not the network. For public or large-scale use, run grading inside a network-less container.
- Teacher bias. Tasks reflect the teacher's style. A planned mitigation is rotating several teachers and averaging.
- Sample size. With
--n 12confidence intervals are wide. Use--n 50or more to compare close models. - Language. Tasks are in JavaScript, which may slightly penalize models that are much stronger in Python.
- Benchmaxxing index: isomorphic variants of HumanEval/MBPP vs. the originals
- IRT-based difficulty calibration and item filtering
- Tournament mode: every model also acts as teacher
- Seeds from recent repos/issues (after model knowledge cutoffs)
- More categories: reasoning over fictional worlds, constrained writing, tool use
Issues and PRs are welcome, especially new categories. Run npm run typecheck && npm test before submitting.