Skip to content

Latest commit

 

History

2 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

unbenchmaxxed

A contamination-proof LLM benchmark that can't be benchmaxxed. Every run, an LLM "teacher" invents brand-new coding tasks, verifies them by actually executing a reference solution, and then grades any model on OpenRouter against hidden tests. There is no fixed dataset to leak, memorize or overfit.

License: MIT Node >= 22 TypeScript


The problem: benchmaxxing

Public LLM benchmarks (HumanEval, MBPP, LiveCodeBench, SWE-bench…) are static. Once published, their problems and solutions end up in training data, and labs tune their models to top the leaderboard. This is benchmaxxing: scores go up while real-world ability barely moves. Data contamination and benchmark overfitting make it hard to tell whether a model reasons or just remembers.

unbenchmaxxed takes the opposite approach:

  • No static dataset. Tasks are generated on the fly from random recipes, so each suite is new.
  • Nothing to leak. Generated tasks stay in .private/ (git-ignored). Only aggregate scores are published.
  • Verifiable by construction. Expected outputs come from running code in a sandbox, not from an LLM's opinion.
  • Any model, one key. Evaluate any model available on OpenRouter, including reasoning models at different effort levels.

Quickstart

git clone https://github.com/CodeSquar/unbenchmaxxed.git
cd unbenchmaxxed
npm install
cp .env.example .env        # add your OPENROUTER_API_KEY
npm test                    # offline self-test (no API calls)
npm run bench -- run --models openai/gpt-5.6-terra,deepseek/deepseek-v4-flash@high --n 12

You'll see a live progress board while tasks are generated, followed by per-model results and a Markdown report in reports/.

How it works

flowchart LR
  A[Random recipe<br/>domain × concept × twist × difficulty] --> B[Teacher LLM]
  B --> C[Spec + reference solution<br/>+ test generator]
  C --> D[Sandbox: run generator<br/>→ test inputs]
  D --> E[Sandbox: run reference<br/>→ expected outputs]
  E --> F{Validation<br/>determinism · variety<br/>reviewer LLM}
  F -- rejected --> B
  F -- ok --> G[Evaluated models<br/>see spec + 3 examples]
  G --> H[Graded on hidden tests]
Loading
  1. Random recipe. An invented domain × algorithmic concept × twist × output type × difficulty (seeds.ts). That makes tens of thousands of combinations before the teacher adds its own variation. No task is written by hand.

  2. Teacher. The teacher LLM writes a specification, a reference solution, and a generateTests(rng) function. It never types test data by hand. A seeded RNG keeps the tests reproducible and allows large stress inputs.

  3. Asymmetric verification. Expected outputs come from executing the reference solution in a sandbox. The teacher only has to write code that works; it doesn't need to be smarter than the models it evaluates.

  4. Validation. Each task must pass these checks:

    • enough valid tests;
    • a deterministic reference;
    • varied outputs, so a constant answer can't pass;
    • an ambiguity audit by a reviewer from a different model family.

    Tasks that fail are discarded and regenerated.

  5. Evaluation. Each model sees the spec plus 3 public examples and is graded on the hidden tests. The time limit is 10× the reference runtime, so inefficient solutions fail.

  6. Report. The report includes:

    • mean score with a 95% bootstrap confidence interval;
    • fully-solved rate;
    • breakdown by difficulty;
    • latency and output tokens.

    Models that also acted as the teacher are flagged ⚠️ (authorship bias).

Usage

Command What it does
npm run bench -- run --models a,b --n 12 Generate a fresh suite and evaluate the models
npm run bench -- generate --n 30 Only generate a suite, to evaluate several models on the same tasks later
npm run bench -- eval --suite <id> --models a,b Evaluate a saved suite
npm run bench -- categories List task categories

Reasoning effort per model

Append @effort to any evaluated model to set its reasoning effort. The same model can be listed several times to compare effort levels side by side:

npm run bench -- eval --suite <id> --models deepseek/deepseek-v4-flash@low,deepseek/deepseek-v4-flash@high,moonshotai/kimi-k3

Valid values are none | minimal | low | medium | high | xhigh. If a model doesn't support reasoning (checked against the OpenRouter catalog), the effort is ignored with a warning.

Options

Flag Default Description
--n 12 Number of tasks
--difficulty mixed easy, medium, hard or mixed
--teacher $TEACHER_MODEL Model that generates tasks
--reviewer $REVIEWER_MODEL Model that audits tasks (--no-review to skip)
--teacher-effort / --reviewer-effort $TEACHER_EFFORT / $REVIEWER_EFFORT Reasoning effort for teacher/reviewer
--concurrency 4 Parallel API calls
--seed random Recipe seed (reproducible task recipes)
--timeout 300 Seconds per model answer

Configuration (.env)

Variable Description
OPENROUTER_API_KEY Your OpenRouter key
TEACHER_MODEL Default teacher model
REVIEWER_MODEL Default reviewer, ideally from a different family than the teacher
TEACHER_EFFORT, REVIEWER_EFFORT Optional reasoning effort

Outputs

  • reports/<suite>.md contains aggregates only and is safe to publish.
  • .private/ holds tasks, model answers and grading details. Never publish it: leaked tasks become training data.

Adding a category

The framework is category-agnostic. To add one:

  1. Create src/categories/<id>/index.ts implementing Category (src/core/types.ts):
    • generate(ctx) creates a task, keeping private grading data in hidden;
    • validate(task, ctx) rejects broken, trivial or ambiguous tasks;
    • grade(task, response) returns { score: 0..1, passed }.
  2. Register it in src/categories/index.ts.

Prefer tasks that are verifiable by construction: fictional worlds with planted facts, constraints checkable by code, simulated environments. Use an LLM judge only for what code can't verify, and then use pairwise comparison with judges from several model families.

FAQ

What is benchmaxxing? Optimizing a model, intentionally or through contaminated training data, to score high on a specific public benchmark without a matching gain in general ability. Static benchmarks make it inevitable over time.

Can't labs just train on this repo? They can train on the generator, but not on the tasks: every suite is new and private. Learning to solve freshly invented, verified problems in general is the skill being measured.

Doesn't the teacher need to be smarter than the evaluated models? No. Expected outputs come from executing the teacher's reference solution, not from the teacher's claims. A mid-tier teacher can produce hard tasks as long as its reference code runs correctly. Tasks where the spec and reference disagree are caught by the reviewer and discarded.

Why JavaScript? It runs in-process with Node's permission model and needs no extra toolchain. Other languages can be added as new categories.

Which teacher should I use? Strong instruction-followers produce the most valid tasks per dollar. Weaker or heavily quantized models work but get more tasks rejected. Avoid evaluating the teacher itself, or read its score with the ⚠️ in mind.

Known limitations

  • Sandbox. Node's Permission Model blocks disk writes, child processes and workers, but not the network. For public or large-scale use, run grading inside a network-less container.
  • Teacher bias. Tasks reflect the teacher's style. A planned mitigation is rotating several teachers and averaging.
  • Sample size. With --n 12 confidence intervals are wide. Use --n 50 or more to compare close models.
  • Language. Tasks are in JavaScript, which may slightly penalize models that are much stronger in Python.

Roadmap

  • Benchmaxxing index: isomorphic variants of HumanEval/MBPP vs. the originals
  • IRT-based difficulty calibration and item filtering
  • Tournament mode: every model also acts as teacher
  • Seeds from recent repos/issues (after model knowledge cutoffs)
  • More categories: reasoning over fictional worlds, constrained writing, tool use

Contributing

Issues and PRs are welcome, especially new categories. Run npm run typecheck && npm test before submitting.

License

MIT

About

Contamination-proof LLM benchmark: fresh, verified coding tasks every run, impossible to benchmaxx

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages