Public periodic AMBER benchmark results of models served by CommandCode (api.commandcode.ai). Cases stay private; results are public. 中文说明:README.md
- One
results/YYYY-Www.mdper issue: same cases, same harness, full library per model; same model name across vendors side by side. - Each issue pins: library size and hashes, per-case defect-hunt score and pass/fail, terminal states, token usage (when the lane reports it) and latency, environment fingerprint, and a qualitative verdict written under evidence discipline.
- Cases, oracles, transcripts and intermediates are never published.
- Sister repos: amber-deepseek (official DeepSeek lane), amber-opencode (OpenCode Go lane), amber-gpt, amber-crof, amber-ollama, amber-devin, amber-workbuddy (WorkBuddy ACP lane). This repo's comparison axis is same-name cross-vendor duels — the same model name on CommandCode / OpenCode Go / the official DeepSeek API can be a different endpoint, and every cross-repo citation carries an explicit date and band declaration.
- Publish only: scores and aggregates, token usage (when reported), speed, qualitative verdicts.
- Never publish: case content, oracles/graders, transcripts, candidate workspaces, anything that could reconstruct a case.
- Every issue pins: model ID, effort band, date (UTC), harness version, per-case bundle hash — verifiable against the public hash index in amber.
- Case numbering is private: public matrices use stable aliases (A-xxxxxxxx, hash-derived) plus bundle hashes only.
- Tone: community measurement, not vendor attacks.
Same model name, same provider, two runs can still score differently — inference parameters, load, and server-side versions drift. Relay/aggregator lanes add an upstream routing layer: the same name may not be the same endpoint. Every conclusion here is dated and banded, and we re-test periodically. A single day's number is a snapshot, not a law.
| Issue | Content | Headline |
|---|---|---|
| 2026-W37 | deepseek/deepseek-v4.1-flash full-library debut (23 cases, GA day) | 17/23 (15/21 public subset); full marks on the UI-build paper (fifth repo to publish a pass); 8/9 on the defense case that used to kill everyone; all three same-name lanes verified genuine v4.1; adversarial-review and vision are its weak faces |
Not affiliated with or sponsored by CommandCode or DeepSeek. Scores are dated, band-specific snapshots, not purchasing advice.