diff --git a/.claude/commands/launch-llms-benchmark.md b/.claude/commands/launch-llms-benchmark.md index a2ffb87..28b732c 100644 --- a/.claude/commands/launch-llms-benchmark.md +++ b/.claude/commands/launch-llms-benchmark.md @@ -1,12 +1,20 @@ --- -description: Benchmark one LLM on the llms-benchmark demo stack - drive it through a coding-agent CLI (opencode, claude or copilot) on the stored scenario, grade the observation report it produced, and propose its row of the results table -argument-hint: " [effort, default medium]" +description: Benchmark one LLM on the llms-benchmark demo stack - drive it through a coding-agent CLI (opencode, claude or copilot) or a vllm-on-tap preset served for the run, on the stored scenario, grade the observation report it produced, and propose its row of the results table +argument-hint: " [effort, default medium] | local [effort, default medium]" --- Run the whole llms-benchmark protocol for one model on one CLI, end to end, and come back with a pull request adding or replacing that row in `.llms-benchmark/README.md`. +Two kinds of run, two tables. **Remote serving**: a provider serves a +model id to the CLI (OpenRouter, Anthropic, Copilot). **Local +serving**: the run serves a vllm-on-tap preset itself, on a vllm-on-tap +environment of this repository (`.vot/environments/`), and drives it +through `opencode` — always opencode, the only CLI that takes a custom +OpenAI-compatible endpoint per launch. Every step below applies to both; +a **Local serving** bullet says what changes for the second. + The question this benchmark answers: **how much of what a model reports, after observing a running stack it has never seen, actually holds up?** The demo stack under `.llms-benchmark/src/` is deliberately defective, @@ -40,6 +48,15 @@ the same way you would grade a colleague's incident report. nothing else changed). A model the CLI cannot run, or cannot run at the requested effort, is a preflight failure, not a row. Below, `` is that argument, passed verbatim to the CLI's flag. +- **Local serving** — `local [effort]`: the + literal `local`; the name of a vllm-on-tap environment + (`.vot/environments/.yaml`); the preset, as vllm-on-tap resolves + it (`.vot/presets/.yaml` first, then the builtins); the effort, + `medium` when omitted, as above. The CLI is `opencode`, never asked. + Preset, effort and **GPU** identify the row; the GPU is read off the + environment, never asked (step 1). Below, `` is the + preset's `served_model_name` (its `name` when absent) and `` + its merged `--max-model-len`. **Never ask for an API key, and never handle one.** Every credential this protocol needs — the OpenRouter provider in opencode, the Claude Code @@ -122,11 +139,39 @@ Steps: `docker compose` reads on its own from that file. Check its presence, never its value, and never print it. The file is gitignored; `.env.example` next to it says what goes in. + - **Local serving**, in place of the CLI-and-model check above (the + opencode binary is still checked, in step 3): + - `.vot/environments/.yaml` exists, and its stack + type is one the vllm-on-tap stack-guide supports; run that type's + *Prerequisites and install* checks (the `vllm-on-tap:vot-config` + skill's step 4), and stop on any failure — installing a tool or + logging in is the user's; + - the preset resolves and validates (the `vllm-on-tap:load-preset` + skill's *Resolve*, *Validate* and *Merge*, for the environment's + stack type) and carries the flags an agent needs: tool calling + (`--enable-auto-tool-choice` with a `--tool-call-parser`) and a + `--max-model-len` of at least 65536 — opencode's system prompt and + tool definitions alone take about 15k tokens, and a report-writing + turn carries the whole observation. Use the model's native + maximum context when the KV cache holds it. The benchmark serves its own + presets, `.vot/presets/odd-.yaml`, tuned for the run: a + preset without the `odd-` prefix, a builtin included, is refused + — copy it under an `odd-` name with the flags it lacks (the row + then names that preset); + - the environment's `otlp_endpoint`, when set, does not point at the + local oddyssey stack: the served model's own spans would land in + the store the run observes; + - **the GPU cell**, read off the environment: on `azure` the serve + profile's GPU, `A100 80 GB`; on `local-vllm-metal` + `sysctl -n machdep.cpu.brand_string` and `hw.memsize` in GB + (`Apple M4 Pro 24 GB`); on `local-vllm` and `local-vllm-docker` + `nvidia-smi --query-gpu=name,memory.total --format=csv,noheader`. 2. **Create the work branch**: `bench/---`, where `` is the model id with `/` and `.` replaced by `-`. Everything the run installs, configures, and produces happens on - this branch, and none of it is what ships. + this branch, and none of it is what ships. **Local serving**: + `bench/local----`. 3. **Install or update the CLI, and the oddyssey package for it.** Record the CLI's version: it goes in the pull request, never in the @@ -195,6 +240,69 @@ Steps: List them before launching: none of the package's nine skills may be there, and a run that lists a skill twice resolves one of them. + **Local serving** — step 3 is opencode's, then the preset is served + for the run: a serve takes ten to thirty minutes to answer (image + pull, weights, compile), and it bills from then on. + - record `.vot/config.yaml`'s `current`, then set it to + `` (vllm-on-tap serves on the current environment; + step 9 puts the recorded value back); + - run the `vllm-on-tap:vot-serve` skill's steps for ``, up to + its report: the base URL and the served name. **From here on the unit + bills until step 9's destroy, whatever happens in between** — a stop, + a failed smoke, a void run all end in step 9. Collect the unit's log to a file from its creation (on `azure`, + `az containerapp logs show --tail 300` every 20 s, new lines only: + history stops at 300 lines, the environment keeps none, and + `--follow` returned nothing on 2026-09-27). Watch it while it starts: on `EngineCore failed to start` destroy it at + once (it restarts in a loop and bills), fix the preset, serve again; + - write the provider file into your own scratch directory, never into + the repository nor opencode's user configuration: + + ```json + {"$schema": "https://opencode.ai/config.json", + "provider": {"vot": {"npm": "@ai-sdk/openai-compatible", "name": "vllm-on-tap", + "options": {"baseURL": "/v1", "apiKey": "{env:VOT_API_KEY}"}, + "models": {"": {"name": "", "tool_call": true, + "limit": {"context": , "output": 32768}}}}}} + ``` + + Add `"permission": {"bash": {"opencode": "deny", "opencode *": "deny"}}`: + a nested `opencode run` uses the user's default model and voids the run. + + A model that takes a reasoning effort (served with a + `--reasoning-parser`, and whose chat template reads + `reasoning_effort`) also gets `"reasoning": true` and + `"variants": {"low": {"reasoningEffort": "low"}, "medium": {"reasoningEffort": "medium"}, "high": {"reasoningEffort": "high"}}` + in its model entry, so step 4's `--variant ` sends the + effort instead of being ignored; + + Keep `output` at 32768: it bounds the reasoning too. + + `OPENCODE_CONFIG=` on a launch line adds the `vot` + provider to that launch only: the user's configuration and + `~/.local/share/opencode/auth.json` (the OpenRouter key) are never + written, so a remote run after a local one needs no switch back + (verified on 2026-09-27: `opencode models vot` listed the model, + `opencode models openrouter` still answered, `auth.json` unchanged); + - `VOT_API_KEY` on `azure` is the app's secret, read inline on the + launch line and never printed: + `VOT_API_KEY="$(az containerapp secret show --name vot- --resource-group --secret-name vllm-api-key --query value -o tsv)"`; + on a local stack type vLLM checks no key, and `VOT_API_KEY=none`; + - **smoke the served model through opencode before any run**, from a + scratch directory: `OPENCODE_CONFIG= VOT_API_KEY=... opencode run --model vot/ --format json "run the shell command date and reply with its output" < /dev/null` + must show a `tool_use` event for `bash` and a `text` event: a serve + without vLLM's tool-call parser answers text only, and a mission on + it never drives. A failed smoke is a preflight failure; step 9 still + destroys the unit. + - measure decode at about 80k tokens of context before the run (a + streamed request, time to first token apart): the runs' prompts reach + 200k. A preset whose run cannot end within 40 minutes is not run. + - a model whose chat template gates thinking (Gemma 4: `enable_thinking`, + off by default) gets it through the model entry's + `"options": {"chat_template_kwargs": {"enable_thinking": true}}`, and + its preset through `--default-chat-template-kwargs`. + - on an A100 (Triton attention) an FP8 KV cache is refused (SM89+); + `int8_per_token_head` starts but decodes far slower at long context. + 4. **Select the model.** Nothing to configure: the provider is already set up (preflight), and the model and effort are passed on the command line in step 6, never persisted into a config file — `opencode`: @@ -202,7 +310,10 @@ Steps: `--model --effort `; `copilot`: `--model --effort `. The three flags name the same effort level; that is what makes two rows of one model at one effort - comparable across CLIs. + comparable across CLIs. **Local serving**: `--model vot/`, + with `--variant ` only when `OPENCODE_CONFIG= opencode models vot --verbose` + lists that variant for the model — otherwise no `--variant` and + `default` in the Effort column, step 1's opencode rule. 5. **Clean what the next run must not read — then recreate the demo stack, never reuse a running one.** Before every run, whatever the @@ -298,6 +409,25 @@ Steps: "" < /dev/null ``` + **Local serving** — the same line, with the provider file and the key + in front, and the preset in the title (step 7 selects on it): + + ``` + OPENCODE_CONFIG= VOT_API_KEY="" \ + caffeinate -i opencode run --model vot/ [--variant ] \ + --format json --auto --title "llms-benchmark " \ + "" < /dev/null + ``` + + **Local serving runs the mission once, not twice, and stops at 40 + minutes** (no provider varies + between two runs). The row is that single run, and the + two-run rules of this step do not apply to it. A void attempt (step + 8's shapes, a run measuring another model) is re-run once on the + same unit, after step 9's teardown (without its destroy) and step 5; + two void attempts and the preset's result is "did not drive the + scenario", with no row. + `claude` — generate the session id yourself and write it down: this CLI takes no title, and step 7 identifies the session by that id: @@ -672,6 +802,15 @@ Steps: If that returns anything other than exactly one row, stop and say so rather than guess. + **Local serving**: the title is `llms-benchmark `, the model + id is `` and `json_extract(model,'$.providerID')` is + `vot`. The tokens are summed exactly as below; **there is no Cost**: + the endpoint bills no tokens (`SUM(cost)` is 0), and the GPU's own + bill runs from the serve to the destroy, the serve's start included, + so it is no figure of the run's. The pull request states + the serve's start and destroy times and, on `azure`, the unit's + GPU-hours; the table has no money column. + **Then sum the whole tree, not the root.** opencode dispatches the observation to a subagent, which gets its own session; on the first run the subagent carried 80% of the spend. Walk `parent_id` @@ -1005,6 +1144,13 @@ Steps: step 10 and let the branch take the file with it; - `git checkout main`, then delete the work branch (`git branch -D`) — it never gets pushed. + - **Local serving, after the last attempt** — the run, or the + preflight or the attempt that failed: run the `vllm-on-tap:vot-destroy` skill's steps for + `` and check the unit is gone (on `azure`, + `az containerapp show --name vot- ...` fails); put the + `current` recorded in step 3 back into `.vot/config.yaml`; delete + the provider file. Never end the command with the unit up: on + `azure` it bills a GPU by the second until it is destroyed. 10. **Open the results PR from a clean base.** - **Open the run's issue first.** Every PR in this repository @@ -1014,7 +1160,10 @@ Steps: - From `main`, freshly pulled, create `docs/llms-benchmark---` and make **one** change: the row in the results tables of `.llms-benchmark/README.md`. - `## Results` holds the two tables below. **A row is identified by + `## Results` holds two subsections: `### Remote serving`, the two + tables below, and `### Local serving`, the two tables of the local + serving section after them. **Local serving**: the branch is + `docs/llms-benchmark-local---`. **A row is identified by model, effort, CLI and provider together.** That key is not in the table yet → append the row; already there → replace that row in place. The same model driven through two CLIs, at two efforts, or @@ -1122,6 +1271,35 @@ Steps: until it is re-run. Changing a weight or a bound re-scores every row and is the maintainer's decision. + **Local serving — its own two tables, under `### Local serving`.** + The same shape, with three changes: **Model** becomes **Preset** (the + preset's name in backticks, linked to its YAML: + `` [``](../.vot/presets/.yaml) ``), **Provider** becomes **GPU** (step 1's cell: `A100 80 GB`, + `Apple M4 Pro 24 GB`), and **Cost** and **$/confirmed** are dropped + — the endpoint bills no tokens (step 7). A row is identified by + preset, effort, CLI and GPU; the CLI is always `opencode`. Headline, + twelve columns: + + ```text + | Rank | Preset | Effort | CLI | GPU | oddyssey | Scoring | Confirmed / reported | Telemetry / Perf / Behavior | Total | Accuracy | seconds/confirmed | + ``` + + and detail, the remote detail table's columns with Preset and GPU in + place of Model and Provider: + + ```text + | Preset | Effort | CLI | GPU | oddyssey | Preflight | Drive | Observation | Turns | Median turn | Input | Output | Cache | Signals | + ``` + + Its Scoring keeps the four remaining axes and their bounds, the + weights renormalized so they still sum to one (the maintainer's + decision of 2026-09-27): **seconds/confirmed** 3/7, **Total** 2/7, + **Accuracy** 1/7, **Confirmed** 1/7. With no cost, a tie in the + rank is broken on the shorter run; the row is step 6's single local + run, and the pull request carries its rulings alone. + The two tables are never merged or ranked against each other: the + remote score carries a cost the local one cannot. + Cost per confirmed finding is the column that answers the question in the README's title: cost and duration alone reward whichever model gives up soonest. The breakdown by kind exists because a run can diff --git a/.gitignore b/.gitignore index 7db7403..4349f7a 100644 --- a/.gitignore +++ b/.gitignore @@ -223,3 +223,5 @@ __marimo__/ apm.lock.yaml # ...except the generated marketplace plugin, which ships its own MCP config !marketplace/oddyssey/.mcp.json +.vot/config.yaml +.vot/run/ diff --git a/.llms-benchmark/README.md b/.llms-benchmark/README.md index d31e37a..ee5974b 100644 --- a/.llms-benchmark/README.md +++ b/.llms-benchmark/README.md @@ -12,6 +12,13 @@ only variables are the model, its effort, the CLI and the provider. ## Results +Two tables, never ranked against each other: **remote serving**, where a +provider serves the model and bills its tokens, and **local serving**, +where the run serves an open-weight model itself and the only bill is +the GPU's. + +### Remote serving + One row per model, effort, CLI and provider, always its latest run. | Rank | Model | Effort | CLI | Provider | oddyssey | Scoring | Confirmed / reported | Telemetry / Perf / Behavior | Total | Cost | Accuracy | $/confirmed | seconds/confirmed | @@ -94,6 +101,32 @@ design. A row measured under an earlier revision of the protocol is marked ⚠︎ and provisional until re-run. The table keeps no history: one row per model, effort, CLI and provider, its latest run. +### Local serving + +One row per preset, effort, CLI and GPU, always its latest run. The model +is a [vllm-on-tap](https://github.com/using-system/vllm-on-tap) preset +served with vLLM for the run, on a GPU in the cloud or on the machine +itself, and driven through opencode. + +| Rank | Preset | Effort | CLI | GPU | oddyssey | Scoring | Confirmed / reported | Telemetry / Perf / Behavior | Total | Accuracy | seconds/confirmed | +| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | +| **#1** | [`odd-qwen3-6-35b-a3b`](../.vot/presets/odd-qwen3-6-35b-a3b.yaml) | default | opencode | A100 80 GB | 1.13.0 | **18.5** | 3 / 9 | 1 / 2 / 0 | **21m38s** | 33% | **433s** | +| **#2** | [`odd-qwen3-8-27b`](../.vot/presets/odd-qwen3-8-27b.yaml) | default | opencode | A100 80 GB | 1.13.0 | 9.0 | **7 / 11** | 4 / 3 / 0 | 150m02s | **64%** | 1286s | + +
+Run detail — phases, turns, tokens + +| Preset | Effort | CLI | GPU | oddyssey | Preflight | Drive | Observation | Turns | Median turn | Input | Output | Cache | Signals | +| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | +| [`odd-qwen3-6-35b-a3b`](../.vot/presets/odd-qwen3-6-35b-a3b.yaml) | default | opencode | A100 80 GB | 1.13.0 | 1m20s | 2m00s | 18m18s | 76 | 7.5s | 9.0M | 37k | — | 4/4 | +| [`odd-qwen3-8-27b`](../.vot/presets/odd-qwen3-8-27b.yaml) | default | opencode | A100 80 GB | 1.13.0 | 12m13s | 2m04s | 135m45s | 56 | 65.7s | 5.0M | 182k | — | 4/4 | + +
+ +- **Preset** is the vllm-on-tap preset served, **GPU** the hardware it ran on (`A100 80 GB` for a serverless Azure GPU, the chip and its memory on a local machine). Preset, effort, CLI and GPU identify a row. +- **No Cost, no $/confirmed**: a served model bills no tokens, and the GPU bills by the hour whatever the run does. +- **Scoring** keeps the four other axes and their bounds, weighted **seconds/confirmed** 3/7, **Total** 2/7, **Accuracy** 1/7 and **Confirmed** 1/7; a tie goes to the shorter run. Every other column reads as in the remote table. + ## How a row is produced ```text @@ -101,10 +134,13 @@ A row measured under an earlier revision of the protocol is marked ⚠︎ and pr /launch-llms-benchmark claude anthropic/claude-haiku-4.5 /launch-llms-benchmark copilot openai/gpt-5.6-luna /launch-llms-benchmark copilot openai/gpt-5.6-sol high +/launch-llms-benchmark local azure-sweden odd-qwen3-8-27b ``` The CLI and the model id, in `vendor/name` form, are required; an optional third argument sets the effort (`medium` by default). Prerequisites, set up once: an OpenRouter provider in opencode, a Claude Code login with the package installed at user scope, or a Copilot CLI login; and `OPENAI_API_KEY` in `docker-compose/llms-benchmark/.env` for the demo agent's own model calls (`.env.example` next to it). +A local serving row takes `local`, a vllm-on-tap environment of this repository (`.vot/environments/`), one of the benchmark's presets (`.vot/presets/odd-*.yaml`), and an optional effort; it needs opencode and the vllm-on-tap plugin. The command serves the preset once, runs the mission on it, and destroys the served model at the end. + The command cleans everything a run must not read (stored reports of the three services, leftovers, the local stack's data), recreates the demo stack, drives the model headless at the requested effort through one `/odd-observe` mission naming the three services, the stored scenario and the local stack, grades the report finding by finding on evidence, and opens the pull request carrying the row and the rulings. ## The stack under observation diff --git a/.vot/environments/azure-sweden.yaml b/.vot/environments/azure-sweden.yaml new file mode 100644 index 0000000..b30654f --- /dev/null +++ b/.vot/environments/azure-sweden.yaml @@ -0,0 +1,5 @@ +name: azure-sweden +stack: azure +config: + location: swedencentral + telemetry_enabled: true diff --git a/.vot/environments/local-metal.yaml b/.vot/environments/local-metal.yaml new file mode 100644 index 0000000..77fc2bf --- /dev/null +++ b/.vot/environments/local-metal.yaml @@ -0,0 +1,4 @@ +name: local-metal +stack: local-vllm-metal +config: + port: 8000 diff --git a/.vot/presets/odd-qwen3-6-35b-a3b.yaml b/.vot/presets/odd-qwen3-6-35b-a3b.yaml new file mode 100644 index 0000000..32a37a6 --- /dev/null +++ b/.vot/presets/odd-qwen3-6-35b-a3b.yaml @@ -0,0 +1,17 @@ +name: odd-qwen3-6-35b-a3b +description: Qwen3.6 35B-A3B FP8 (MoE, ~3B active) tuned for the oddyssey llms-benchmark on one A100 80 GB - FP8 weights through Marlin, MTP speculative decoding, 256k context (the model's native maximum), tool calling and reasoning parsers, text only; NVIDIA stacks only +model: Qwen/Qwen3.6-35B-A3B-FP8 +served_model_name: qwen3.6-35b-a3b +gpu_memory_gb: 48 +vllm_args: + --max-model-len: 262144 + --gpu-memory-utilization: 0.92 + --max-num-seqs: 8 + --max-num-batched-tokens: 32768 + --enable-prefix-caching: true + --language-model-only: true + --speculative-config: '{"method": "mtp", "num_speculative_tokens": 3}' + --enable-auto-tool-choice: true + --tool-call-parser: qwen3_coder + --reasoning-parser: qwen3 +env: [] diff --git a/.vot/presets/odd-qwen3-8-27b.yaml b/.vot/presets/odd-qwen3-8-27b.yaml new file mode 100644 index 0000000..93625e8 --- /dev/null +++ b/.vot/presets/odd-qwen3-8-27b.yaml @@ -0,0 +1,17 @@ +name: odd-qwen3-8-27b +description: Qwen3.8 27B FP8 tuned for the oddyssey llms-benchmark on one A100 80 GB - FP8 weights through Marlin, MTP speculative decoding, 256k context (the model's native maximum), tool calling and reasoning parsers, text only; NVIDIA stacks only +model: Qwen/Qwen3.8-27B-FP8 +served_model_name: qwen3.8-27b +gpu_memory_gb: 40 +vllm_args: + --max-model-len: 262144 + --gpu-memory-utilization: 0.92 + --max-num-seqs: 8 + --max-num-batched-tokens: 32768 + --enable-prefix-caching: true + --language-model-only: true + --speculative-config: '{"method": "mtp", "num_speculative_tokens": 3}' + --enable-auto-tool-choice: true + --tool-call-parser: qwen3_coder + --reasoning-parser: qwen3 +env: []