From 6921ae9c3d1ef1fb71b32facbee8dd217aa34cfa Mon Sep 17 00:00:00 2001 From: Dominik Simonik Date: Fri, 18 Sep 2026 14:46:04 +0200 Subject: [PATCH 01/18] feat(evals): mock the very-good-cli MCP server Adds evals/mocks/very-good-cli/ with fixed mocks for create, packages_get, test, and packages_check_licenses, plus the real tools/list response as _tools.json so mocked tools carry their production schemas. Registration is confirmed: a run now reports mocked: very-good-cli(create=fixed, packages_check_licenses=fixed, packages_get=fixed, test=fixed) No case calls them yet. Every prompt still says the session cannot run anything and every allowed_tools still lists only [Read, Glob, Grep, Skill], so the tool-driven skills stay graded on narration until both change. The SessionStart hook still fires under mocks. check_vgv_cli gates on `command -v very_good`, not on MCP availability, so the three prompts that tell the model to ignore the startup notice keep that sentence. CI: evals/mocks/ is a shared input, so a change to it widens scope to all skills the way _fixture/ does. The scope job previously emitted any changed evals/ subdirectory as a --tag, which would have made `mocks` a tag matching no cases; it now requires the directory to hold */case.yaml. That guard runs as an if-statement because the && chain it replaced could return non-zero from the loop body and trip `set -e`. Co-Authored-By: Claude Opus 5 --- .github/workflows/evals.yaml | 8 +- config/cspell.json | 1 + evals/README.md | 81 ++++++- evals/mocks/very-good-cli/_tools.json | 205 ++++++++++++++++++ evals/mocks/very-good-cli/create.md | 10 + .../very-good-cli/packages_check_licenses.md | 26 +++ evals/mocks/very-good-cli/packages_get.md | 9 + evals/mocks/very-good-cli/test.md | 17 ++ 8 files changed, 349 insertions(+), 8 deletions(-) create mode 100644 evals/mocks/very-good-cli/_tools.json create mode 100644 evals/mocks/very-good-cli/create.md create mode 100644 evals/mocks/very-good-cli/packages_check_licenses.md create mode 100644 evals/mocks/very-good-cli/packages_get.md create mode 100644 evals/mocks/very-good-cli/test.md diff --git a/.github/workflows/evals.yaml b/.github/workflows/evals.yaml index 645abd6..4bdfeb5 100644 --- a/.github/workflows/evals.yaml +++ b/.github/workflows/evals.yaml @@ -82,12 +82,14 @@ jobs: echo "$CHANGED" | { grep -oE '^skills/[^/]+/' || true; } | cut -d/ -f2 echo "$CHANGED" | { grep -oE '^evals/[^_][^/]*/' || true; } | cut -d/ -f2 } | sort -u | while read -r s; do - [ -n "$s" ] && [ -d "evals/$s" ] && echo "$s" + if [ -n "$s" ] && compgen -G "evals/$s/*/case.yaml" >/dev/null; then + echo "$s" + fi done ) - if echo "$CHANGED" | grep -qE '^evals/_fixture/'; then - echo "fixture changed — widening to all skills" + if echo "$CHANGED" | grep -qE '^evals/(_fixture|mocks)/'; then + echo "shared eval input changed — widening to all skills" SKILLS=$(all_skills) fi fi diff --git a/config/cspell.json b/config/cspell.json index cf6cadb..99c4f7b 100644 --- a/config/cspell.json +++ b/config/cspell.json @@ -13,6 +13,7 @@ "Bienvenido", "bypassable", "Codex", + "compgen", "CSPRNG", "cupertino", "Cupertino", diff --git a/evals/README.md b/evals/README.md index b9f30f4..a66c0e5 100644 --- a/evals/README.md +++ b/evals/README.md @@ -103,6 +103,10 @@ threshold is `1.0`, so **pass `--threshold 0.8` or every imperfect case exits 1* There are **no custom-code graders**. A check that needs to execute something has no home. +A grader reads `last_message` unless its `target` says otherwise. The other targets are +`trace`, `files`, `{ source: file, path: }`, and `mock_calls`, which is every call +made to a [mocked MCP tool](#mocking-the-mcp-servers). + ### Routing graders Ninety-nine of the hundred cases carry one. @@ -146,7 +150,9 @@ Negative controls use the same grader with `min: 0` and `max: 0`. injects "Very Good CLI is not installed" into every with-plugin run, because the sandbox has no `very_good` on PATH. Tool-driven cases answer with that blocker instead of the question. Say in the prompt that the CLI is installed and that a startup notice saying - otherwise should be ignored. + otherwise should be ignored. **Mocking the server does not silence it.** `check_vgv_cli` + tests `command -v very_good`, so it reports on the PATH and never on whether the MCP + tools are reachable. A mocked run carries the same warning an unmocked one does. Beyond that: write prompts as a user would send them, name no skill in a prompt, grade mechanically where you can, include the cases where the skill must say no, keep a negative @@ -199,6 +205,62 @@ holding the path, and runs then fail at scaffold time. CI runs on Linux and is u --- +## Mocking the MCP servers + +A run never starts the plugin's real MCP servers. `evals/mocks//.md` registers +a stand-in under the server's own name from `.mcp.json`, and a mocked tool is callable +without an `--allow-tools` grant. A server with no mock directory is not started at all and +its tools are absent, which every run reports on a `mocked:` line. + +`very-good-cli` is mocked. `dart` is not, so runs still print +`plugin_vgv-ai-flutter-plugin_dart[not started: no mock]`. + +```text +evals/mocks/very-good-cli/ +├── _tools.json # the real tools/list response +├── create.md +├── packages_check_licenses.md +├── packages_get.md +└── test.md +``` + +A mock file is frontmatter plus a body, and the body is the tool result: + +| Key | Default | Purpose | +| ------------ | ------- | --------------------------------------------------------------- | +| `type` | `fixed` | `agent` instead plays the server through a judge-model call | +| `expect` | unset | Per-input guard: a type name, a literal, a list, or a `/regex/` | +| `error` | `false` | `fixed` only. Return the body as a tool error | +| `abort_when` | unset | `agent` only | + +`{{input.}}` substitutes a call argument into the body, and +`{{file:fixtures/}}` inserts a file from a `fixtures/` directory beside the mock. +`_server.md` answers several tools from one `agent` mock; a `.md` for the same tool +wins. A case's own `mocks/` directory overrides the suite's file by file. + +**`expect` aborts, it does not fail.** A call that violates it ends the run at score 0, +reported as `aborted`, with no failing grader to read. Keep it to what the real server's +schema already enforces and grade argument *choices* with `tool_used` or with a `regex` +against the `mock_calls` target, which carries every mocked call, its input, and the +answer. + +Every mock here is `type: fixed`. An `agent` mock answers through the judge model, so it +costs money, varies run to run, and needs a recording adopted from +`results//mock-recordings/` into `.replay/` before CI repeats. + +`_tools.json` is the real `tools/list` response, so a mocked tool carries the real +descriptions and input schemas rather than a permissive placeholder. Regenerate it after a +Very Good CLI release by speaking MCP to `very_good mcp` over stdio and saving the +`tools/list` result. A stale one teaches the model a schema the CLI no longer has. + +`packages_check_licenses` returns a deliberately mixed result, one `GPL-3.0` and one +`unknown` among twelve permissive licenses, so a case has something real to flag. + +Keep the bodies as raw tool output. A mock that already names what a skill teaches hands +the answer to the no-plugin arm, exactly as a non-neutral fixture does. + +--- + ## Running ```bash @@ -244,7 +306,13 @@ scoped by `--tag` to the changed skills, with-plugin arm only, and `continue-on- A regression is therefore reported after it lands, and a case that has stopped discriminating goes unnoticed until you re-check with `include_baseline`. -- Changing `_fixture/` widens the scope to all 15 skills. +- Changing `_fixture/` or `mocks/` widens the scope to all 15 skills. Both are shared + inputs, so a change to either can move any case. +- The scope job's `find` writes `-exec dirname {} \;` rather than the shorter + `-printf '%h\n'`. `-printf` is a GNU extension that BSD `find` does not have, so the + short form passes in CI and fails for anyone running the same job on macOS. +- A directory under `evals/` only becomes a `--tag` if it actually holds `*/case.yaml`. + `mocks/` and `results/` sit there without being skills. - CI needs `ANTHROPIC_API_KEY`, having no Claude Code session, and `--trust-plugin`, because a run with no terminal cannot answer the trust prompt. - `--ablation none` is the only mode where a `tool_used: Skill` grader is scored by @@ -263,9 +331,12 @@ discriminating goes unnoticed until you re-check with `include_baseline`. a `tracePath` into a sandbox deleted unless `--keep-temp` is passed. `report.html` does show the judged text. - **Judge calibration.** Most graders are `llm` with no human-labelled gold set. -- **Tool execution.** The six tool-driven skills are graded on the decisions they narrate. - `claude plugin eval` can mock MCP servers under `evals/mocks//.md`, which - would grade them on the calls they actually make. Nothing here uses it yet. +- **Tool execution.** `very-good-cli` is now mocked, so its four tools can be called + inside a run and graded with `tool_used` or against `mock_calls`. No case does yet: every + prompt still says the session cannot run anything, and every `allowed_tools` still lists + only `[Read, Glob, Grep, Skill]`. Until both change, the tool-driven skills stay graded on + the calls they narrate rather than the calls they make. The `dart` server has no mock at + all. - **Stable routing.** Whether a skill activates is nondeterministic, which is why routing is a `tool_used` grader rather than inferred from content. - **Prose in a `SKILL.md`.** Deliberate. An earlier version asserted a hundred `contains` diff --git a/evals/mocks/very-good-cli/_tools.json b/evals/mocks/very-good-cli/_tools.json new file mode 100644 index 0000000..032c6bd --- /dev/null +++ b/evals/mocks/very-good-cli/_tools.json @@ -0,0 +1,205 @@ +{ + "tools": [ + { + "name": "create", + "description": "Create a very good Dart or Flutter project in seconds based on the provided template. Each template has a corresponding sub-command.\n ", + "inputSchema": { + "type": "object", + "properties": { + "subcommand": { + "type": "string", + "description": "The available subcommands to provide an specific template, are:\napp_ui_package - Generate a Very Good App UI package.\ndart_cli - Generate a Very Good Dart CLI application.\ndart_package - Generate a Very Good Dart package.\ndocs_site - Generate a Very Good documentation site.\nflame_game - Generate a Very Good Flame game.\nflutter_app - Generate a Very Good Flutter application.\nflutter_package - Generate a Very Good Flutter package.\nflutter_plugin - Generate a Very Good Flutter plugin.\n", + "enum": [ + "app_ui_package", + "flame_game", + "flutter_app", + "flutter_package", + "flutter_plugin", + "dart_cli", + "dart_package", + "docs_site" + ] + }, + "name": { + "type": "string", + "description": "Project name" + }, + "description": { + "type": "string", + "description": "The description for this new project.\n(defaults to \"A Very Good Project created by Very Good CLI.\")" + }, + "org_name": { + "type": "string", + "description": "The organization for this new project.\n(defaults to \"com.example.verygoodcore\")" + }, + "output_directory": { + "type": "string", + "description": "The desired output directory when creating a new project." + }, + "application_id": { + "type": "string", + "description": "The bundle identifier on iOS or application id on Android. (defaults to .)" + }, + "platforms": { + "type": "string", + "description": "Comma-separated platforms. 'Example: \"android,ios,web\".\nThe values for platforms are: android, ios, web, macos, linux, and windows.\nIf is omitted, then all platforms are enabled by default.\nOnly available for subcommands: flutter_plugin with all values) and flame_game (only android and ios)\n " + }, + "publishable": { + "type": "boolean", + "description": "Whether package is intended for publishing (flutter_package, dart_package only)" + }, + "workspace": { + "type": "boolean", + "description": "Whether the generated project should resolve its dependencies from a parent Pub workspace." + }, + "executable-name": { + "type": "string", + "description": "CLI custom executable name (dart_cli only)" + }, + "template": { + "type": "string", + "description": "The template used to generate this new project.\nThe values are:\ncore - Generate a Very Good Flutter application.\nIf is omitted, then core will be selected.\n" + } + }, + "required": [ + "subcommand", + "name" + ] + } + }, + { + "name": "test", + "description": "Run tests in a Dart or Flutter project.", + "inputSchema": { + "type": "object", + "properties": { + "directory": { + "type": "string", + "description": "The package to test (defaults to current directory). Can be absolute or relative path to project root. A package path belongs here, not in \"paths\"." + }, + "paths": { + "type": "array", + "description": "Test files or directories to run, relative to the package selected by \"directory\" (e.g. ['test/src/foo_test.dart', 'test/widgets']). These are targets inside that one package: to test a different package, set \"directory\" to it rather than putting its path here. When omitted, the whole suite runs. Cannot be combined with \"recursive\". Note that targeting specific paths disables the test optimization step.", + "items": { + "type": "string" + } + }, + "dart": { + "type": "boolean", + "description": "Whether to run Dart tests. If not specified, Flutter tests will be run if a Flutter project is detected." + }, + "coverage": { + "type": "boolean", + "description": "Whether to collect coverage information." + }, + "recursive": { + "type": "boolean", + "description": "Run tests recursively for all nested packages. Cannot be combined with \"paths\", because each package runs in its own working directory and a path is only meaningful within one of them." + }, + "optimization": { + "type": "boolean", + "description": "Whether to apply optimizations for test performance.\nAutomatically disabled when --platform is specified.\nAdd the `skip_very_good_optimization` tag to specific test files to disable them individually.\n(defaults to on)" + }, + "concurrency": { + "type": "string", + "description": "The number of concurrent test suites run.\nAutomatically set to 1 when --platform is specified.\n(defaults to \"4\")" + }, + "tags": { + "type": "string", + "description": "Run only tests associated with the specified tags." + }, + "exclude_coverage": { + "type": "string", + "description": "A glob which will be used to exclude files that match from the coverage (e.g. '**/*.g.dart')." + }, + "exclude_tags": { + "type": "string", + "description": "Run only tests that do not have the specified tags." + }, + "min_coverage": { + "type": "string", + "description": "Whether to enforce a minimum coverage percentage." + }, + "test_randomize_ordering_seed": { + "type": "string", + "description": "The seed to randomize the execution order of test cases within test files." + }, + "update_goldens": { + "type": "boolean", + "description": "Whether \"matchesGoldenFile()\" calls within your test methods should update the golden files." + }, + "force_ansi": { + "type": "boolean", + "description": "Whether to force ansi output. If not specified, it will maintain the default behavior based on stdout and stderr." + }, + "dart-define": { + "type": "string", + "description": "Additional key-value pairs that will be available as constants from the String.fromEnvironment, bool.fromEnvironment, int.fromEnvironment, and double.fromEnvironment constructors.\nMultiple defines can be passed by repeating \"--dart-define\" multiple times.\n(e.g., foo=bar)\n" + }, + "dart-define-from-file": { + "type": "string", + "description": "The path of a .json or .env file containing key-value pairs that will be available as environment variables. These can be accessed using the String.fromEnvironment, bool.fromEnvironment, and int.fromEnvironment constructors.\nMultiple defines can be passed by repeating \"--dart-define-from-file\" multiple times. Entries from \"--dart-define\" with identical keys take precedence over entries from these files." + }, + "platform": { + "type": "string", + "description": "The platform to run tests on.\nThe available values are: chrome, vm, android, ios.\nOnly one value can be selected.\n " + }, + "run_skipped": { + "type": "boolean", + "description": "Run skipped tests instead of skipping them. Only applies to Dart tests (dart: true)." + }, + "check_ignore": { + "type": "boolean", + "description": "Whether to check for and respect coverage ignore comments (e.g. // coverage:ignore-line). Only applies to Dart tests (dart: true)." + }, + "show_uncovered": { + "type": "boolean", + "description": "Whether to list uncovered lines when coverage is below 100%. Implicitly enables coverage collection when used alone. Useful for identifying which lines still need tests after a min_coverage failure." + }, + "timeout_seconds": { + "type": "integer", + "description": "Maximum seconds to wait for the test run before killing the Flutter test process. Flutter tests can hang indefinitely when pumpAndSettle() is called without a timeout argument. When omitted, no timeout is applied." + } + } + } + }, + { + "name": "packages_get", + "description": " Install or update a Dart/Flutter package dependencies.\n Use after creating a project or modifying pubspec.yaml.\n Supports recursive installation and package exclusion.", + "inputSchema": { + "type": "object", + "properties": { + "directory": { + "type": "string", + "description": "Target directory path (defaults to current directory). Can be absolute or relative path to project root." + }, + "recursive": { + "type": "boolean", + "description": "Install dependencies for all nested packages recursively. Useful for monorepos or projects with multiple packages." + }, + "ignore": { + "type": "string", + "description": "Comma-separated list of package names to skip. Example: \"package1,package2\". Useful to avoid processing problematic packages." + } + } + } + }, + { + "name": "packages_check_licenses", + "description": " Verify package licenses for compliance and validation in a Dart or Flutter project.\n Identifies license types (MIT, BSD, Apache, etc.) for all\n dependencies. Use to ensure license compatibility.", + "inputSchema": { + "type": "object", + "properties": { + "directory": { + "type": "string", + "description": "Target directory path (defaults to current directory). Path to the project root containing pubspec.yaml." + }, + "licenses": { + "type": "boolean", + "description": "Verify all package licenses (defaults to true). Currently the only supported check type. Reports license types for all dependencies." + } + } + } + } + ] +} diff --git a/evals/mocks/very-good-cli/create.md b/evals/mocks/very-good-cli/create.md new file mode 100644 index 0000000..05fd86a --- /dev/null +++ b/evals/mocks/very-good-cli/create.md @@ -0,0 +1,10 @@ +--- +expect: + subcommand: [app_ui_package, flame_game, flutter_app, flutter_package, flutter_plugin, dart_cli, dart_package, docs_site] + name: string + org_name: string + output_directory: string +--- + +✓ Generated 87 file(s) from the {{input.subcommand}} template. +Created {{input.name}}. diff --git a/evals/mocks/very-good-cli/packages_check_licenses.md b/evals/mocks/very-good-cli/packages_check_licenses.md new file mode 100644 index 0000000..24fa77d --- /dev/null +++ b/evals/mocks/very-good-cli/packages_check_licenses.md @@ -0,0 +1,26 @@ +--- +expect: + directory: string + licenses: boolean +--- + +Retrieved 14 licenses from 14 packages. + +| package | version | license | +| ------------------------ | ------- | ------------ | +| async | 2.13.0 | BSD-3-Clause | +| characters | 1.4.1 | BSD-3-Clause | +| collection | 1.19.1 | BSD-3-Clause | +| cupertino_icons | 1.0.8 | MIT | +| glyph_atlas | 0.7.1 | GPL-3.0 | +| http | 1.5.0 | BSD-3-Clause | +| material_color_utilities | 0.13.0 | Apache-2.0 | +| meta | 1.17.0 | BSD-3-Clause | +| path | 1.9.1 | BSD-3-Clause | +| refresh_indicator_x | 2.4.0 | unknown | +| source_span | 1.10.1 | BSD-3-Clause | +| string_scanner | 1.4.1 | BSD-3-Clause | +| term_glyph | 1.2.2 | BSD-3-Clause | +| vector_math | 2.2.0 | BSD-3-Clause | + +License types: BSD-3-Clause (10), MIT (1), Apache-2.0 (1), GPL-3.0 (1), unknown (1). diff --git a/evals/mocks/very-good-cli/packages_get.md b/evals/mocks/very-good-cli/packages_get.md new file mode 100644 index 0000000..0f70fe3 --- /dev/null +++ b/evals/mocks/very-good-cli/packages_get.md @@ -0,0 +1,9 @@ +--- +expect: + directory: string + recursive: boolean + ignore: string +--- + +Running "flutter pub get"... +✓ Retrieved dependencies (2.1s) diff --git a/evals/mocks/very-good-cli/test.md b/evals/mocks/very-good-cli/test.md new file mode 100644 index 0000000..1d9e883 --- /dev/null +++ b/evals/mocks/very-good-cli/test.md @@ -0,0 +1,17 @@ +--- +expect: + directory: string + recursive: boolean + coverage: boolean + min_coverage: string + optimization: boolean +--- + +Running "flutter test"... +00:04 +38: All tests passed! + +Collected coverage to coverage/lcov.info. +lines......: 94.2% (356 of 378 lines) + +lib/src/login/login_bloc.dart 81.0% 17 of 21 lines +lib/src/settings/settings_page.dart 72.7% 8 of 11 lines From 5d1ff8500968fefa61e90772f3700bf03c5da7ff Mon Sep 17 00:00:00 2001 From: Dominik Simonik Date: Fri, 18 Sep 2026 15:06:04 +0200 Subject: [PATCH 02/18] docs(evals): record what the mock and bloc runs measured Measured on Claude Code 2.1.270, model and judge claude-sonnet-5. bloc-writes-sealed-events-and-states: --runs 3 scored 1.00, 3/3, every grader passing. The CI reading of 0.60 with four simultaneous failures was run-to-run variance. Case, graders and skills/bloc/SKILL.md unchanged. Two things block a mocked very-good-cli call, both before any grader runs: 1. A mocked tool needs an --allow-tools grant. The published docs say it does not. Listing it in a case's allowed_tools is necessary and not sufficient. The tool name is the bare mcp__very-good-cli__ form. 2. check-vgv-cli.sh denies every mcp__*very-good-cli__* call in the sandbox. `command -v very_good` succeeds there, since the run inherits the host PATH; `very_good --version` returns nothing under the run's throwaway $HOME. check_vgv_cli reads that as not_installed and denies, so the mock is never reached. Resolving 2 means changing what the hook gates on, which is shipped plugin behavior, so it is documented rather than decided here. Co-Authored-By: Claude Opus 5 --- evals/BASELINE.md | 25 ++++++++++++++++++++++--- evals/README.md | 44 +++++++++++++++++++++++++++++++++++++++++--- 2 files changed, 63 insertions(+), 6 deletions(-) diff --git a/evals/BASELINE.md b/evals/BASELINE.md index b9ab8f1..6b9d378 100644 --- a/evals/BASELINE.md +++ b/evals/BASELINE.md @@ -34,6 +34,23 @@ The previous sweep, under a Haiku judge and before the refusal-shaped descriptio measured 5 misses out of 83, and the one before that measured 14. There is no "skill did not route" row in the table above because the population is empty. +## `bloc-writes-sealed-events-and-states`, re-measured 2026-09-18 + +The CI run on 2026-09-18 scored it 0.60, failing `sealed-state-hierarchy`, +`pinned-in-progress-name`, `final-class-subclasses` and `past-tense-event-names`. A +`--runs 3` re-measurement on the same model and judge CI uses scored **1.00, 3/3, every +grader passing on every run**. + +```bash +claude plugin eval . --scaffold --trust-plugin --ablation none --runs 3 \ + --threshold 0.8 --model claude-sonnet-5 --judge-model claude-sonnet-5 \ + --case bloc-writes-sealed-events-and-states +``` + +Four graders failing at once looked like a real miss and was not. The case, its graders and +`skills/bloc/SKILL.md` are unchanged. Treat it as the worked example of why one CI reading +is not a measurement. + ## Reading the two numbers that look bad **15 cases show Δ <= 0.** All 15 are negative controls, and that is the design: a model @@ -105,6 +122,8 @@ editing, because a single reading of a case is not a measurement. the switch, so it never explains a routing miss, but every answer after the skill fires ran on Haiku while the no-plugin arm ran on Sonnet. Its Δ is understated by an unmeasured amount. -- **Tool-driven skills are graded on narration.** The MCP servers are not mocked, so the - six skills that drive tools are measured on the calls they describe, not the calls they - make. +- **Tool-driven skills are graded on narration.** `very-good-cli` now has mocks, but they + are unreachable: the plugin's own `check-vgv-cli.sh` PreToolUse hook denies every + `mcp__*very-good-cli__*` call inside a run. The skills that drive tools are still + measured on the calls they describe, not the calls they make. See `README.md` → + Mocking the MCP servers. diff --git a/evals/README.md b/evals/README.md index a66c0e5..35a043c 100644 --- a/evals/README.md +++ b/evals/README.md @@ -208,9 +208,47 @@ holding the path, and runs then fail at scaffold time. CI runs on Linux and is u ## Mocking the MCP servers A run never starts the plugin's real MCP servers. `evals/mocks//.md` registers -a stand-in under the server's own name from `.mcp.json`, and a mocked tool is callable -without an `--allow-tools` grant. A server with no mock directory is not started at all and -its tools are absent, which every run reports on a `mocked:` line. +a stand-in under the server's own name from `.mcp.json`. A server with no mock directory is +not started at all and its tools are absent, which every run reports on a `mocked:` line. + +**Two things block a mocked call, and both bite before any grader runs.** + +**The grant.** The published docs say a mocked tool "is allowed without an `--allow-tools` +grant". Measured on Claude Code 2.1.270, it is not. Listing the tool in a case's +`allowed_tools` is necessary and not sufficient, and without the operator grant the run +reports: + +```text +not granted (missing --allow-tools grant, or a malformed entry): +mcp__very-good-cli__packages_check_licenses +``` + +So any case that drives a mocked tool needs both, and the tool name is the **bare** server +form, not the plugin-namespaced one: + +```bash +claude plugin eval . --scaffold --allow-tools 'mcp__very-good-cli__packages_check_licenses' +``` + +**The plugin's own PreToolUse hook.** `check-vgv-cli.sh` matches `mcp__.*very-good-cli__.*` +and denies the call outright when `check_vgv_cli` does not return `ok`. In the sandbox it +never does, so every mocked `very-good-cli` call comes back as: + +```text +Very Good CLI is not installed. This tool requires Very Good CLI >= 1.3.0. +``` + +That is the hook's deny message, not the mock. The mock is never reached. `command -v +very_good` actually *succeeds* inside a run, because the sandbox inherits the host PATH; +what fails is the next line, `very_good --version`, which returns nothing under the run's +throwaway `$HOME`. On a machine where `dart` is a version-manager shim it resolves through +`$HOME`, and the CI runner has no `very_good` at all. Either way `check_vgv_cli` reports +`not_installed` and the call is denied. + +**Until that is resolved, the `very-good-cli` mocks are registered but unreachable**, and +the tool-driven cases stay graded on narration. Fixing it means changing what the hook +gates on, which is shipped plugin behavior rather than an eval concern, so it is not +decided here. `very-good-cli` is mocked. `dart` is not, so runs still print `plugin_vgv-ai-flutter-plugin_dart[not started: no mock]`. From 42c21b28e881f47c7e1c6707487ad96a49b553ef Mon Sep 17 00:00:00 2001 From: Dominik Simonik Date: Fri, 18 Sep 2026 15:16:28 +0200 Subject: [PATCH 03/18] fix(evals): only guard mock inputs the real schema requires A driven call aborted the run at score 0: aborted by mock very-good-cli/packages_check_licenses: the model's call violates expect: directory = (missing) is not a string `expect` treats a missing key as a violation, so listing an optional field makes the mock reject correct calls. Only `create` has required fields (`subcommand`, `name`); the other three tools have none, so their guards are gone. Argument choices belong in a grader, not in `expect`, which aborts. Verified against PR #155: with `unverifiable`, the PreToolUse hook stands aside and the model reaches the mock, calling it with {"licenses": true}. Before #155 the hook denied the call and the mock was never reached. Also corrects the tool name in the docs. The bare mcp__very-good-cli__ form works for --allow-tools and allowed_tools, but the name the model actually invokes, and the one a tool_used grader needs, is the plugin-namespaced mcp__plugin_vgv-ai-flutter-plugin_very-good-cli__. Co-Authored-By: Claude Opus 5 --- evals/BASELINE.md | 9 ++-- evals/README.md | 44 +++++++++---------- evals/mocks/very-good-cli/create.md | 2 - .../very-good-cli/packages_check_licenses.md | 6 --- evals/mocks/very-good-cli/packages_get.md | 7 --- evals/mocks/very-good-cli/test.md | 9 ---- 6 files changed, 25 insertions(+), 52 deletions(-) diff --git a/evals/BASELINE.md b/evals/BASELINE.md index 6b9d378..099c331 100644 --- a/evals/BASELINE.md +++ b/evals/BASELINE.md @@ -122,8 +122,7 @@ editing, because a single reading of a case is not a measurement. the switch, so it never explains a routing miss, but every answer after the skill fires ran on Haiku while the no-plugin arm ran on Sonnet. Its Δ is understated by an unmeasured amount. -- **Tool-driven skills are graded on narration.** `very-good-cli` now has mocks, but they - are unreachable: the plugin's own `check-vgv-cli.sh` PreToolUse hook denies every - `mcp__*very-good-cli__*` call inside a run. The skills that drive tools are still - measured on the calls they describe, not the calls they make. See `README.md` → - Mocking the MCP servers. +- **Tool-driven skills are graded on narration.** `very-good-cli` has mocks and they are + reachable, but no case drives one yet, so these skills are still measured on the calls + they describe rather than the calls they make. See `README.md` → Mocking the MCP + servers. diff --git a/evals/README.md b/evals/README.md index 35a043c..e282f35 100644 --- a/evals/README.md +++ b/evals/README.md @@ -223,32 +223,30 @@ not granted (missing --allow-tools grant, or a malformed entry): mcp__very-good-cli__packages_check_licenses ``` -So any case that drives a mocked tool needs both, and the tool name is the **bare** server -form, not the plugin-namespaced one: - ```bash claude plugin eval . --scaffold --allow-tools 'mcp__very-good-cli__packages_check_licenses' ``` -**The plugin's own PreToolUse hook.** `check-vgv-cli.sh` matches `mcp__.*very-good-cli__.*` -and denies the call outright when `check_vgv_cli` does not return `ok`. In the sandbox it -never does, so every mocked `very-good-cli` call comes back as: +**Two different tool names are in play, and mixing them up costs a grader.** The *bare* +`mcp__very-good-cli__` form works for the `--allow-tools` grant and for a case's +`allowed_tools`. The name the model actually invokes, and therefore the name a +`tool_used` grader has to carry, is the plugin-namespaced one: ```text -Very Good CLI is not installed. This tool requires Very Good CLI >= 1.3.0. +mcp__plugin_vgv-ai-flutter-plugin_very-good-cli__packages_check_licenses ``` -That is the hook's deny message, not the mock. The mock is never reached. `command -v -very_good` actually *succeeds* inside a run, because the sandbox inherits the host PATH; -what fails is the next line, `very_good --version`, which returns nothing under the run's -throwaway `$HOME`. On a machine where `dart` is a version-manager shim it resolves through -`$HOME`, and the CI runner has no `very_good` at all. Either way `check_vgv_cli` reports -`not_installed` and the call is denied. +**The plugin's own PreToolUse hook used to eat every call.** `check-vgv-cli.sh` matches +`mcp__.*very-good-cli__.*` and denied outright whenever `check_vgv_cli` did not return +`ok`, so the model received the hook's "Very Good CLI is not installed" text as the tool +result and the mock was never reached. In a run `command -v very_good` *succeeds*, because +the sandbox inherits the host PATH; `very_good --version` then returns nothing under the +run's throwaway `$HOME`, because the installed `very_good` is a shim that execs `dart`. -**Until that is resolved, the `very-good-cli` mocks are registered but unreachable**, and -the tool-driven cases stay graded on narration. Fixing it means changing what the hook -gates on, which is shipped plugin behavior rather than an eval concern, so it is not -decided here. +`check_vgv_cli` now returns `unverifiable` for exactly that case and both hooks stand +aside, so the mocks are reachable and the SessionStart notice does not fire either. +**Mocked cases therefore require a plugin that has the `unverifiable` status**; against an +older build every `very-good-cli` mock is dead. `very-good-cli` is mocked. `dart` is not, so runs still print `plugin_vgv-ai-flutter-plugin_dart[not started: no mock]`. @@ -369,12 +367,12 @@ discriminating goes unnoticed until you re-check with `include_baseline`. a `tracePath` into a sandbox deleted unless `--keep-temp` is passed. `report.html` does show the judged text. - **Judge calibration.** Most graders are `llm` with no human-labelled gold set. -- **Tool execution.** `very-good-cli` is now mocked, so its four tools can be called - inside a run and graded with `tool_used` or against `mock_calls`. No case does yet: every - prompt still says the session cannot run anything, and every `allowed_tools` still lists - only `[Read, Glob, Grep, Skill]`. Until both change, the tool-driven skills stay graded on - the calls they narrate rather than the calls they make. The `dart` server has no mock at - all. +- **Tool execution.** `very-good-cli` is mocked and a mocked call has been driven end to + end, so its four tools can be graded with `tool_used` or against `mock_calls`. No case + does yet: every prompt still says the session cannot run anything, and every + `allowed_tools` still lists only `[Read, Glob, Grep, Skill]`. Until both change, the + tool-driven skills stay graded on the calls they narrate. The `dart` server has no mock + at all. - **Stable routing.** Whether a skill activates is nondeterministic, which is why routing is a `tool_used` grader rather than inferred from content. - **Prose in a `SKILL.md`.** Deliberate. An earlier version asserted a hundred `contains` diff --git a/evals/mocks/very-good-cli/create.md b/evals/mocks/very-good-cli/create.md index 05fd86a..6ac1b3b 100644 --- a/evals/mocks/very-good-cli/create.md +++ b/evals/mocks/very-good-cli/create.md @@ -2,8 +2,6 @@ expect: subcommand: [app_ui_package, flame_game, flutter_app, flutter_package, flutter_plugin, dart_cli, dart_package, docs_site] name: string - org_name: string - output_directory: string --- ✓ Generated 87 file(s) from the {{input.subcommand}} template. diff --git a/evals/mocks/very-good-cli/packages_check_licenses.md b/evals/mocks/very-good-cli/packages_check_licenses.md index 24fa77d..bc1514a 100644 --- a/evals/mocks/very-good-cli/packages_check_licenses.md +++ b/evals/mocks/very-good-cli/packages_check_licenses.md @@ -1,9 +1,3 @@ ---- -expect: - directory: string - licenses: boolean ---- - Retrieved 14 licenses from 14 packages. | package | version | license | diff --git a/evals/mocks/very-good-cli/packages_get.md b/evals/mocks/very-good-cli/packages_get.md index 0f70fe3..05add4f 100644 --- a/evals/mocks/very-good-cli/packages_get.md +++ b/evals/mocks/very-good-cli/packages_get.md @@ -1,9 +1,2 @@ ---- -expect: - directory: string - recursive: boolean - ignore: string ---- - Running "flutter pub get"... ✓ Retrieved dependencies (2.1s) diff --git a/evals/mocks/very-good-cli/test.md b/evals/mocks/very-good-cli/test.md index 1d9e883..64868ed 100644 --- a/evals/mocks/very-good-cli/test.md +++ b/evals/mocks/very-good-cli/test.md @@ -1,12 +1,3 @@ ---- -expect: - directory: string - recursive: boolean - coverage: boolean - min_coverage: string - optimization: boolean ---- - Running "flutter test"... 00:04 +38: All tests passed! From 7bbb6fd556dea52bf2ce61618acd2e8e52833c4c Mon Sep 17 00:00:00 2001 From: Dominik Simonik Date: Fri, 18 Sep 2026 15:46:37 +0200 Subject: [PATCH 04/18] docs(evals): drop the startup-notice workaround from the one prompt that needed it With check_vgv_cli returning `unverifiable`, warn-missing-mcp.sh emits nothing, so the sentence telling the model to ignore a startup notice has nothing left to guard against. Removed from create-project-scopes-dependency-install-to-the-new-project, which measured 3/3 at 1.00 afterwards, including the no-flutter-create grader that the notice originally tripped. The other two prompts named as carrying this workaround do not. They say the session cannot reach a real toolchain, which is true regardless of the hook, so they are unchanged. Co-Authored-By: Claude Opus 5 --- evals/README.md | 10 +++++++--- .../prompt.md | 2 +- 2 files changed, 8 insertions(+), 4 deletions(-) diff --git a/evals/README.md b/evals/README.md index e282f35..ce75aef 100644 --- a/evals/README.md +++ b/evals/README.md @@ -150,9 +150,13 @@ Negative controls use the same grader with `min: 0` and `max: 0`. injects "Very Good CLI is not installed" into every with-plugin run, because the sandbox has no `very_good` on PATH. Tool-driven cases answer with that blocker instead of the question. Say in the prompt that the CLI is installed and that a startup notice saying - otherwise should be ignored. **Mocking the server does not silence it.** `check_vgv_cli` - tests `command -v very_good`, so it reports on the PATH and never on whether the MCP - tools are reachable. A mocked run carries the same warning an unmocked one does. + otherwise should be ignored. **Mocking the server does not silence it**, because + `check_vgv_cli` tests `command -v very_good` rather than whether the MCP tools are + reachable. What does silence it is the `unverifiable` status: the binary resolves but its + version cannot be read, and the hook then emits nothing. `create-project-scopes-dependency-install-to-the-new-project` + has had the workaround sentence removed on that basis, measured 3/3 at 1.00. The other + two prompts keep theirs, which guard against a different thing: the sandbox has no real + toolchain, and no hook change fixes that. Beyond that: write prompts as a user would send them, name no skill in a prompt, grade mechanically where you can, include the cases where the skill must say no, keep a negative diff --git a/evals/create-project/create-project-scopes-dependency-install-to-the-new-project/prompt.md b/evals/create-project/create-project-scopes-dependency-install-to-the-new-project/prompt.md index ed831e2..ccb27c5 100644 --- a/evals/create-project/create-project-scopes-dependency-install-to-the-new-project/prompt.md +++ b/evals/create-project/create-project-scopes-dependency-install-to-the-new-project/prompt.md @@ -6,4 +6,4 @@ tags: [create-project] description: The create-then-install order, with the install scoped to the created project via directory, plus the organization prompt. --- -I want to scaffold a new Flutter app named storefront, placed at apps/storefront inside my existing monorepo, and then install its dependencies. Which tools would you call, in order, and what arguments would you pass to each? Tell me anything you still need from me. Do not run anything yet. Very Good CLI is installed and on PATH, so treat its tools as available and ignore any startup notice saying otherwise. +I want to scaffold a new Flutter app named storefront, placed at apps/storefront inside my existing monorepo, and then install its dependencies. Which tools would you call, in order, and what arguments would you pass to each? Tell me anything you still need from me. Do not run anything yet. Very Good CLI is installed and on PATH, so treat its tools as available. From c16c92cbd5c46401e73719a5d1d88d235992f73a Mon Sep 17 00:00:00 2001 From: Dominik Simonik Date: Fri, 18 Sep 2026 16:11:17 +0200 Subject: [PATCH 05/18] fix(evals): make the six weakest cases discriminate MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit All numbers measured on claude-sonnet-5 for both model and judge. Adds signal rather than cutting it. A grader that passes in both arms earns no delta but still fails if the skill regresses later, so the free ones that pin a Core Standard or an anti-pattern stay. Four graders are deleted, each for a reason that holds without a score: `uses-pump-app` and `names-all-rights-reserved` duplicate a rubric beside them, `declares-weather-repository` restates the prompt, and `uses-test-widgets` is table stakes for any widget test. Two skills were at fault rather than their cases: - skills/bloc/SKILL.md gains a Core Standard for `build:` constructing the bloc under test. That was the only real difference between the two arms of bloc-tests-with-bloc-test-and-mocktail, which moves Δ +0.38 -> +0.75. - bloc-writes-sealed-events-and-states was grading one of two approaches the skill sanctions. The skill documents a Subclass and a Single Class approach and selects on whether states carry different data; the graders only accept the first, so a status-enum answer followed the skill and failed anyway. The prompt now supplies states that carry different data. Readings went 1.00/0.70/0.60 on the unchanged case to 1.00/1.00/1.00 at --runs 3. license-compliance-refuses-to-clear-missing-licenses graded several ways of saying no, which any model says. It now also grades the skill's risk categorization: Δ +0.43 -> +0.86 at --runs 3 on both arms. testing-uses-pump-app-in-widget-tests referred to helpers the fixture does not have, so the no-plugin arm sometimes asked a question instead of answering. Its prompt is now self-contained. Corrects the rule in BASELINE.md that one no-plugin run can disqualify a grader. It cannot: accessibility-declines-gesture-detector-tap-target failed all three content graders without the plugin on one run and passed all three on the next. Co-Authored-By: Claude Opus 5 --- evals/BASELINE.md | 74 ++++++++++++++----- evals/README.md | 45 +++++++++++ .../graders/build-returns-a-fresh-bloc.md | 12 +++ .../prompt.md | 4 +- .../no-match-text-direction-on-icon.md | 1 + .../graders/declares-weather-repository.md | 4 - .../graders/names-all-rights-reserved.md | 5 -- .../graders/rates-both-packages-high-risk.md | 11 +++ .../graders/uses-pump-app.md | 4 - .../graders/uses-test-widgets.md | 4 - .../graders/wraps-with-pump-app.md | 8 +- .../prompt.md | 21 +++++- skills/bloc/SKILL.md | 1 + 13 files changed, 153 insertions(+), 41 deletions(-) create mode 100644 evals/bloc/bloc-tests-with-bloc-test-and-mocktail/graders/build-returns-a-fresh-bloc.md delete mode 100644 evals/layered-architecture/layered-architecture-transforms-models-in-the-repository/graders/declares-weather-repository.md delete mode 100644 evals/license-compliance/license-compliance-refuses-to-clear-missing-licenses/graders/names-all-rights-reserved.md create mode 100644 evals/license-compliance/license-compliance-refuses-to-clear-missing-licenses/graders/rates-both-packages-high-risk.md delete mode 100644 evals/testing/testing-uses-pump-app-in-widget-tests/graders/uses-pump-app.md delete mode 100644 evals/testing/testing-uses-pump-app-in-widget-tests/graders/uses-test-widgets.md diff --git a/evals/BASELINE.md b/evals/BASELINE.md index 099c331..1f63da6 100644 --- a/evals/BASELINE.md +++ b/evals/BASELINE.md @@ -11,9 +11,11 @@ claude plugin eval . --trust-plugin --scaffold \ --no-publish --max-cost-usd 45 -j 4 ``` -`--runs 1` is deliberate for a two-arm sweep: one no-plugin pass is enough to disqualify a -grader, which is what the sweep is for. Anything about a *single* case, whether it -regressed or whether it routes reliably, needs `--runs 3`, because both arms are noisy. +`--runs 1` is deliberate for a two-arm sweep over the whole suite: it is a shortlist, not a +verdict. It is **not** enough to disqualify a grader. Measured 2026-09-18, the no-plugin arm +of `accessibility-declines-gesture-detector-tap-target` failed all three content graders on +one run and passed all three on the next. Anything about a single case or a single grader +needs `--runs 3` on both arms. ## Full two-arm run, 2026-09-17, Claude Code 2.1.270 @@ -34,22 +36,31 @@ The previous sweep, under a Haiku judge and before the refusal-shaped descriptio measured 5 misses out of 83, and the one before that measured 14. There is no "skill did not route" row in the table above because the population is empty. -## `bloc-writes-sealed-events-and-states`, re-measured 2026-09-18 +## `bloc-writes-sealed-events-and-states`, fixed 2026-09-18 -The CI run on 2026-09-18 scored it 0.60, failing `sealed-state-hierarchy`, -`pinned-in-progress-name`, `final-class-subclasses` and `past-tense-event-names`. A -`--runs 3` re-measurement on the same model and judge CI uses scored **1.00, 3/3, every -grader passing on every run**. +The CI run scored it 0.60, failing `sealed-state-hierarchy`, `pinned-in-progress-name`, +`final-class-subclasses` and `past-tense-event-names`. A first `--runs 3` scored it 1.00, +3/3, which looked like variance. It was not: a later run scored 0.70, so the readings were +1.00, 0.70 and 0.60 on an unchanged case. -```bash -claude plugin eval . --scaffold --trust-plugin --ablation none --runs 3 \ - --threshold 0.8 --model claude-sonnet-5 --judge-model claude-sonnet-5 \ - --case bloc-writes-sealed-events-and-states -``` +**The case was grading one of two approaches the skill sanctions.** `skills/bloc/SKILL.md` +documents a Subclass Approach and a Single Class Approach, and says to choose by whether +the states carry different data. The graders only accept the first, so a model that wrote +the status-enum state was following the skill and failing the case anyway. + +The prompt now supplies states that carry different data, which is the skill's own rule for +choosing subclasses: a spinner while in flight, a `User` on success, an error message on +failure. Measured after the change, `--runs 3` on both arms: + +| | runs | mean | +| ------- | --------------------- | ----: | +| with | 1.00, 1.00, 1.00 | 1.00 | +| without | 0.70, 0.40, 0.60 | 0.57 | -Four graders failing at once looked like a real miss and was not. The case, its graders and -`skills/bloc/SKILL.md` are unchanged. Treat it as the worked example of why one CI reading -is not a measurement. +No grader failed in any with-arm run. An intermediate attempt that asked for "the whole +request lifecycle the UI will render" instead made it consistently *worse*, 0.70 three +times out of three, by steering harder toward the enum. The wording has to select the +approach the way the skill selects it, not describe the feature. ## Reading the two numbers that look bad @@ -68,9 +79,34 @@ cases that clear less than Δ 0.50, which is where a free grader actually costs - `license-compliance-refuses-to-clear-missing-licenses` - `testing-uses-pump-app-in-widget-tests` -A grader is only worth keeping if it can fail in the no-plugin arm. Those six cases are -the shortlist for the next grader pass. The remaining 93 free graders sit on cases that -discriminate through their other graders, which is untidy rather than disqualifying. +Those six cases were the shortlist for the grader pass below. + +## Grader pass over the six weakest cases, 2026-09-18 + +A free grader is still a regression test, so this pass mostly **added** signal rather than +cutting. Four graders were deleted, each for a reason that does not depend on a score: two +exact duplicates of a rubric beside them, one that restates the prompt +(`class WeatherRepository`), and one that is table stakes for the format (`testWidgets`). +An earlier draft of this pass deleted fifteen on a single free reading and had to restore +eleven. + +| case | Δ before | Δ after | +| ---------------------------------------------------------- | -------: | ------: | +| `license-compliance-refuses-to-clear-missing-licenses` | +0.43 | +0.86 | +| `internationalization-uses-directional-insets-for-rtl` | +0.50 | +0.86 | +| `bloc-tests-with-bloc-test-and-mocktail` | +0.38 | +0.75 | +| `layered-architecture-transforms-models-in-the-repository` | +0.67 | +0.57 | +| `testing-uses-pump-app-in-widget-tests` | +0.38 | +0.43 | +| `accessibility-declines-gesture-detector-tap-target` | +0.86 | +0.50 | + +`license-compliance` and `internationalization` were measured at `--runs 3` on both arms. +The rest are single-run readings and move several tenths on their own, so read the last two +rows as noise rather than regression: nothing was removed from `accessibility` that its +remaining graders do not still cover. + +Two skills changed because the grader pass found the skill at fault rather than the case. +`skills/bloc/SKILL.md` gained a Core Standard for `build:` constructing the bloc under test, +which is what separated the two arms of `bloc-tests-with-bloc-test-and-mocktail`. ## Per-skill mean Δ diff --git a/evals/README.md b/evals/README.md index ce75aef..042fc74 100644 --- a/evals/README.md +++ b/evals/README.md @@ -163,6 +163,51 @@ mechanically where you can, include the cases where the skill must say no, keep control's rubric to the absence of the skill's vocabulary, and check a grader fails in the without-arm before trusting it. +### Graders that cannot fail + +A grader that passes in the no-plugin arm measures the model, not the skill. Find them by +running the case two-arm and comparing the two `graders` lists. + +**One two-arm run cannot disqualify a grader.** The no-plugin arm swings hard between runs. +`accessibility-declines-gesture-detector-tap-target` had all three content graders *failing* +without the plugin on one run and all three *passing* on the next, with nothing changed in +between. Use `--runs 3` on both arms before calling a grader free, and prefer reasons that +do not depend on a score at all: redundancy, restating the prompt, and what the two arms' +text actually differs on. + +**Adding beats deleting.** A free grader still fails if the skill later regresses, so it is +a regression test even when it earns no Δ. Keep every free grader that pins a Core Standard +or an anti-pattern: `no-mockito`, `no-left-anchored-insets` and `no-hand-rolled-icon-mirroring` +all pass in both arms today and all catch a real regression tomorrow. A grader pass over +these six cases deleted fifteen of them on a single free reading and had to put eleven back. + +Delete only for a reason that holds without a score: + +1. **A genuine duplicate.** `uses-pump-app` matched `pumpApp` beside a rubric judging the + same thing; `names-all-rights-reserved` was the regex of the rubric next to it. +2. **It restates the prompt.** `class WeatherRepository` passes whenever the model read the + question. +3. **It is table stakes for the format.** `testWidgets` appears in every widget test ever + written. + +Everything else gets *added to*, not cut. Read both arms' output side by side, find what +only the plugin produced, grade that, and weight it so the case turns on it. +`license-compliance-refuses-to-clear-missing-licenses` graded several ways of saying "no", +which any model says; what the plugin added was the skill's risk categorization, so that +became a weighted grader beside the ones already there. + +If a case has no discriminating grader even then, the skill may genuinely teach nothing the +model does not already do, and the honest fix is the skill rather than the case. +`bloc-tests-with-bloc-test-and-mocktail` was that: both arms reached for `bloc_test` and +`mocktail` unaided, and the only real difference was that the plugin built a fresh bloc +inside `build:` while the bare model shared one from `setUp`. The skill's examples showed +that and its Core Standards never said it, so the standard was added and the grader now +tests it. That moved the case from Δ +0.38 to +0.75. + +**Check that both arms answered.** A no-plugin arm that asks a clarifying question instead +of doing the work makes every grader look discriminating for one run. That is a prompt that +is not self-contained, not a result. + ### Writing an `llm` rubric Frontmatter is only `type: llm`. The body is the rubric, written as concrete PASS and FAIL diff --git a/evals/bloc/bloc-tests-with-bloc-test-and-mocktail/graders/build-returns-a-fresh-bloc.md b/evals/bloc/bloc-tests-with-bloc-test-and-mocktail/graders/build-returns-a-fresh-bloc.md new file mode 100644 index 0000000..96fb44e --- /dev/null +++ b/evals/bloc/bloc-tests-with-bloc-test-and-mocktail/graders/build-returns-a-fresh-bloc.md @@ -0,0 +1,12 @@ +--- +type: llm +weight: 3 +--- + +Judge only the `build:` callbacks of the `blocTest` calls. + +PASS if every `build:` constructs the bloc inside the callback, for example +`build: () => LoginBloc(authRepository: authRepository)` or `build: LoginBloc.new`. + +FAIL if any `build:` returns a bloc that was constructed outside the callback, for example +one assigned to a variable in `setUp` and referenced as `build: () => loginBloc`. diff --git a/evals/bloc/bloc-writes-sealed-events-and-states/prompt.md b/evals/bloc/bloc-writes-sealed-events-and-states/prompt.md index b759d8d..464c945 100644 --- a/evals/bloc/bloc-writes-sealed-events-and-states/prompt.md +++ b/evals/bloc/bloc-writes-sealed-events-and-states/prompt.md @@ -3,7 +3,7 @@ max_turns: 12 timeout_seconds: 600 allowed_tools: [Read, Glob, Grep, Skill] tags: [bloc] -description: "The house shape of a bloc: sealed event and state hierarchies with `final class` subclasses, Equatable with props, past-tense event names, and the state names SKILL.md pins for a login flow." +description: "The house shape of a bloc: states carrying different data select the subclass approach, so sealed hierarchies with `final class` subclasses, Equatable with props, past-tense event names, and the state names SKILL.md pins for a login flow." --- -Write a LoginBloc for email and password authentication with submit and logout events, and success and failure states. Output Dart code only. +Write a LoginBloc for email and password authentication, with a submit event and a logout event. While the request is in flight the button shows a spinner. On success the UI needs the signed-in User object; on failure it needs the error message to display. Nothing else is on screen before the first attempt. Output Dart code only. diff --git a/evals/internationalization/internationalization-uses-directional-insets-for-rtl/graders/no-match-text-direction-on-icon.md b/evals/internationalization/internationalization-uses-directional-insets-for-rtl/graders/no-match-text-direction-on-icon.md index 87499aa..4bad5e9 100644 --- a/evals/internationalization/internationalization-uses-directional-insets-for-rtl/graders/no-match-text-direction-on-icon.md +++ b/evals/internationalization/internationalization-uses-directional-insets-for-rtl/graders/no-match-text-direction-on-icon.md @@ -2,4 +2,5 @@ type: regex pattern: 'Icon\(\s*Icons\.[A-Za-z0-9_]+,\s*matchTextDirection' match: not_contains +weight: 3 --- diff --git a/evals/layered-architecture/layered-architecture-transforms-models-in-the-repository/graders/declares-weather-repository.md b/evals/layered-architecture/layered-architecture-transforms-models-in-the-repository/graders/declares-weather-repository.md deleted file mode 100644 index 6aaaf9f..0000000 --- a/evals/layered-architecture/layered-architecture-transforms-models-in-the-repository/graders/declares-weather-repository.md +++ /dev/null @@ -1,4 +0,0 @@ ---- -type: regex -pattern: 'class WeatherRepository' ---- diff --git a/evals/license-compliance/license-compliance-refuses-to-clear-missing-licenses/graders/names-all-rights-reserved.md b/evals/license-compliance/license-compliance-refuses-to-clear-missing-licenses/graders/names-all-rights-reserved.md deleted file mode 100644 index f12bf84..0000000 --- a/evals/license-compliance/license-compliance-refuses-to-clear-missing-licenses/graders/names-all-rights-reserved.md +++ /dev/null @@ -1,5 +0,0 @@ ---- -type: regex -pattern: 'all rights reserved' -flags: i ---- diff --git a/evals/license-compliance/license-compliance-refuses-to-clear-missing-licenses/graders/rates-both-packages-high-risk.md b/evals/license-compliance/license-compliance-refuses-to-clear-missing-licenses/graders/rates-both-packages-high-risk.md new file mode 100644 index 0000000..03746b4 --- /dev/null +++ b/evals/license-compliance/license-compliance-refuses-to-clear-missing-licenses/graders/rates-both-packages-high-risk.md @@ -0,0 +1,11 @@ +--- +type: llm +weight: 3 +--- + +PASS if both unlicensed packages are given an explicit risk level and that level is the +highest the response uses, for example "High risk", "High", or "critical", whether in a +table column or in prose. + +FAIL if the response attaches no risk level to them, or rates them anything below its own +top level. diff --git a/evals/testing/testing-uses-pump-app-in-widget-tests/graders/uses-pump-app.md b/evals/testing/testing-uses-pump-app-in-widget-tests/graders/uses-pump-app.md deleted file mode 100644 index 4a94320..0000000 --- a/evals/testing/testing-uses-pump-app-in-widget-tests/graders/uses-pump-app.md +++ /dev/null @@ -1,4 +0,0 @@ ---- -type: regex -pattern: 'pumpApp' ---- diff --git a/evals/testing/testing-uses-pump-app-in-widget-tests/graders/uses-test-widgets.md b/evals/testing/testing-uses-pump-app-in-widget-tests/graders/uses-test-widgets.md deleted file mode 100644 index 9f07bb0..0000000 --- a/evals/testing/testing-uses-pump-app-in-widget-tests/graders/uses-test-widgets.md +++ /dev/null @@ -1,4 +0,0 @@ ---- -type: regex -pattern: 'testWidgets' ---- diff --git a/evals/testing/testing-uses-pump-app-in-widget-tests/graders/wraps-with-pump-app.md b/evals/testing/testing-uses-pump-app-in-widget-tests/graders/wraps-with-pump-app.md index ce78fff..8be4adb 100644 --- a/evals/testing/testing-uses-pump-app-in-widget-tests/graders/wraps-with-pump-app.md +++ b/evals/testing/testing-uses-pump-app-in-widget-tests/graders/wraps-with-pump-app.md @@ -1,7 +1,11 @@ --- type: llm +weight: 3 --- -PASS if the test wraps the widget with a pumpApp helper. +PASS if the test puts the widget on screen through a `pumpApp` helper on the tester, for +example `await tester.pumpApp(const LoginView());`. -FAIL if it calls pumpWidget(MaterialApp(...)) inline in the test body. +FAIL if it calls `pumpWidget` directly in the test body, including +`await tester.pumpWidget(MaterialApp(home: LoginView()));`, or defines its own inline +wrapper instead of using the helper. diff --git a/evals/testing/testing-uses-pump-app-in-widget-tests/prompt.md b/evals/testing/testing-uses-pump-app-in-widget-tests/prompt.md index 61ee1c5..fbb4eb6 100644 --- a/evals/testing/testing-uses-pump-app-in-widget-tests/prompt.md +++ b/evals/testing/testing-uses-pump-app-in-widget-tests/prompt.md @@ -6,4 +6,23 @@ tags: [testing] description: "Widget tests wrap through the shared `pumpApp` helper and drive the view with a MockBloc or MockCubit rather than a real bloc." --- -Write a widget test for a LoginView that shows an error SnackBar when the bloc emits LoginFailure, using our existing widget-test helpers. Output Dart code only. +Write a widget test for `LoginView`. It should assert that a SnackBar carrying the error +message appears when the bloc emits a failure state. Our project already has the usual +shared widget-test helpers available under `test/helpers/`. Output Dart code only. + +These are the types, and this description is authoritative over anything on disk: + +```dart +class LoginView extends StatelessWidget { const LoginView({super.key}); } + +class LoginBloc extends Bloc { ... } + +sealed class LoginState extends Equatable { const LoginState(); } + +final class LoginInitial extends LoginState {} + +final class LoginFailure extends LoginState { + const LoginFailure(this.errorMessage); + final String errorMessage; +} +``` diff --git a/skills/bloc/SKILL.md b/skills/bloc/SKILL.md index ea8c21b..ff45cb3 100644 --- a/skills/bloc/SKILL.md +++ b/skills/bloc/SKILL.md @@ -27,6 +27,7 @@ Apply these standards to ALL Bloc/Cubit work: - **Use `blocTest()` from `package:bloc_test`** for all Bloc and Cubit tests — never raw `test()` with manual stream assertions - **Use `package:mocktail` for mocking** — never `package:mockito` +- **`build:` constructs the bloc under test** — return a new instance from the callback, never one built in `setUp` and shared, so every `blocTest` starts from a clean bloc and its lifecycle stays with `blocTest` - **No bloc-to-bloc direct dependencies** — blocs communicate through the UI or shared repositories - **Page/View separation** — Page provides the Bloc/Cubit via `BlocProvider`, View consumes via `BlocBuilder`/`BlocListener` - **Sealed classes for events and multi-state types** — enables exhaustive pattern matching with Dart 3 `switch` From 3527eea9c64df756f5374f5a7812f2bd93adbe64 Mon Sep 17 00:00:00 2001 From: Dominik Simonik Date: Fri, 18 Sep 2026 16:16:37 +0200 Subject: [PATCH 06/18] docs(evals): scope the mock claims to eval runs "very-good-cli is mocked" read as a property of the plugin. The mocks exist only for `claude plugin eval`: they are not shipped behavior, they do not affect a real session, and a user's very-good-cli tools still reach the real `very_good mcp` server. Co-Authored-By: Claude Opus 5 --- evals/BASELINE.md | 4 ++-- evals/README.md | 11 +++++++---- 2 files changed, 9 insertions(+), 6 deletions(-) diff --git a/evals/BASELINE.md b/evals/BASELINE.md index 1f63da6..0d3e20c 100644 --- a/evals/BASELINE.md +++ b/evals/BASELINE.md @@ -158,7 +158,7 @@ editing, because a single reading of a case is not a measurement. the switch, so it never explains a routing miss, but every answer after the skill fires ran on Haiku while the no-plugin arm ran on Sonnet. Its Δ is understated by an unmeasured amount. -- **Tool-driven skills are graded on narration.** `very-good-cli` has mocks and they are - reachable, but no case drives one yet, so these skills are still measured on the calls +- **Tool-driven skills are graded on narration.** `very-good-cli` has a stand-in inside an + eval run, which changes nothing outside one, and it is reachable, but no case drives it yet, so these skills are still measured on the calls they describe rather than the calls they make. See `README.md` → Mocking the MCP servers. diff --git a/evals/README.md b/evals/README.md index 042fc74..3e45970 100644 --- a/evals/README.md +++ b/evals/README.md @@ -297,7 +297,10 @@ aside, so the mocks are reachable and the SessionStart notice does not fire eith **Mocked cases therefore require a plugin that has the `unverifiable` status**; against an older build every `very-good-cli` mock is dead. -`very-good-cli` is mocked. `dart` is not, so runs still print +These mocks exist only for eval runs. They are not shipped behavior, they do not affect a +real session, and a user's `very-good-cli` tools still go to the real +[`very_good mcp`](https://pub.dev/packages/very_good_cli) server as always. Inside a run, +`very-good-cli` has a stand-in and `dart` does not, so runs still print `plugin_vgv-ai-flutter-plugin_dart[not started: no mock]`. ```text @@ -416,9 +419,9 @@ discriminating goes unnoticed until you re-check with `include_baseline`. a `tracePath` into a sandbox deleted unless `--keep-temp` is passed. `report.html` does show the judged text. - **Judge calibration.** Most graders are `llm` with no human-labelled gold set. -- **Tool execution.** `very-good-cli` is mocked and a mocked call has been driven end to - end, so its four tools can be graded with `tool_used` or against `mock_calls`. No case - does yet: every prompt still says the session cannot run anything, and every +- **Tool execution.** `very-good-cli` has a stand-in *inside an eval run* and a mocked call + has been driven end to end, so its four tools can be graded with `tool_used` or against + `mock_calls`. No case does yet: every prompt still says the session cannot run anything, and every `allowed_tools` still lists only `[Read, Glob, Grep, Skill]`. Until both change, the tool-driven skills stay graded on the calls they narrate. The `dart` server has no mock at all. From 36ee305f1d509d54bd8ed039f12581d0f977692a Mon Sep 17 00:00:00 2001 From: Dominik Simonik Date: Fri, 18 Sep 2026 16:21:19 +0200 Subject: [PATCH 07/18] docs(evals): drop the rest of the harness insulation from the create-project prompt The prompt also told the model the CLI was installed and on PATH so it would treat the tools as available. With the hook no longer denying, that is doing nothing either: removing the whole sentence measured 3/3 at 1.00 on --runs 3. A prompt should describe the user's situation, not work around the harness. Co-Authored-By: Claude Opus 5 --- evals/README.md | 15 ++++++++------- .../prompt.md | 2 +- 2 files changed, 9 insertions(+), 8 deletions(-) diff --git a/evals/README.md b/evals/README.md index 3e45970..fe62d82 100644 --- a/evals/README.md +++ b/evals/README.md @@ -150,13 +150,14 @@ Negative controls use the same grader with `min: 0` and `max: 0`. injects "Very Good CLI is not installed" into every with-plugin run, because the sandbox has no `very_good` on PATH. Tool-driven cases answer with that blocker instead of the question. Say in the prompt that the CLI is installed and that a startup notice saying - otherwise should be ignored. **Mocking the server does not silence it**, because - `check_vgv_cli` tests `command -v very_good` rather than whether the MCP tools are - reachable. What does silence it is the `unverifiable` status: the binary resolves but its - version cannot be read, and the hook then emits nothing. `create-project-scopes-dependency-install-to-the-new-project` - has had the workaround sentence removed on that basis, measured 3/3 at 1.00. The other - two prompts keep theirs, which guard against a different thing: the sandbox has no real - toolchain, and no hook change fixes that. + otherwise should be ignored — **but check whether it is still needed before writing one.** + Mocking the server does not silence the hook, because `check_vgv_cli` tests + `command -v very_good` rather than whether the MCP tools are reachable. The + `unverifiable` status does silence it: the binary resolves, its version cannot be read, + and the hook emits nothing. `create-project-scopes-dependency-install-to-the-new-project` + carried two sentences of insulation against the old behavior and now carries none, + measured 3/3 at 1.00 without them. The other two prompts keep theirs, which guard against + a different thing: the sandbox has no real toolchain, and no hook change fixes that. Beyond that: write prompts as a user would send them, name no skill in a prompt, grade mechanically where you can, include the cases where the skill must say no, keep a negative diff --git a/evals/create-project/create-project-scopes-dependency-install-to-the-new-project/prompt.md b/evals/create-project/create-project-scopes-dependency-install-to-the-new-project/prompt.md index ccb27c5..9f190d2 100644 --- a/evals/create-project/create-project-scopes-dependency-install-to-the-new-project/prompt.md +++ b/evals/create-project/create-project-scopes-dependency-install-to-the-new-project/prompt.md @@ -6,4 +6,4 @@ tags: [create-project] description: The create-then-install order, with the install scoped to the created project via directory, plus the organization prompt. --- -I want to scaffold a new Flutter app named storefront, placed at apps/storefront inside my existing monorepo, and then install its dependencies. Which tools would you call, in order, and what arguments would you pass to each? Tell me anything you still need from me. Do not run anything yet. Very Good CLI is installed and on PATH, so treat its tools as available. +I want to scaffold a new Flutter app named storefront, placed at apps/storefront inside my existing monorepo, and then install its dependencies. Which tools would you call, in order, and what arguments would you pass to each? Tell me anything you still need from me. Do not run anything yet. From 6f2ee876774184f949e49ade801389b08ceed969 Mon Sep 17 00:00:00 2001 From: Dominik Simonik Date: Fri, 18 Sep 2026 16:30:49 +0200 Subject: [PATCH 08/18] docs(evals): correct the mock findings against the published reference Re-read https://code.claude.com/docs/en/plugin-evals and re-tested each claim. The grant claim was wrong. A mocked tool needs no --allow-tools grant and no allowed_tools entry, exactly as the docs say. Verified with allowed_tools: [Read, Glob, Grep, Skill] and no grant: the model still called packages_check_licenses with {"directory": ".", "licenses": true}. The earlier "not granted (missing --allow-tools grant, or a malformed entry)" was the second half of that message. The bare mcp__very-good-cli__ form is not a valid entry in a case; the name is mcp__plugin____. The runner cannot tell a missing grant from a bad name, so it offers both. Two behaviors are genuinely undocumented and are now recorded: - `expect` treats a missing field as a violation, so naming an optional argument aborts a correct call. The reference presents it as a type guard. - `_tools.json` is load-bearing. Renaming an argument in it and changing nothing else made the model call the tool with the renamed argument, which confirms both that it is parsed and that the saved `tools/list` result is the shape it wants. Frontmatter is optional; three of the four mocks have none. Co-Authored-By: Claude Opus 5 --- evals/README.md | 56 ++++++++++++++++++++++++++++++------------------- 1 file changed, 35 insertions(+), 21 deletions(-) diff --git a/evals/README.md b/evals/README.md index fe62d82..d718b4c 100644 --- a/evals/README.md +++ b/evals/README.md @@ -261,31 +261,31 @@ A run never starts the plugin's real MCP servers. `evals/mocks//.m a stand-in under the server's own name from `.mcp.json`. A server with no mock directory is not started at all and its tools are absent, which every run reports on a `mocked:` line. -**Two things block a mocked call, and both bite before any grader runs.** +**A mocked tool needs no grant, and must not be listed in `allowed_tools`.** It is callable +with `allowed_tools: [Read, Glob, Grep, Skill]` and no `--allow-tools` on the command line, +exactly as the docs say. Verified: with neither, the model still called +`packages_check_licenses` with `{"directory": ".", "licenses": true}`. -**The grant.** The published docs say a mocked tool "is allowed without an `--allow-tools` -grant". Measured on Claude Code 2.1.270, it is not. Listing the tool in a case's -`allowed_tools` is necessary and not sufficient, and without the operator grant the run -reports: +**Use the plugin-namespaced tool name everywhere.** The name is +`mcp__plugin____`, so here: ```text -not granted (missing --allow-tools grant, or a malformed entry): -mcp__very-good-cli__packages_check_licenses -``` - -```bash -claude plugin eval . --scaffold --allow-tools 'mcp__very-good-cli__packages_check_licenses' +mcp__plugin_vgv-ai-flutter-plugin_very-good-cli__packages_check_licenses ``` -**Two different tool names are in play, and mixing them up costs a grader.** The *bare* -`mcp__very-good-cli__` form works for the `--allow-tools` grant and for a case's -`allowed_tools`. The name the model actually invokes, and therefore the name a -`tool_used` grader has to carry, is the plugin-namespaced one: +That is the name a `tool_used` grader has to carry. The bare `mcp__very-good-cli__` +form, which the skills' own `allowed-tools` use for a real session, is **not** a valid +entry here: putting it in a case's `allowed_tools` produces a warning that reads as if the +tool needed a grant, when it is really telling you the name does not exist. ```text -mcp__plugin_vgv-ai-flutter-plugin_very-good-cli__packages_check_licenses +not granted (missing --allow-tools grant, or a malformed entry): +mcp__very-good-cli__packages_check_licenses ``` +Both halves of that message are offered because the runner cannot tell them apart. Check +the name before reaching for `--allow-tools`. + **The plugin's own PreToolUse hook used to eat every call.** `check-vgv-cli.sh` matches `mcp__.*very-good-cli__.*` and denied outright whenever `check_vgv_cli` did not return `ok`, so the model received the hook's "Very Good CLI is not installed" text as the tool @@ -313,7 +313,8 @@ evals/mocks/very-good-cli/ └── test.md ``` -A mock file is frontmatter plus a body, and the body is the tool result: +A mock file is an optional frontmatter block plus a body, and the body is the tool result. +Three of the four here have no frontmatter at all, which is the same as `type: fixed`: | Key | Default | Purpose | | ------------ | ------- | --------------------------------------------------------------- | @@ -327,6 +328,17 @@ A mock file is frontmatter plus a body, and the body is the tool result: `_server.md` answers several tools from one `agent` mock; a `.md` for the same tool wins. A case's own `mocks/` directory overrides the suite's file by file. +**`expect` treats a missing field as a violation**, which the reference does not say. It +reads as a type guard, so naming an optional argument looks harmless, and it is not: + +```text +aborted by mock very-good-cli/packages_check_licenses: the model's call violates +expect: directory = (missing) is not a string +``` + +Only `create` has required arguments (`subcommand`, `name`), so only `create.md` carries an +`expect`. Guard what the real schema requires and nothing else. + **`expect` aborts, it does not fail.** A call that violates it ends the run at score 0, reported as `aborted`, with no failing grader to read. Keep it to what the real server's schema already enforces and grade argument *choices* with `tool_used` or with a `regex` @@ -337,10 +349,12 @@ Every mock here is `type: fixed`. An `agent` mock answers through the judge mode costs money, varies run to run, and needs a recording adopted from `results//mock-recordings/` into `.replay/` before CI repeats. -`_tools.json` is the real `tools/list` response, so a mocked tool carries the real -descriptions and input schemas rather than a permissive placeholder. Regenerate it after a -Very Good CLI release by speaking MCP to `very_good mcp` over stdio and saving the -`tools/list` result. A stale one teaches the model a schema the CLI no longer has. +`_tools.json` is the real `tools/list` result, the object with the `tools` array, so a +mocked tool carries the real descriptions and input schemas rather than a permissive +placeholder. It is load-bearing: renaming an argument in it and changing nothing else made +the model call the tool with the renamed argument. Regenerate it after a Very Good CLI +release by speaking MCP to `very_good mcp` over stdio and saving the `tools/list` result. A +stale one teaches the model a schema the CLI no longer has. `packages_check_licenses` returns a deliberately mixed result, one `GPL-3.0` and one `unknown` among twelve permissive licenses, so a case has something real to flag. From ae7ec2b5640e2253febf3c62486c6d3b4c7df6fc Mon Sep 17 00:00:00 2001 From: Dominik Simonik Date: Fri, 18 Sep 2026 16:38:20 +0200 Subject: [PATCH 09/18] docs(evals): note that real servers can be started, and match the reference's wording A run starts no real MCP server *unless asked*, with --allow-real-servers or --mocks off. The unqualified sentence read as if it were impossible. Co-Authored-By: Claude Opus 5 --- evals/README.md | 7 ++++--- 1 file changed, 4 insertions(+), 3 deletions(-) diff --git a/evals/README.md b/evals/README.md index d718b4c..1e1d0b8 100644 --- a/evals/README.md +++ b/evals/README.md @@ -257,8 +257,9 @@ holding the path, and runs then fail at scaffold time. CI runs on Linux and is u ## Mocking the MCP servers -A run never starts the plugin's real MCP servers. `evals/mocks//.md` registers -a stand-in under the server's own name from `.mcp.json`. A server with no mock directory is +A run never starts the plugin's real MCP servers unless you ask, with `--allow-real-servers` +or `--mocks off`. Neither is used here. `evals/mocks//.md` registers a +stand-in under the server's own name from `.mcp.json`. A server with no mock directory is not started at all and its tools are absent, which every run reports on a `mocked:` line. **A mocked tool needs no grant, and must not be listed in `allowed_tools`.** It is callable @@ -306,7 +307,7 @@ real session, and a user's `very-good-cli` tools still go to the real ```text evals/mocks/very-good-cli/ -├── _tools.json # the real tools/list response +├── _tools.json # the real tools/list result ├── create.md ├── packages_check_licenses.md ├── packages_get.md From 2b99785b365889f7c1465d2de83d3465fdb03c2a Mon Sep 17 00:00:00 2001 From: Dominik Simonik Date: Mon, 21 Sep 2026 10:43:25 +0200 Subject: [PATCH 10/18] docs(evals): trim the tool-execution gap to the gap The bullet restated the whole mock setup inside "What this does not cover", which the Mocking the MCP servers section already covers. It now states the limitation and links there. Co-Authored-By: Claude Opus 5 --- evals/README.md | 9 +++------ 1 file changed, 3 insertions(+), 6 deletions(-) diff --git a/evals/README.md b/evals/README.md index 1e1d0b8..505ca26 100644 --- a/evals/README.md +++ b/evals/README.md @@ -435,12 +435,9 @@ discriminating goes unnoticed until you re-check with `include_baseline`. a `tracePath` into a sandbox deleted unless `--keep-temp` is passed. `report.html` does show the judged text. - **Judge calibration.** Most graders are `llm` with no human-labelled gold set. -- **Tool execution.** `very-good-cli` has a stand-in *inside an eval run* and a mocked call - has been driven end to end, so its four tools can be graded with `tool_used` or against - `mock_calls`. No case does yet: every prompt still says the session cannot run anything, and every - `allowed_tools` still lists only `[Read, Glob, Grep, Skill]`. Until both change, the - tool-driven skills stay graded on the calls they narrate. The `dart` server has no mock - at all. +- **Tool execution.** No case drives a tool, so the skills that would are graded on the + calls they narrate. Closing that gap is what + [Mocking the MCP servers](#mocking-the-mcp-servers) is for, and nothing uses it yet. - **Stable routing.** Whether a skill activates is nondeterministic, which is why routing is a `tool_used` grader rather than inferred from content. - **Prose in a `SKILL.md`.** Deliberate. An earlier version asserted a hundred `contains` From c6d3e4adee58d29c5b27b35daf6b660e4dbb65e4 Mon Sep 17 00:00:00 2001 From: Dominik Simonik Date: Mon, 21 Sep 2026 13:38:43 +0200 Subject: [PATCH 11/18] docs(evals): record the full one-arm run and walk back the bloc claim 100 cases against this branch with #155, one arm, pinned to claude-sonnet-5: mean 0.961, 94/100 at threshold, 83 perfect, $13.35, no usage errors. bloc-writes-sealed-events-and-states was called fixed on one --runs 3. It scored 0.70 in the full run with the corrected prompt in place, losing the same three graders as before. Seven post-fix runs now read 1.00 six times and 0.70 once, against one in three before the fix, so it is improved and not cured. Making it airtight means the skill naming which state approach a request lifecycle with differing payloads takes. Of the six cases under threshold, four were re-measured at --runs 3. green-gate-refuses-to-carry-green-forward (0.91) and create-project-asks-for-organization-when-required (0.83) clear or nearly clear on average, ui-package-declines-hand-rolled-button hit the turn cap rather than failing on content, and green-gate-budgets-per-package matches its 2026-09-17 score exactly. layered-architecture-wires-repositories-in-bootstrap is the one real finding at 0.61, failing constructs-in-bootstrap on all three runs. It is untouched by this branch. Co-Authored-By: Claude Opus 5 --- evals/BASELINE.md | 48 ++++++++++++++++++++++++++++++++++++++++++----- 1 file changed, 43 insertions(+), 5 deletions(-) diff --git a/evals/BASELINE.md b/evals/BASELINE.md index 0d3e20c..a772156 100644 --- a/evals/BASELINE.md +++ b/evals/BASELINE.md @@ -36,7 +36,39 @@ The previous sweep, under a Haiku judge and before the refusal-shaped descriptio measured 5 misses out of 83, and the one before that measured 14. There is no "skill did not route" row in the table above because the population is empty. -## `bloc-writes-sealed-events-and-states`, fixed 2026-09-18 +## Full one-arm run, 2026-09-21, this branch with #155 + +100 cases, 1 run each, with-plugin arm only, `partial: false`, **$13.35**, 10 minutes at +`-j 4`, models and judge pinned to `claude-sonnet-5`. + +| | value | +| ------------------------------- | ---------: | +| Mean case score | **0.961** | +| Cases at or above threshold 0.8 | **94/100** | +| Cases scoring a perfect 1.00 | **83/100** | +| Usage or auth errors | **0** | + +Not comparable with the 2026-09-17 sweep above: that one was two-arm, and +`--ablation none` weights the routing graders differently. + +Six cases came in under 0.8. One is a harness error rather than a result, and the four +that were re-measured at `--runs 3` split three ways: + +| case | 1-run | `--runs 3` | reading | +| -------------------------------------------------------- | ----: | ---------: | ----------------------------------------- | +| `ui-package-declines-hand-rolled-button` | 0.43 | — | hit the 12-turn cap; no content verdict | +| `layered-architecture-wires-repositories-in-bootstrap` | 0.50 | **0.61** | real, and consistently short | +| `green-gate-budgets-per-package-across-a-monorepo` | 0.62 | — | matches its 2026-09-17 score exactly | +| `bloc-writes-sealed-events-and-states` | 0.70 | **1.00** | improved, not cured; see below | +| `green-gate-refuses-to-carry-green-forward` | 0.71 | **0.91** | clears threshold on average | +| `create-project-asks-for-organization-when-required` | 0.75 | **0.83** | borderline; `asks-for-organization` flaps | + +`layered-architecture-wires-repositories-in-bootstrap` is the one worth acting on: +`constructs-in-bootstrap` failed all three runs and two runs also lost `path-dependencies` +and `uses-repository-provider`. It is untouched by this branch, so it is either drift since +2026-09-17 or a single-run miss in that sweep. + +## `bloc-writes-sealed-events-and-states`, improved 2026-09-18 The CI run scored it 0.60, failing `sealed-state-hierarchy`, `pinned-in-progress-name`, `final-class-subclasses` and `past-tense-event-names`. A first `--runs 3` scored it 1.00, @@ -57,10 +89,16 @@ failure. Measured after the change, `--runs 3` on both arms: | with | 1.00, 1.00, 1.00 | 1.00 | | without | 0.70, 0.40, 0.60 | 0.57 | -No grader failed in any with-arm run. An intermediate attempt that asked for "the whole -request lifecycle the UI will render" instead made it consistently *worse*, 0.70 three -times out of three, by steering harder toward the enum. The wording has to select the -approach the way the skill selects it, not describe the feature. +An intermediate attempt that asked for "the whole request lifecycle the UI will render" +instead made it consistently *worse*, 0.70 three times out of three, by steering harder +toward the enum. The wording has to select the approach the way the skill selects it, not +describe the feature. + +**Improved rather than cured.** Across seven post-fix runs the case scored 1.00 six times +and 0.70 once, against one in three before, and the 0.70 lost the same three graders as +ever. The prompt biases the model toward the subclass approach; it does not force it. If it +needs to be airtight, the skill has to say which approach a request lifecycle with +differing payloads takes, rather than leaving it to the selection rule. ## Reading the two numbers that look bad From f7c78a486be5c7e18932597544569a67f8840db7 Mon Sep 17 00:00:00 2001 From: Dominik Simonik Date: Mon, 21 Sep 2026 14:05:16 +0200 Subject: [PATCH 12/18] docs(evals): cut the process narrative out of the eval docs Both files had drifted into a record of how the work was done rather than what a future reader needs. BASELINE.md dropped a section headed "this branch with #155", a reference that means nothing once merged, and a six-row triage table that belonged in the pull request. The bloc case's history is now the finding and the caveat rather than a blow-by-blow. The grader pass keeps its before/after numbers and loses the prose around them. README.md loses the forensics on why the hook denied mocked calls, which is fixed rather than something to act on, and the account of an earlier draft of the grader pass. The tool-naming trap is the same warning in a third of the space. BASELINE 202 -> 162 lines, README 450 -> 431. Co-Authored-By: Claude Opus 5 --- evals/BASELINE.md | 88 +++++++++++++---------------------------------- evals/README.md | 47 ++++++++----------------- 2 files changed, 38 insertions(+), 97 deletions(-) diff --git a/evals/BASELINE.md b/evals/BASELINE.md index a772156..36e56f1 100644 --- a/evals/BASELINE.md +++ b/evals/BASELINE.md @@ -36,10 +36,10 @@ The previous sweep, under a Haiku judge and before the refusal-shaped descriptio measured 5 misses out of 83, and the one before that measured 14. There is no "skill did not route" row in the table above because the population is empty. -## Full one-arm run, 2026-09-21, this branch with #155 +## Full one-arm run, 2026-09-21 -100 cases, 1 run each, with-plugin arm only, `partial: false`, **$13.35**, 10 minutes at -`-j 4`, models and judge pinned to `claude-sonnet-5`. +100 cases, 1 run each, with-plugin arm, `partial: false`, **$13.35**, 10 minutes at `-j 4`, +models and judge pinned to `claude-sonnet-5`. | | value | | ------------------------------- | ---------: | @@ -48,57 +48,30 @@ measured 5 misses out of 83, and the one before that measured 14. There is no | Cases scoring a perfect 1.00 | **83/100** | | Usage or auth errors | **0** | -Not comparable with the 2026-09-17 sweep above: that one was two-arm, and -`--ablation none` weights the routing graders differently. +Not comparable with the two-arm sweep above: `--ablation none` weights the routing graders +differently. -Six cases came in under 0.8. One is a harness error rather than a result, and the four -that were re-measured at `--runs 3` split three ways: - -| case | 1-run | `--runs 3` | reading | -| -------------------------------------------------------- | ----: | ---------: | ----------------------------------------- | -| `ui-package-declines-hand-rolled-button` | 0.43 | — | hit the 12-turn cap; no content verdict | -| `layered-architecture-wires-repositories-in-bootstrap` | 0.50 | **0.61** | real, and consistently short | -| `green-gate-budgets-per-package-across-a-monorepo` | 0.62 | — | matches its 2026-09-17 score exactly | -| `bloc-writes-sealed-events-and-states` | 0.70 | **1.00** | improved, not cured; see below | -| `green-gate-refuses-to-carry-green-forward` | 0.71 | **0.91** | clears threshold on average | -| `create-project-asks-for-organization-when-required` | 0.75 | **0.83** | borderline; `asks-for-organization` flaps | - -`layered-architecture-wires-repositories-in-bootstrap` is the one worth acting on: -`constructs-in-bootstrap` failed all three runs and two runs also lost `path-dependencies` -and `uses-repository-provider`. It is untouched by this branch, so it is either drift since -2026-09-17 or a single-run miss in that sweep. +Six cases came in under 0.8. Re-measured at `--runs 3`, only one holds: +**`layered-architecture-wires-repositories-in-bootstrap` at 0.61**, losing +`constructs-in-bootstrap` on all three runs. `green-gate-refuses-to-carry-green-forward` +(0.91) and `create-project-asks-for-organization-when-required` (0.83) clear or nearly +clear, `ui-package-declines-hand-rolled-button` hit the turn cap rather than failing on +content, and `green-gate-budgets-per-package-across-a-monorepo` matches its 2026-09-17 +score. ## `bloc-writes-sealed-events-and-states`, improved 2026-09-18 -The CI run scored it 0.60, failing `sealed-state-hierarchy`, `pinned-in-progress-name`, -`final-class-subclasses` and `past-tense-event-names`. A first `--runs 3` scored it 1.00, -3/3, which looked like variance. It was not: a later run scored 0.70, so the readings were -1.00, 0.70 and 0.60 on an unchanged case. - -**The case was grading one of two approaches the skill sanctions.** `skills/bloc/SKILL.md` -documents a Subclass Approach and a Single Class Approach, and says to choose by whether -the states carry different data. The graders only accept the first, so a model that wrote -the status-enum state was following the skill and failing the case anyway. - -The prompt now supplies states that carry different data, which is the skill's own rule for -choosing subclasses: a spinner while in flight, a `User` on success, an error message on -failure. Measured after the change, `--runs 3` on both arms: +The case was grading one of two approaches the skill sanctions. `skills/bloc/SKILL.md` +documents a Subclass Approach and a Single Class Approach and chooses by whether the states +carry different data; the graders accept only the first, so a status-enum answer followed +the skill and failed anyway. The prompt now supplies states that carry different data, +which is the skill's own rule. -| | runs | mean | -| ------- | --------------------- | ----: | -| with | 1.00, 1.00, 1.00 | 1.00 | -| without | 0.70, 0.40, 0.60 | 0.57 | - -An intermediate attempt that asked for "the whole request lifecycle the UI will render" -instead made it consistently *worse*, 0.70 three times out of three, by steering harder -toward the enum. The wording has to select the approach the way the skill selects it, not -describe the feature. - -**Improved rather than cured.** Across seven post-fix runs the case scored 1.00 six times -and 0.70 once, against one in three before, and the 0.70 lost the same three graders as -ever. The prompt biases the model toward the subclass approach; it does not force it. If it -needs to be airtight, the skill has to say which approach a request lifecycle with -differing payloads takes, rather than leaving it to the selection rule. +Readings went from 1.00/0.70/0.60 on the unchanged case to six 1.00s and one 0.70 across +seven post-fix runs. **Improved, not cured.** Making it airtight means the skill naming +which approach a request lifecycle with differing payloads takes, rather than leaving it to +the selection rule. An attempt to fix it by asking for "the whole request lifecycle the UI +will render" made it consistently worse, 0.70 three times out of three. ## Reading the two numbers that look bad @@ -121,13 +94,6 @@ Those six cases were the shortlist for the grader pass below. ## Grader pass over the six weakest cases, 2026-09-18 -A free grader is still a regression test, so this pass mostly **added** signal rather than -cutting. Four graders were deleted, each for a reason that does not depend on a score: two -exact duplicates of a rubric beside them, one that restates the prompt -(`class WeatherRepository`), and one that is table stakes for the format (`testWidgets`). -An earlier draft of this pass deleted fifteen on a single free reading and had to restore -eleven. - | case | Δ before | Δ after | | ---------------------------------------------------------- | -------: | ------: | | `license-compliance-refuses-to-clear-missing-licenses` | +0.43 | +0.86 | @@ -137,14 +103,8 @@ eleven. | `testing-uses-pump-app-in-widget-tests` | +0.38 | +0.43 | | `accessibility-declines-gesture-detector-tap-target` | +0.86 | +0.50 | -`license-compliance` and `internationalization` were measured at `--runs 3` on both arms. -The rest are single-run readings and move several tenths on their own, so read the last two -rows as noise rather than regression: nothing was removed from `accessibility` that its -remaining graders do not still cover. - -Two skills changed because the grader pass found the skill at fault rather than the case. -`skills/bloc/SKILL.md` gained a Core Standard for `build:` constructing the bloc under test, -which is what separated the two arms of `bloc-tests-with-bloc-test-and-mocktail`. +The first two were measured at `--runs 3` on both arms; the rest are single-run and move +several tenths on their own, so the last two rows are noise rather than regression. ## Per-skill mean Δ diff --git a/evals/README.md b/evals/README.md index 505ca26..36d1f0f 100644 --- a/evals/README.md +++ b/evals/README.md @@ -179,8 +179,7 @@ text actually differs on. **Adding beats deleting.** A free grader still fails if the skill later regresses, so it is a regression test even when it earns no Δ. Keep every free grader that pins a Core Standard or an anti-pattern: `no-mockito`, `no-left-anchored-insets` and `no-hand-rolled-icon-mirroring` -all pass in both arms today and all catch a real regression tomorrow. A grader pass over -these six cases deleted fifteen of them on a single free reading and had to put eleven back. +all pass in both arms today and all catch a real regression tomorrow. Delete only for a reason that holds without a score: @@ -267,37 +266,19 @@ with `allowed_tools: [Read, Glob, Grep, Skill]` and no `--allow-tools` on the co exactly as the docs say. Verified: with neither, the model still called `packages_check_licenses` with `{"directory": ".", "licenses": true}`. -**Use the plugin-namespaced tool name everywhere.** The name is -`mcp__plugin____`, so here: - -```text -mcp__plugin_vgv-ai-flutter-plugin_very-good-cli__packages_check_licenses -``` - -That is the name a `tool_used` grader has to carry. The bare `mcp__very-good-cli__` -form, which the skills' own `allowed-tools` use for a real session, is **not** a valid -entry here: putting it in a case's `allowed_tools` produces a warning that reads as if the -tool needed a grant, when it is really telling you the name does not exist. - -```text -not granted (missing --allow-tools grant, or a malformed entry): -mcp__very-good-cli__packages_check_licenses -``` - -Both halves of that message are offered because the runner cannot tell them apart. Check -the name before reaching for `--allow-tools`. - -**The plugin's own PreToolUse hook used to eat every call.** `check-vgv-cli.sh` matches -`mcp__.*very-good-cli__.*` and denied outright whenever `check_vgv_cli` did not return -`ok`, so the model received the hook's "Very Good CLI is not installed" text as the tool -result and the mock was never reached. In a run `command -v very_good` *succeeds*, because -the sandbox inherits the host PATH; `very_good --version` then returns nothing under the -run's throwaway `$HOME`, because the installed `very_good` is a shim that execs `dart`. - -`check_vgv_cli` now returns `unverifiable` for exactly that case and both hooks stand -aside, so the mocks are reachable and the SessionStart notice does not fire either. -**Mocked cases therefore require a plugin that has the `unverifiable` status**; against an -older build every `very-good-cli` mock is dead. +**Use the plugin-namespaced tool name**, `mcp__plugin____`, which +here is `mcp__plugin_vgv-ai-flutter-plugin_very-good-cli__packages_check_licenses`. That is +what a `tool_used` grader carries. The bare `mcp__very-good-cli__` form the skills +use for a real session is not valid in a case: put it in `allowed_tools` and you get +`not granted (missing --allow-tools grant, or a malformed entry)`, which is the runner +saying the name is unknown rather than that a grant is needed. + +**The mocks need `check_vgv_cli` to return `unverifiable`.** In a run `very_good` resolves +on PATH but `very_good --version` answers nothing, because it is a shim that execs `dart` +and `$HOME` is throwaway. Without the `unverifiable` status the PreToolUse hook reads that +as "not installed" and denies the call, so the model gets the hook's text as the tool +result and the mock is never reached. That status also keeps the SessionStart notice +quiet. These mocks exist only for eval runs. They are not shipped behavior, they do not affect a real session, and a user's `very-good-cli` tools still go to the real From d0822ac3c2483a4c8139695b521fa9cab3e2d513 Mon Sep 17 00:00:00 2001 From: Dominik Simonik Date: Mon, 21 Sep 2026 14:09:22 +0200 Subject: [PATCH 13/18] docs(evals): rewrite both eval docs as reference, not narrative MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Both had turned into an account of how the work was done. They are now what a reader needs to act, and nothing else. BASELINE.md is numbers: two run tables, cases below threshold, per-skill Δ, the grader pass, and caveats as one line each. The case histories and the essays around the tables are gone. 202 -> 111 lines, about what it was before this branch despite carrying an extra run and a new table. README.md keeps every fact that changes what you would do and drops the justification around them. The mocking section is the layout, the frontmatter keys, and four non-obvious behaviors as bullets. Graders that cannot fail is the rule and three reasons to delete. 450 -> 354 lines. No facts removed: tool naming, the expect-aborts-on-missing-field trap, _tools.json, arm: both, weight: 3, mock_calls, the -printf portability note and the measurements are all still there. Co-Authored-By: Claude Opus 5 --- evals/BASELINE.md | 149 ++++++++--------------- evals/README.md | 295 +++++++++++++++++----------------------------- 2 files changed, 158 insertions(+), 286 deletions(-) diff --git a/evals/BASELINE.md b/evals/BASELINE.md index 36e56f1..4f76481 100644 --- a/evals/BASELINE.md +++ b/evals/BASELINE.md @@ -1,8 +1,7 @@ # Measured baseline -Observed numbers from `claude plugin eval`, not estimates. Re-measure after any change to -a skill description, a skill body, a prompt, or the fixture, and date new numbers rather -than editing these. +Observed numbers from `claude plugin eval`. Re-measure after any change to a skill +description, a skill body, a prompt, or the fixture. Date new numbers, do not edit these. ```bash claude plugin eval . --trust-plugin --scaffold \ @@ -11,15 +10,12 @@ claude plugin eval . --trust-plugin --scaffold \ --no-publish --max-cost-usd 45 -j 4 ``` -`--runs 1` is deliberate for a two-arm sweep over the whole suite: it is a shortlist, not a -verdict. It is **not** enough to disqualify a grader. Measured 2026-09-18, the no-plugin arm -of `accessibility-declines-gesture-detector-tap-target` failed all three content graders on -one run and passed all three on the next. Anything about a single case or a single grader -needs `--runs 3` on both arms. +`--runs 1` over the whole suite is a shortlist, not a verdict. Confirm anything about a +single case or a single grader with `--runs 3` on both arms. -## Full two-arm run, 2026-09-17, Claude Code 2.1.270 +## Two-arm, 2026-09-17 -100 cases, 200 agent runs, 0 run errors, `partial: false`, $24.93, 30 minutes at `-j 4`. +100 cases, 200 runs, 0 errors, $24.93, 30 minutes at `-j 4`. | | value | | ------------------------------- | ---------: | @@ -31,83 +27,38 @@ needs `--runs 3` on both arms. | Positive cases with Δ <= 0 | **0/84** | | Suite score | **0.971** | -**Routing is solved for now.** Every one of the 84 positive cases activated its skill. -The previous sweep, under a Haiku judge and before the refusal-shaped description fixes, -measured 5 misses out of 83, and the one before that measured 14. There is no -"skill did not route" row in the table above because the population is empty. +All 15 cases with Δ <= 0 are negative controls, which pass for free without the plugin. +Exclude them when comparing arms. -## Full one-arm run, 2026-09-21 +## One-arm, 2026-09-21 -100 cases, 1 run each, with-plugin arm, `partial: false`, **$13.35**, 10 minutes at `-j 4`, -models and judge pinned to `claude-sonnet-5`. +100 cases, 1 run each, 0 errors, $13.35, 10 minutes at `-j 4`. | | value | | ------------------------------- | ---------: | | Mean case score | **0.961** | | Cases at or above threshold 0.8 | **94/100** | | Cases scoring a perfect 1.00 | **83/100** | -| Usage or auth errors | **0** | -Not comparable with the two-arm sweep above: `--ablation none` weights the routing graders -differently. +Not comparable with the two-arm run: `--ablation none` weights routing graders differently. -Six cases came in under 0.8. Re-measured at `--runs 3`, only one holds: -**`layered-architecture-wires-repositories-in-bootstrap` at 0.61**, losing -`constructs-in-bootstrap` on all three runs. `green-gate-refuses-to-carry-green-forward` -(0.91) and `create-project-asks-for-organization-when-required` (0.83) clear or nearly -clear, `ui-package-declines-hand-rolled-button` hit the turn cap rather than failing on -content, and `green-gate-budgets-per-package-across-a-monorepo` matches its 2026-09-17 -score. - -## `bloc-writes-sealed-events-and-states`, improved 2026-09-18 - -The case was grading one of two approaches the skill sanctions. `skills/bloc/SKILL.md` -documents a Subclass Approach and a Single Class Approach and chooses by whether the states -carry different data; the graders accept only the first, so a status-enum answer followed -the skill and failed anyway. The prompt now supplies states that carry different data, -which is the skill's own rule. - -Readings went from 1.00/0.70/0.60 on the unchanged case to six 1.00s and one 0.70 across -seven post-fix runs. **Improved, not cured.** Making it airtight means the skill naming -which approach a request lifecycle with differing payloads takes, rather than leaving it to -the selection rule. An attempt to fix it by asking for "the whole request lifecycle the UI -will render" made it consistently worse, 0.70 three times out of three. - -## Reading the two numbers that look bad - -**15 cases show Δ <= 0.** All 15 are negative controls, and that is the design: a model -with no plugin passes a "must not invoke the skill" check for free, so its without-arm -score is already near perfect and there is no lift to measure. Among the 84 positive -cases, nothing scored Δ <= 0. Exclude negative controls whenever you compare arms. - -**121 content graders pass with no plugin loaded.** 28 of those sit on the six positive -cases that clear less than Δ 0.50, which is where a free grader actually costs signal: - -- `accessibility-declines-gesture-detector-tap-target` -- `bloc-tests-with-bloc-test-and-mocktail` -- `internationalization-uses-directional-insets-for-rtl` -- `layered-architecture-transforms-models-in-the-repository` -- `license-compliance-refuses-to-clear-missing-licenses` -- `testing-uses-pump-app-in-widget-tests` +## Cases below threshold -Those six cases were the shortlist for the grader pass below. +Six from the 2026-09-21 run, four re-measured at `--runs 3`. -## Grader pass over the six weakest cases, 2026-09-18 - -| case | Δ before | Δ after | -| ---------------------------------------------------------- | -------: | ------: | -| `license-compliance-refuses-to-clear-missing-licenses` | +0.43 | +0.86 | -| `internationalization-uses-directional-insets-for-rtl` | +0.50 | +0.86 | -| `bloc-tests-with-bloc-test-and-mocktail` | +0.38 | +0.75 | -| `layered-architecture-transforms-models-in-the-repository` | +0.67 | +0.57 | -| `testing-uses-pump-app-in-widget-tests` | +0.38 | +0.43 | -| `accessibility-declines-gesture-detector-tap-target` | +0.86 | +0.50 | - -The first two were measured at `--runs 3` on both arms; the rest are single-run and move -several tenths on their own, so the last two rows are noise rather than regression. +| case | 1 run | `--runs 3` | reading | +| ------------------------------------------------------ | ----: | ---------: | ------------------- | +| `ui-package-declines-hand-rolled-button` | 0.43 | — | hit the turn cap | +| `layered-architecture-wires-repositories-in-bootstrap` | 0.50 | **0.61** | real, act on this | +| `green-gate-budgets-per-package-across-a-monorepo` | 0.62 | — | matches 2026-09-17 | +| `bloc-writes-sealed-events-and-states` | 0.70 | **1.00** | improved, not cured | +| `green-gate-refuses-to-carry-green-forward` | 0.71 | **0.91** | clears on average | +| `create-project-asks-for-organization-when-required` | 0.75 | **0.83** | borderline | ## Per-skill mean Δ +Two-arm run, 2026-09-17. + | skill | mean Δ | | -------------------------- | -----: | | ui-package | +0.79 | @@ -126,37 +77,35 @@ several tenths on their own, so the last two rows are noise rather than regressi | static-security | +0.56 | | create-project | +0.56 | -The spread is narrow and every skill clears +0.55, so no skill is currently carrying the -suite or dragging on it. `create-project` sits last partly because its `SKILL.md` pins -`model: haiku`, so its with-arm answers on a weaker model than its baseline. See the -caveats. +Every skill clears +0.55, so none is carrying or dragging the suite. -## Cases below threshold +## Grader pass, 2026-09-18 + +The 2026-09-17 run found 121 content graders passing with no plugin loaded, 28 of them on +these six cases. -| case | score | Δ | -| ----------------------------------------------------- | ----: | ----: | -| `static-security-refuses-platform-channel-biometrics` | 0.50 | +0.50 | -| `green-gate-budgets-per-package-across-a-monorepo` | 0.62 | +0.62 | -| `create-project-does-not-over-ask` | 0.75 | +0.50 | +| case | Δ before | Δ after | +| ---------------------------------------------------------- | -------: | ------: | +| `license-compliance-refuses-to-clear-missing-licenses` | +0.43 | +0.86 | +| `internationalization-uses-directional-insets-for-rtl` | +0.50 | +0.86 | +| `bloc-tests-with-bloc-test-and-mocktail` | +0.38 | +0.75 | +| `layered-architecture-transforms-models-in-the-repository` | +0.67 | +0.57 | +| `testing-uses-pump-app-in-widget-tests` | +0.38 | +0.43 | +| `accessibility-declines-gesture-detector-tap-target` | +0.86 | +0.50 | -All three routed, and all three carry a healthy Δ, so the skill is firing and helping. -They are graded strictly rather than broken. Confirm any of them with `--runs 3` before -editing, because a single reading of a case is not a measurement. +The first two are `--runs 3` on both arms. The rest are single-run, so the last two rows +are noise rather than regression. ## Caveats -- **Single run per arm.** Treat any one case's Δ as a shortlist entry, not a verdict. -- **Sonnet judges.** The suite moved off the Haiku judge after it marked two correct - answers wrong on a 100-case run, both clean passes under Sonnet. Sonnet is not - infallible either: read the judge's votes in the report before editing a skill. -- **Not comparable with anything before 2026-09-17.** Earlier numbers in this file's - history were measured under a Haiku judge and before the harness-insulation fixes to - three prompts, so they understate the suite by an unmeasured amount. -- **`create-project` pins `model: haiku`** in its `SKILL.md`. Routing is decided before - the switch, so it never explains a routing miss, but every answer after the skill fires - ran on Haiku while the no-plugin arm ran on Sonnet. Its Δ is understated by an - unmeasured amount. -- **Tool-driven skills are graded on narration.** `very-good-cli` has a stand-in inside an - eval run, which changes nothing outside one, and it is reachable, but no case drives it yet, so these skills are still measured on the calls - they describe rather than the calls they make. See `README.md` → Mocking the MCP - servers. +- **The no-plugin arm swings between runs.** One two-arm run cannot disqualify a grader. + `accessibility-declines-gesture-detector-tap-target` failed all three content graders + without the plugin on one run and passed all three on the next. +- **Sonnet judges.** Read the judge's votes in the report before editing a skill. +- **Nothing before 2026-09-17 is comparable.** Earlier numbers ran under a Haiku judge. +- **`create-project` pins `model: haiku`.** Its with-arm answers on a weaker model than its + baseline, so its Δ is understated. Routing is decided before the switch. +- **`bloc-writes-sealed-events-and-states` still flaps.** Six 1.00s and one 0.70 over seven + runs, against one in three before the prompt fix. Curing it means `skills/bloc/SKILL.md` + naming which state approach applies. +- **Tool-driven skills are graded on narration.** No case drives a mocked tool yet. diff --git a/evals/README.md b/evals/README.md index 36d1f0f..31673fa 100644 --- a/evals/README.md +++ b/evals/README.md @@ -9,8 +9,7 @@ claude plugin eval . --scaffold --tag bloc # one skill claude plugin eval . --scaffold --ablation none # with-plugin arm only, half the cost ``` -- Claude Code **>= 2.1.269**. Earlier builds answer `plugin eval is currently in early - access`. +- Claude Code **>= 2.1.269**. - No API key locally. Runs authenticate the same way your normal session does. - `--scaffold` is **not optional**. Without it the fixture never runs and cases fail for unrelated reasons. @@ -24,13 +23,11 @@ claude plugin eval . --scaffold --ablation none # with-plugin arm only, ha | with | The plugin loaded | What the plugin produces | | without | No plugin loaded at all | What the bare model produces | -`Δ` is the with-arm score minus the without-arm score. A grader that passes in both arms -is measuring the model, not the skill. +`Δ` is the with-arm score minus the without-arm score. A grader that passes in both arms is +measuring the model, not the skill. -Each run gets a throwaway home directory, workspace, and Claude Code configuration. Your -settings, `CLAUDE.md`, MCP servers, other plugins, memory, and skills are all absent, and -the eval directory is unreadable from inside a run, so a case cannot read its own graders -or its siblings. +Each run gets a throwaway home directory, workspace, and Claude Code configuration, and the +eval directory is unreadable from inside a run. Exclude negative controls when you compare arms. A model with no plugin passes a "must not invoke the skill" check for free. @@ -62,12 +59,11 @@ tags: [bloc] description: Sealed hierarchies, Equatable, and the pinned state names for a LoginBloc. --- -Write a LoginBloc for email and password authentication with submit and logout events, -and success and failure states. Output Dart code only. +Write a LoginBloc for email and password authentication with submit and logout events. +Output Dart code only. ``` -Every case repeats that frontmatter verbatim. `claude plugin eval` has no shared-defaults -mechanism, so it is forced rather than duplication worth removing. +Every case repeats that frontmatter verbatim; there is no shared-defaults mechanism. `case.yaml`: @@ -101,7 +97,7 @@ threshold is `1.0`, so **pass `--threshold 0.8` or every imperfect case exits 1* | `llm` | `criteria`, `focus` | Frontmatter is just `type: llm`; the file body is the rubric | | `baseline` | `baseline_file`, `criteria` | Unused here | -There are **no custom-code graders**. A check that needs to execute something has no home. +There are **no custom-code graders**. A grader reads `last_message` unless its `target` says otherwise. The other targets are `trace`, `files`, `{ source: file, path: }`, and `mock_calls`, which is every call @@ -109,9 +105,7 @@ made to a [mocked MCP tool](#mocking-the-mcp-servers). ### Routing graders -Ninety-nine of the hundred cases carry one. -`create-project-infers-dart-package-for-api-client` is the exception, and routing is -covered by that skill's other six cases. The shape is always: +Ninety-nine of the hundred cases carry one: ```markdown --- @@ -124,14 +118,13 @@ arm: both ``` **`arm: both` is mandatory.** Without it a two-arm run drops every `tool_used: Skill` -grader from the score and reports it as an indicator, so a case whose skill never fires -still scores `1.00`. The cost is that the without-arm is then penalized for not having the -plugin, which inflates `Δ`. Read the two arm scores rather than `Δ` alone. +grader from the score, so a case whose skill never fires still scores `1.00`. It also +penalizes the without-arm, which inflates `Δ`, so read the two arm scores rather than `Δ`. -**`weight: 3` is mandatory** and nothing applies it for you. It is what sinks a case on a -routing miss alone. It does not protect the content graders: against `n` content graders -of weight 1, a single content miss scores `(n + 2) / (n + 3)`, which clears 0.8 for every -`n >= 2`. When one grader carries a case's whole point, weight that grader too. +**`weight: 3` is mandatory** and nothing applies it for you. It sinks a case on a routing +miss alone. It does not protect the content graders: against `n` content graders of weight +1, a single content miss scores `(n + 2) / (n + 3)`, which clears 0.8 for every `n >= 2`. +When one grader carries a case's whole point, weight that grader too. Negative controls use the same grader with `min: 0` and `max: 0`. @@ -146,72 +139,46 @@ Negative controls use the same grader with `min: 0` and `max: 0`. the conventions the case exists to measure. - **Prompts must be self-contained.** Paste in any class a prompt refers to, and name pasted text as authoritative when it describes state not on disk. -- **The plugin's own SessionStart hook fires inside the run.** `warn-missing-mcp.sh` - injects "Very Good CLI is not installed" into every with-plugin run, because the sandbox - has no `very_good` on PATH. Tool-driven cases answer with that blocker instead of the - question. Say in the prompt that the CLI is installed and that a startup notice saying - otherwise should be ignored — **but check whether it is still needed before writing one.** - Mocking the server does not silence the hook, because `check_vgv_cli` tests - `command -v very_good` rather than whether the MCP tools are reachable. The - `unverifiable` status does silence it: the binary resolves, its version cannot be read, - and the hook emits nothing. `create-project-scopes-dependency-install-to-the-new-project` - carried two sentences of insulation against the old behavior and now carries none, - measured 3/3 at 1.00 without them. The other two prompts keep theirs, which guard against - a different thing: the sandbox has no real toolchain, and no hook change fixes that. +- **Check both arms answered.** A no-plugin arm that asks a clarifying question instead of + doing the work makes every grader look discriminating. That is a prompt that is not + self-contained, not a result. +- **The plugin's SessionStart hook can fire inside a run.** `warn-missing-mcp.sh` injects + "Very Good CLI is not installed" whenever `check_vgv_cli` returns `not_installed`, and a + tool-driven case then answers with that blocker instead of the question. The + `unverifiable` status keeps it quiet. Two prompts still say the session cannot reach a + real toolchain, which is a different problem and true regardless. Beyond that: write prompts as a user would send them, name no skill in a prompt, grade -mechanically where you can, include the cases where the skill must say no, keep a negative -control's rubric to the absence of the skill's vocabulary, and check a grader fails in the -without-arm before trusting it. +mechanically where you can, include the cases where the skill must say no, and keep a +negative control's rubric to the absence of the skill's vocabulary. ### Graders that cannot fail A grader that passes in the no-plugin arm measures the model, not the skill. Find them by running the case two-arm and comparing the two `graders` lists. -**One two-arm run cannot disqualify a grader.** The no-plugin arm swings hard between runs. -`accessibility-declines-gesture-detector-tap-target` had all three content graders *failing* -without the plugin on one run and all three *passing* on the next, with nothing changed in -between. Use `--runs 3` on both arms before calling a grader free, and prefer reasons that -do not depend on a score at all: redundancy, restating the prompt, and what the two arms' -text actually differs on. +**One two-arm run cannot disqualify a grader.** Use `--runs 3` on both arms, and prefer +reasons that do not depend on a score at all. **Adding beats deleting.** A free grader still fails if the skill later regresses, so it is a regression test even when it earns no Δ. Keep every free grader that pins a Core Standard -or an anti-pattern: `no-mockito`, `no-left-anchored-insets` and `no-hand-rolled-icon-mirroring` -all pass in both arms today and all catch a real regression tomorrow. +or an anti-pattern. Delete only for a reason that holds without a score: -Delete only for a reason that holds without a score: - -1. **A genuine duplicate.** `uses-pump-app` matched `pumpApp` beside a rubric judging the - same thing; `names-all-rights-reserved` was the regex of the rubric next to it. +1. **A genuine duplicate** of a grader beside it. 2. **It restates the prompt.** `class WeatherRepository` passes whenever the model read the question. -3. **It is table stakes for the format.** `testWidgets` appears in every widget test ever - written. - -Everything else gets *added to*, not cut. Read both arms' output side by side, find what -only the plugin produced, grade that, and weight it so the case turns on it. -`license-compliance-refuses-to-clear-missing-licenses` graded several ways of saying "no", -which any model says; what the plugin added was the skill's risk categorization, so that -became a weighted grader beside the ones already there. - -If a case has no discriminating grader even then, the skill may genuinely teach nothing the -model does not already do, and the honest fix is the skill rather than the case. -`bloc-tests-with-bloc-test-and-mocktail` was that: both arms reached for `bloc_test` and -`mocktail` unaided, and the only real difference was that the plugin built a fresh bloc -inside `build:` while the bare model shared one from `setUp`. The skill's examples showed -that and its Core Standards never said it, so the standard was added and the grader now -tests it. That moved the case from Δ +0.38 to +0.75. - -**Check that both arms answered.** A no-plugin arm that asks a clarifying question instead -of doing the work makes every grader look discriminating for one run. That is a prompt that -is not self-contained, not a result. +3. **It is table stakes for the format**, such as `testWidgets` in a widget test. + +Everything else gets added to. Read both arms' output side by side, find what only the +plugin produced, grade that, and weight it so the case turns on it. + +If nothing discriminates even then, the skill may teach nothing the model does not already +do, and the fix is the skill rather than the case. ### Writing an `llm` rubric -Frontmatter is only `type: llm`. The body is the rubric, written as concrete PASS and FAIL -conditions with an example of each where the wording is open to reading: +Frontmatter is only `type: llm`. The body is the rubric, as concrete PASS and FAIL +conditions with an example of each: ```markdown --- @@ -224,9 +191,8 @@ verb, for example LoginSubmitted. FAIL if any event name is imperative, such as SubmitLogin. ``` -Match the skill's own vocabulary. A rubric asking for "rounds" against a skill that teaches -"iteration count" fails on correct answers. Keep `llm` graders for short output; for -anything long a `regex` reads the whole thing the same way every time. +Match the skill's own vocabulary. Keep `llm` graders for short output; for anything long a +`regex` reads the whole thing the same way every time. --- @@ -235,56 +201,28 @@ anything long a `regex` reads the whole thing the same way every time. Every run starts in an empty workspace. `_fixture/fixture.sh` recreates the neutral Flutter skeleton, hooked up through `context.scaffold_script`, and runs **only with `--scaffold`**. -`context.scaffold_script` will not take a path that leaves the case directory: - -```text -path "../../_fixture/fixture.sh" escapes the case directory (`..` or an absolute path) -— it must name something inside it -``` - -It does resolve a symlink inside the case directory, so each case's `fixture.sh` is a -symlink to the one script. A new case needs its own: +`context.scaffold_script` will not take a path that leaves the case directory, but it does +resolve a symlink inside it, so each case's `fixture.sh` is a symlink to the one script. A +new case needs its own: ```bash ln -s ../../_fixture/fixture.sh evals///fixture.sh ``` **On Windows**, a checkout without `core.symlinks=true` turns each link into a text file -holding the path, and runs then fail at scaffold time. CI runs on Linux and is unaffected. +and runs fail at scaffold time. CI runs on Linux and is unaffected. --- ## Mocking the MCP servers -A run never starts the plugin's real MCP servers unless you ask, with `--allow-real-servers` -or `--mocks off`. Neither is used here. `evals/mocks//.md` registers a -stand-in under the server's own name from `.mcp.json`. A server with no mock directory is -not started at all and its tools are absent, which every run reports on a `mocked:` line. - -**A mocked tool needs no grant, and must not be listed in `allowed_tools`.** It is callable -with `allowed_tools: [Read, Glob, Grep, Skill]` and no `--allow-tools` on the command line, -exactly as the docs say. Verified: with neither, the model still called -`packages_check_licenses` with `{"directory": ".", "licenses": true}`. - -**Use the plugin-namespaced tool name**, `mcp__plugin____`, which -here is `mcp__plugin_vgv-ai-flutter-plugin_very-good-cli__packages_check_licenses`. That is -what a `tool_used` grader carries. The bare `mcp__very-good-cli__` form the skills -use for a real session is not valid in a case: put it in `allowed_tools` and you get -`not granted (missing --allow-tools grant, or a malformed entry)`, which is the runner -saying the name is unknown rather than that a grant is needed. - -**The mocks need `check_vgv_cli` to return `unverifiable`.** In a run `very_good` resolves -on PATH but `very_good --version` answers nothing, because it is a shim that execs `dart` -and `$HOME` is throwaway. Without the `unverifiable` status the PreToolUse hook reads that -as "not installed" and denies the call, so the model gets the hook's text as the tool -result and the mock is never reached. That status also keeps the SessionStart notice -quiet. - -These mocks exist only for eval runs. They are not shipped behavior, they do not affect a -real session, and a user's `very-good-cli` tools still go to the real -[`very_good mcp`](https://pub.dev/packages/very_good_cli) server as always. Inside a run, -`very-good-cli` has a stand-in and `dart` does not, so runs still print -`plugin_vgv-ai-flutter-plugin_dart[not started: no mock]`. +`evals/mocks//.md` registers a stand-in under the server's own name from +`.mcp.json`. A server with no mock directory is not started and its tools are absent, which +every run reports on a `mocked:` line. Real servers start only with `--allow-real-servers` +or `--mocks off`, neither of which is used here. + +This applies to eval runs only. A real session still reaches the real `very_good mcp` +server. ```text evals/mocks/very-good-cli/ @@ -295,8 +233,8 @@ evals/mocks/very-good-cli/ └── test.md ``` -A mock file is an optional frontmatter block plus a body, and the body is the tool result. -Three of the four here have no frontmatter at all, which is the same as `type: fixed`: +Frontmatter is optional and the body is the tool result. Three of the four here have no +frontmatter, which is the same as `type: fixed`. | Key | Default | Purpose | | ------------ | ------- | --------------------------------------------------------------- | @@ -310,39 +248,35 @@ Three of the four here have no frontmatter at all, which is the same as `type: f `_server.md` answers several tools from one `agent` mock; a `.md` for the same tool wins. A case's own `mocks/` directory overrides the suite's file by file. -**`expect` treats a missing field as a violation**, which the reference does not say. It -reads as a type guard, so naming an optional argument looks harmless, and it is not: - -```text -aborted by mock very-good-cli/packages_check_licenses: the model's call violates -expect: directory = (missing) is not a string -``` - -Only `create` has required arguments (`subcommand`, `name`), so only `create.md` carries an -`expect`. Guard what the real schema requires and nothing else. - -**`expect` aborts, it does not fail.** A call that violates it ends the run at score 0, -reported as `aborted`, with no failing grader to read. Keep it to what the real server's -schema already enforces and grade argument *choices* with `tool_used` or with a `regex` -against the `mock_calls` target, which carries every mocked call, its input, and the -answer. +Four things that are not obvious: + +- **A mocked tool needs no `--allow-tools` grant and no `allowed_tools` entry.** Name it in + `allowed_tools` and you get `not granted (missing --allow-tools grant, or a malformed + entry)`, which means the name is unknown, not that a grant is missing. +- **Use the plugin-namespaced name** in a `tool_used` grader: + `mcp__plugin_vgv-ai-flutter-plugin_very-good-cli__`. The bare + `mcp__very-good-cli__` form the skills use for a real session is not valid here. +- **`expect` treats a missing field as a violation**, and a violation aborts the run at + score 0 with no failing grader to read. Guard only what the real schema requires: here + that is `create`, so only `create.md` carries an `expect`. Grade argument *choices* with + `tool_used` or a `regex` against `mock_calls`. +- **`_tools.json` drives the schema the model sees.** It is the `tools/list` result, the + object with the `tools` array. Regenerate it after a Very Good CLI release by speaking + MCP to `very_good mcp` over stdio; a stale one teaches a schema the CLI no longer has. + +The mocks are reachable only when `check_vgv_cli` returns `unverifiable`. In a run +`very_good` resolves on PATH but `very_good --version` answers nothing, because it is a +shim that execs `dart` under a throwaway `$HOME`. Read as `not_installed`, the PreToolUse +hook denies the call and the model gets the hook's text as the tool result. Every mock here is `type: fixed`. An `agent` mock answers through the judge model, so it costs money, varies run to run, and needs a recording adopted from `results//mock-recordings/` into `.replay/` before CI repeats. -`_tools.json` is the real `tools/list` result, the object with the `tools` array, so a -mocked tool carries the real descriptions and input schemas rather than a permissive -placeholder. It is load-bearing: renaming an argument in it and changing nothing else made -the model call the tool with the renamed argument. Regenerate it after a Very Good CLI -release by speaking MCP to `very_good mcp` over stdio and saving the `tools/list` result. A -stale one teaches the model a schema the CLI no longer has. - -`packages_check_licenses` returns a deliberately mixed result, one `GPL-3.0` and one -`unknown` among twelve permissive licenses, so a case has something real to flag. - -Keep the bodies as raw tool output. A mock that already names what a skill teaches hands -the answer to the no-plugin arm, exactly as a non-neutral fixture does. +Keep mock bodies as raw tool output. A mock that names what a skill teaches hands the +answer to the no-plugin arm, exactly as a non-neutral fixture does. +`packages_check_licenses` returns one `GPL-3.0` and one `unknown` among twelve permissive +licenses, so a case has something real to flag. --- @@ -358,29 +292,25 @@ $E --runs 3 --threshold 0.8 # is a red case real? ``` Read the two arm scores, not the total. The without-arm is supposed to score badly. -[BASELINE.md](BASELINE.md) records the last full two-arm measurement. +[BASELINE.md](BASELINE.md) records the measurements. **One run is not a measurement**, and these are not a merge gate. -- `--runs 3` before believing a red case. Cases drift between 2/3 and 3/3 on their own. +- `--runs 3` before believing a red case. - A case that hit its turn cap or timed out scores 0 with **no failing grader**, which reads exactly like a content failure. Check the `NOTES` column, or `cases[].arms.with[].error` in the JSON. - A usage limit hit mid-suite makes every later run fail the same way without marking the document `partial`. Check the errors. -- The suite judges with `claude-sonnet-5`, not the default small model, which marked two - correct answers wrong on a 100-case run. Sonnet is not infallible either: when a case - routes but scores badly, read the judge's votes in the report before editing the skill. +- The suite judges with `claude-sonnet-5`. When a case routes but scores badly, read the + judge's votes in the report before editing the skill. -Measured on full runs: **$13.14** for 100 cases in one arm, **$24.93** for both arms, at -roughly **$0.12 per run**. Per-case cost varies several-fold. At `--runs 3` a two-arm sweep -is six runs per case, so budget around **$75**. `--ablation none` halves it, -`--max-cost-usd` bounds it, and `-j` up to 8 shortens wall clock. +Costs: **$13** for 100 cases in one arm, **$25** for both, roughly **$0.12 per run**. At +`--runs 3` a two-arm sweep is six runs per case, so budget around **$75**. `-j` up to 8 +shortens wall clock. Every run writes `results//` with `aggregate-result.json` and a self-contained -`report.html`. The report is where you find out *why* a run scored low: failed graders are -expanded, and an `llm` grader shows the judge's votes and the text it judged. `results/` is -gitignored. +`report.html`, which is where you find out *why* a run scored low. `results/` is gitignored. --- @@ -388,44 +318,37 @@ gitignored. `.github/workflows/evals.yaml` runs after a merge to `main`, never on a pull request, scoped by `--tag` to the changed skills, with-plugin arm only, and `continue-on-error`. -A regression is therefore reported after it lands, and a case that has stopped -discriminating goes unnoticed until you re-check with `include_baseline`. - -- Changing `_fixture/` or `mocks/` widens the scope to all 15 skills. Both are shared - inputs, so a change to either can move any case. -- The scope job's `find` writes `-exec dirname {} \;` rather than the shorter - `-printf '%h\n'`. `-printf` is a GNU extension that BSD `find` does not have, so the - short form passes in CI and fails for anyone running the same job on macOS. -- A directory under `evals/` only becomes a `--tag` if it actually holds `*/case.yaml`. - `mocks/` and `results/` sit there without being skills. -- CI needs `ANTHROPIC_API_KEY`, having no Claude Code session, and `--trust-plugin`, - because a run with no terminal cannot answer the trust prompt. -- `--ablation none` is the only mode where a `tool_used: Skill` grader is scored by - default. This suite sets `arm: both` so routing scores in either mode, but the two modes - weight the baseline differently. Compare runs from one mode at a time. -- The job has a one-hour ceiling. 100 cases in one arm at `--runs 1` measured roughly 35 - minutes at `-j 4`. A full two-arm run at `--runs 3` is 600 runs and does not fit. + +- Changing `_fixture/` or `mocks/` widens the scope to all 15 skills. +- A directory under `evals/` only becomes a `--tag` if it holds `*/case.yaml`. `mocks/` and + `results/` sit there without being skills. +- The scope job's `find` uses `-exec dirname {} \;` rather than `-printf '%h\n'`, which is + GNU-only and fails on macOS. +- CI needs `ANTHROPIC_API_KEY`, having no Claude Code session, and `--trust-plugin`. +- `--ablation none` and `--ablation with-without` weight the baseline differently. Compare + runs from one mode at a time. +- The job has a one-hour ceiling. 100 cases in one arm measured roughly 35 minutes at + `-j 4`. A two-arm run at `--runs 3` is 600 runs and does not fit. --- ## What this does not cover -- **Dart syntax.** The previous harness parsed every fenced `dart` block through a custom - JavaScript assertion. Native evals have no custom-code graders, so 30 cases across 11 - skills lost that check. The response text is not in `aggregate-result.json` either, only - a `tracePath` into a sandbox deleted unless `--keep-temp` is passed. `report.html` does - show the judged text. +- **Dart syntax.** Native evals have no custom-code graders, so 30 cases across 11 skills + lost the fenced-block parse the previous harness did. The response text is not in + `aggregate-result.json`, only a `tracePath` into a sandbox deleted unless `--keep-temp` + is passed. `report.html` does show the judged text. - **Judge calibration.** Most graders are `llm` with no human-labelled gold set. - **Tool execution.** No case drives a tool, so the skills that would are graded on the - calls they narrate. Closing that gap is what - [Mocking the MCP servers](#mocking-the-mcp-servers) is for, and nothing uses it yet. + calls they narrate. [Mocking the MCP servers](#mocking-the-mcp-servers) is the way to + close that, and nothing uses it yet. - **Stable routing.** Whether a skill activates is nondeterministic, which is why routing is a `tool_used` grader rather than inferred from content. -- **Prose in a `SKILL.md`.** Deliberate. An earlier version asserted a hundred `contains` +- **Prose in a `SKILL.md`.** Deliberate: an earlier version asserted a hundred `contains` patterns against skill bodies, so a copy-edit failed the gate. -- **The skills' own surfaces.** `skills_lint` catches a link pointing at a missing file and - a `name` that does not match its directory. Nothing catches a link resolving to the - *wrong* file, or an `allowed-tools` name that does not exist. Four invariants are - convention alone: `create-project` must not declare `Bash`, `green-gate` must declare it - for parsing `coverage/lcov.info`, `.mcp.json` must keep `--enable dart_format`, and +- **The skills' own surfaces.** `skills_lint` catches a link to a missing file and a `name` + that does not match its directory. Nothing catches a link resolving to the *wrong* file, + or an `allowed-tools` name that does not exist. Four invariants are convention alone: + `create-project` must not declare `Bash`, `green-gate` must declare it for parsing + `coverage/lcov.info`, `.mcp.json` must keep `--enable dart_format`, and `flutter-reviewer` must declare no write tools. From c89dca1c25959de7f3ebd0d0e2f9fa31714deee1 Mon Sep 17 00:00:00 2001 From: Dominik Simonik Date: Mon, 21 Sep 2026 14:46:24 +0200 Subject: [PATCH 14/18] fix(evals): delete BASELINE.md and fix the case it flagged at 0.61 MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit BASELINE.md recorded numbers nobody re-runs after a change, from single runs noisy enough that reading them as a verdict is a mistake this branch made twice. Nothing diffed against it and git already records measurements with the context that makes them readable. Its three durable caveats move to README.md: the no-plugin arm's swing between runs, the Sonnet judge, and create-project's haiku pin. AGENTS.md drops it from the structure block. The one finding it carried is fixed here. layered-architecture-wires-repositories-in-bootstrap asked to wire AuthRepository and UserRepository without saying they were local packages, and the fixture has no packages/ directory, so the model either invented the whole monorepo and passed or treated them as classes in lib/ and lost constructs-in-bootstrap, path-dependencies and uses-repository-provider. Naming the packages alone made it worse: the model checked the tree, found nothing, and asked which it was instead of answering. The prompt now also names the description authoritative over disk, the convention two other prompts already use. At --runs 3 on both arms: with 1.00/1.00/1.00, without 0.33, Δ +0.67, up from 0.61 with three graders failing. Co-Authored-By: Claude Opus 5 --- AGENTS.md | 1 - evals/BASELINE.md | 111 ------------------ evals/README.md | 7 +- .../prompt.md | 2 +- 4 files changed, 7 insertions(+), 114 deletions(-) delete mode 100644 evals/BASELINE.md diff --git a/AGENTS.md b/AGENTS.md index 860d2e5..2afbd9f 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -17,7 +17,6 @@ docs/ # Gitignored, local only plan/ # Planning and design documents evals/ # `claude plugin eval` suite — all 15 skills, 100 cases README.md # Case format, grader reference, how to add a case - BASELINE.md # Last full two-arm measurement: per-skill Δ and routing _fixture/ fixture.sh # The only copy of the neutral Flutter skeleton; every case symlinks here / # One directory per skill, named for the skill it covers diff --git a/evals/BASELINE.md b/evals/BASELINE.md deleted file mode 100644 index 4f76481..0000000 --- a/evals/BASELINE.md +++ /dev/null @@ -1,111 +0,0 @@ -# Measured baseline - -Observed numbers from `claude plugin eval`. Re-measure after any change to a skill -description, a skill body, a prompt, or the fixture. Date new numbers, do not edit these. - -```bash -claude plugin eval . --trust-plugin --scaffold \ - --ablation with-without --runs 1 --threshold 0.8 \ - --model claude-sonnet-5 --judge-model claude-sonnet-5 \ - --no-publish --max-cost-usd 45 -j 4 -``` - -`--runs 1` over the whole suite is a shortlist, not a verdict. Confirm anything about a -single case or a single grader with `--runs 3` on both arms. - -## Two-arm, 2026-09-17 - -100 cases, 200 runs, 0 errors, $24.93, 30 minutes at `-j 4`. - -| | value | -| ------------------------------- | ---------: | -| Mean Δ, all 100 cases | **+0.65** | -| Mean Δ where the skill routed | **+0.76** | -| Routing misses (positive cases) | **0/84** | -| Cases at or above threshold 0.8 | **97/100** | -| Cases scoring a perfect 1.00 | **85/100** | -| Positive cases with Δ <= 0 | **0/84** | -| Suite score | **0.971** | - -All 15 cases with Δ <= 0 are negative controls, which pass for free without the plugin. -Exclude them when comparing arms. - -## One-arm, 2026-09-21 - -100 cases, 1 run each, 0 errors, $13.35, 10 minutes at `-j 4`. - -| | value | -| ------------------------------- | ---------: | -| Mean case score | **0.961** | -| Cases at or above threshold 0.8 | **94/100** | -| Cases scoring a perfect 1.00 | **83/100** | - -Not comparable with the two-arm run: `--ablation none` weights routing graders differently. - -## Cases below threshold - -Six from the 2026-09-21 run, four re-measured at `--runs 3`. - -| case | 1 run | `--runs 3` | reading | -| ------------------------------------------------------ | ----: | ---------: | ------------------- | -| `ui-package-declines-hand-rolled-button` | 0.43 | — | hit the turn cap | -| `layered-architecture-wires-repositories-in-bootstrap` | 0.50 | **0.61** | real, act on this | -| `green-gate-budgets-per-package-across-a-monorepo` | 0.62 | — | matches 2026-09-17 | -| `bloc-writes-sealed-events-and-states` | 0.70 | **1.00** | improved, not cured | -| `green-gate-refuses-to-carry-green-forward` | 0.71 | **0.91** | clears on average | -| `create-project-asks-for-organization-when-required` | 0.75 | **0.83** | borderline | - -## Per-skill mean Δ - -Two-arm run, 2026-09-17. - -| skill | mean Δ | -| -------------------------- | -----: | -| ui-package | +0.79 | -| material-theming | +0.73 | -| very-good-analysis-upgrade | +0.71 | -| animations | +0.69 | -| dart-flutter-sdk-upgrade | +0.69 | -| testing | +0.66 | -| green-gate | +0.65 | -| internationalization | +0.63 | -| layered-architecture | +0.63 | -| bloc | +0.63 | -| license-compliance | +0.61 | -| navigation | +0.58 | -| accessibility | +0.58 | -| static-security | +0.56 | -| create-project | +0.56 | - -Every skill clears +0.55, so none is carrying or dragging the suite. - -## Grader pass, 2026-09-18 - -The 2026-09-17 run found 121 content graders passing with no plugin loaded, 28 of them on -these six cases. - -| case | Δ before | Δ after | -| ---------------------------------------------------------- | -------: | ------: | -| `license-compliance-refuses-to-clear-missing-licenses` | +0.43 | +0.86 | -| `internationalization-uses-directional-insets-for-rtl` | +0.50 | +0.86 | -| `bloc-tests-with-bloc-test-and-mocktail` | +0.38 | +0.75 | -| `layered-architecture-transforms-models-in-the-repository` | +0.67 | +0.57 | -| `testing-uses-pump-app-in-widget-tests` | +0.38 | +0.43 | -| `accessibility-declines-gesture-detector-tap-target` | +0.86 | +0.50 | - -The first two are `--runs 3` on both arms. The rest are single-run, so the last two rows -are noise rather than regression. - -## Caveats - -- **The no-plugin arm swings between runs.** One two-arm run cannot disqualify a grader. - `accessibility-declines-gesture-detector-tap-target` failed all three content graders - without the plugin on one run and passed all three on the next. -- **Sonnet judges.** Read the judge's votes in the report before editing a skill. -- **Nothing before 2026-09-17 is comparable.** Earlier numbers ran under a Haiku judge. -- **`create-project` pins `model: haiku`.** Its with-arm answers on a weaker model than its - baseline, so its Δ is understated. Routing is decided before the switch. -- **`bloc-writes-sealed-events-and-states` still flaps.** Six 1.00s and one 0.70 over seven - runs, against one in three before the prompt fix. Curing it means `skills/bloc/SKILL.md` - naming which state approach applies. -- **Tool-driven skills are graded on narration.** No case drives a mocked tool yet. diff --git a/evals/README.md b/evals/README.md index 31673fa..7683b09 100644 --- a/evals/README.md +++ b/evals/README.md @@ -292,7 +292,6 @@ $E --runs 3 --threshold 0.8 # is a red case real? ``` Read the two arm scores, not the total. The without-arm is supposed to score badly. -[BASELINE.md](BASELINE.md) records the measurements. **One run is not a measurement**, and these are not a merge gate. @@ -304,6 +303,12 @@ Read the two arm scores, not the total. The without-arm is supposed to score bad document `partial`. Check the errors. - The suite judges with `claude-sonnet-5`. When a case routes but scores badly, read the judge's votes in the report before editing the skill. +- **The no-plugin arm swings between runs.** One two-arm run cannot disqualify a grader: + `accessibility-declines-gesture-detector-tap-target` failed all three content graders + without the plugin on one run and passed all three on the next. +- **`create-project` pins `model: haiku`.** Its with-arm answers on a weaker model than its + baseline, so its Δ reads low. Routing is decided before the switch, so the pin never + explains a routing miss. Costs: **$13** for 100 cases in one arm, **$25** for both, roughly **$0.12 per run**. At `--runs 3` a two-arm sweep is six runs per case, so budget around **$75**. `-j` up to 8 diff --git a/evals/layered-architecture/layered-architecture-wires-repositories-in-bootstrap/prompt.md b/evals/layered-architecture/layered-architecture-wires-repositories-in-bootstrap/prompt.md index 8fb0700..70c39cf 100644 --- a/evals/layered-architecture/layered-architecture-wires-repositories-in-bootstrap/prompt.md +++ b/evals/layered-architecture/layered-architecture-wires-repositories-in-bootstrap/prompt.md @@ -6,4 +6,4 @@ tags: [layered-architecture] description: "Bootstrap wiring: clients and repositories constructed in a main entrypoint, passed into App, provided to the tree, and depended on by path rather than by version." --- -Wire my AuthRepository and UserRepository into the app so every feature can reach them. Show the Dart code and the pubspec changes. +My monorepo has two local packages that nothing is wired to yet: `packages/auth_repository`, exposing `AuthRepository`, and `packages/user_repository`, exposing `UserRepository`. They are not in your working directory, so answer from this description rather than inspecting what is on disk. Wire them into the app so every feature can reach them. Show the Dart code and the pubspec changes. From 53bf3ffe07a16fa50868378d49ee08924afab033 Mon Sep 17 00:00:00 2001 From: Dominik Simonik Date: Mon, 21 Sep 2026 15:15:13 +0200 Subject: [PATCH 15/18] feat(evals): grade license-compliance on the call it makes, not the prose MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The mocks had no consumer, so nothing exercised them in CI and their value was unproven. license-compliance-runs-check-with-full-license-info now drives one instead of describing it. The prompt told the model it could not run anything and asked what it would call. It now asks for the audit. Three graders that read prose are replaced: a regex for the tool's name and one for `licenses: true` become tool_used on the call and on its input, and a rubric about what the model said it would retrieve becomes a rubric on whether the two packages the mock flags, a GPL-3.0 and one with no license, reach the report with a risk level. At --runs 3 on both arms: 1.00/1.00/1.00 with the plugin, 0.00 without. That Δ is not content lift. A mocked tool is absent in the no-plugin arm and nothing excludes these graders from the score the way it does for tool_used: Skill, so the without arm fails them for free. README.md now says so, and that the remaining tool-driven skills are still graded on narration. Co-Authored-By: Claude Opus 5 --- evals/README.md | 13 ++++++++++--- .../graders/asks-for-full-license-info.md | 5 +++++ .../graders/calls-the-license-check.md | 5 +++++ .../graders/describes-the-compliance-report.md | 8 ++++++-- .../graders/names-check-licenses.md | 4 ---- .../graders/reports-the-flagged-packages.md | 12 ++++++++++++ .../graders/requests-full-license-info.md | 4 ---- .../graders/retrieves-every-dependency-license.md | 7 ------- .../prompt.md | 4 ++-- 9 files changed, 40 insertions(+), 22 deletions(-) create mode 100644 evals/license-compliance/license-compliance-runs-check-with-full-license-info/graders/asks-for-full-license-info.md create mode 100644 evals/license-compliance/license-compliance-runs-check-with-full-license-info/graders/calls-the-license-check.md delete mode 100644 evals/license-compliance/license-compliance-runs-check-with-full-license-info/graders/names-check-licenses.md create mode 100644 evals/license-compliance/license-compliance-runs-check-with-full-license-info/graders/reports-the-flagged-packages.md delete mode 100644 evals/license-compliance/license-compliance-runs-check-with-full-license-info/graders/requests-full-license-info.md delete mode 100644 evals/license-compliance/license-compliance-runs-check-with-full-license-info/graders/retrieves-every-dependency-license.md diff --git a/evals/README.md b/evals/README.md index 7683b09..62ca976 100644 --- a/evals/README.md +++ b/evals/README.md @@ -278,6 +278,13 @@ answer to the no-plugin arm, exactly as a non-neutral fixture does. `packages_check_licenses` returns one `GPL-3.0` and one `unknown` among twelve permissive licenses, so a case has something real to flag. +**A mocked tool is not there in the no-plugin arm**, so a `tool_used` grader on one fails +for free and takes any grader that needs the tool's output with it. Unlike +`tool_used: Skill`, nothing excludes these from the score, so Δ reads as if the plugin +supplied the content when it mostly supplied the tool. Read the with-arm score. +`license-compliance-runs-check-with-full-license-info` is the worked example: 1.00 with the +plugin and 0.00 without, on three runs each. + --- ## Running @@ -344,9 +351,9 @@ scoped by `--tag` to the changed skills, with-plugin arm only, and `continue-on- `aggregate-result.json`, only a `tracePath` into a sandbox deleted unless `--keep-temp` is passed. `report.html` does show the judged text. - **Judge calibration.** Most graders are `llm` with no human-labelled gold set. -- **Tool execution.** No case drives a tool, so the skills that would are graded on the - calls they narrate. [Mocking the MCP servers](#mocking-the-mcp-servers) is the way to - close that, and nothing uses it yet. +- **Tool execution, mostly.** One case drives a mocked tool, + `license-compliance-runs-check-with-full-license-info`. The other tool-driven skills are + still graded on the calls they narrate. - **Stable routing.** Whether a skill activates is nondeterministic, which is why routing is a `tool_used` grader rather than inferred from content. - **Prose in a `SKILL.md`.** Deliberate: an earlier version asserted a hundred `contains` diff --git a/evals/license-compliance/license-compliance-runs-check-with-full-license-info/graders/asks-for-full-license-info.md b/evals/license-compliance/license-compliance-runs-check-with-full-license-info/graders/asks-for-full-license-info.md new file mode 100644 index 0000000..1720779 --- /dev/null +++ b/evals/license-compliance/license-compliance-runs-check-with-full-license-info/graders/asks-for-full-license-info.md @@ -0,0 +1,5 @@ +--- +type: tool_used +tool: mcp__plugin_vgv-ai-flutter-plugin_very-good-cli__packages_check_licenses +input_match: '"licenses"\s*:\s*true' +--- diff --git a/evals/license-compliance/license-compliance-runs-check-with-full-license-info/graders/calls-the-license-check.md b/evals/license-compliance/license-compliance-runs-check-with-full-license-info/graders/calls-the-license-check.md new file mode 100644 index 0000000..0581340 --- /dev/null +++ b/evals/license-compliance/license-compliance-runs-check-with-full-license-info/graders/calls-the-license-check.md @@ -0,0 +1,5 @@ +--- +type: tool_used +tool: mcp__plugin_vgv-ai-flutter-plugin_very-good-cli__packages_check_licenses +weight: 3 +--- diff --git a/evals/license-compliance/license-compliance-runs-check-with-full-license-info/graders/describes-the-compliance-report.md b/evals/license-compliance/license-compliance-runs-check-with-full-license-info/graders/describes-the-compliance-report.md index e3c52b2..90ab7a1 100644 --- a/evals/license-compliance/license-compliance-runs-check-with-full-license-info/graders/describes-the-compliance-report.md +++ b/evals/license-compliance/license-compliance-runs-check-with-full-license-info/graders/describes-the-compliance-report.md @@ -2,6 +2,10 @@ type: llm --- -PASS if the response describes the written report the audit will produce, and that description has both of these: flagged dependencies held in their own section or table, separate from the compliant ones, and a risk level attached to every flagged dependency (a `Risk` column, or wording such as high/medium/low risk per flagged package). A blank template, a skeleton with placeholder rows, or a plain statement of what the report will contain all count. The response is describing a report it has not run yet, so it does not have to name real packages or fill in real risk levels. +PASS if the report keeps flagged dependencies in their own section or table, separate from +the compliant ones, and attaches a risk level to every flagged dependency, whether as a +`Risk` column or as wording such as high or medium risk per package. -FAIL only if one of those two is missing: the response describes no report at all (for example it promises only to "list the licenses" or to return a pass/fail verdict), the report it describes keeps flagged and compliant dependencies in one undifferentiated list, or the report it describes attaches no risk level to flagged dependencies. +FAIL if flagged and compliant dependencies sit in one undifferentiated list, if the answer +is only a pass/fail verdict or a count, or if no risk level is attached to the flagged +dependencies. diff --git a/evals/license-compliance/license-compliance-runs-check-with-full-license-info/graders/names-check-licenses.md b/evals/license-compliance/license-compliance-runs-check-with-full-license-info/graders/names-check-licenses.md deleted file mode 100644 index 8dc219d..0000000 --- a/evals/license-compliance/license-compliance-runs-check-with-full-license-info/graders/names-check-licenses.md +++ /dev/null @@ -1,4 +0,0 @@ ---- -type: regex -pattern: 'packages[ _]check[ _]licenses' ---- diff --git a/evals/license-compliance/license-compliance-runs-check-with-full-license-info/graders/reports-the-flagged-packages.md b/evals/license-compliance/license-compliance-runs-check-with-full-license-info/graders/reports-the-flagged-packages.md new file mode 100644 index 0000000..9e23be2 --- /dev/null +++ b/evals/license-compliance/license-compliance-runs-check-with-full-license-info/graders/reports-the-flagged-packages.md @@ -0,0 +1,12 @@ +--- +type: llm +--- + +The scan flagged two dependencies: `glyph_atlas` under GPL-3.0 and `refresh_indicator_x` +with no license detected. + +PASS if the report names both of them as needing attention and attaches a risk level to +each, whether in a `Risk` column or in prose. + +FAIL if either is missing, if either is presented as compliant, or if no risk level is +attached to them. diff --git a/evals/license-compliance/license-compliance-runs-check-with-full-license-info/graders/requests-full-license-info.md b/evals/license-compliance/license-compliance-runs-check-with-full-license-info/graders/requests-full-license-info.md deleted file mode 100644 index 808e8ee..0000000 --- a/evals/license-compliance/license-compliance-runs-check-with-full-license-info/graders/requests-full-license-info.md +++ /dev/null @@ -1,4 +0,0 @@ ---- -type: regex -pattern: 'licenses[^\n]{0,24}\btrue\b' ---- diff --git a/evals/license-compliance/license-compliance-runs-check-with-full-license-info/graders/retrieves-every-dependency-license.md b/evals/license-compliance/license-compliance-runs-check-with-full-license-info/graders/retrieves-every-dependency-license.md deleted file mode 100644 index f813fc0..0000000 --- a/evals/license-compliance/license-compliance-runs-check-with-full-license-info/graders/retrieves-every-dependency-license.md +++ /dev/null @@ -1,7 +0,0 @@ ---- -type: llm ---- - -PASS if the response says it will retrieve the full license information for every dependency. - -FAIL if it offers only a pass/fail verdict or a count of packages. diff --git a/evals/license-compliance/license-compliance-runs-check-with-full-license-info/prompt.md b/evals/license-compliance/license-compliance-runs-check-with-full-license-info/prompt.md index 2815e9d..345c35e 100644 --- a/evals/license-compliance/license-compliance-runs-check-with-full-license-info/prompt.md +++ b/evals/license-compliance/license-compliance-runs-check-with-full-license-info/prompt.md @@ -3,7 +3,7 @@ max_turns: 12 timeout_seconds: 600 allowed_tools: [Read, Glob, Grep, Skill] tags: [license-compliance] -description: The audit is the license check with full license information, and the deliverable is the prescribed risk-rated report. +description: Runs the license check with full license information and turns the scan output into the prescribed risk-rated report. --- -We ship this Flutter app next week and I need a license audit of its dependencies first. You can't run anything in this session, so tell me exactly what you would call, with what arguments, and what the finished audit will contain. +We ship this Flutter app next week and I need a license audit of its dependencies first. Run the audit and give me the finished report. From d8613c319e34afcee98e385180adf132c9ca9454 Mon Sep 17 00:00:00 2001 From: Dominik Simonik Date: Mon, 21 Sep 2026 15:32:08 +0200 Subject: [PATCH 16/18] feat(evals): grade every very-good-cli case on the call, not the prose Eight cases now drive or assert a mocked tool, up from one. All measured at --runs 3 with the plugin loaded. Calls the tool: license-compliance-runs-check-with-full-license-info 1.00 license-compliance-scopes-check-to-monorepo-subdirectory 1.00 ui-package-scaffolds-with-app-ui-package-template 1.00 (was 0.88) create-project-plans-dependency-install 1.00 create-project-does-not-over-ask 1.00 (was 0.75) Asserts the tool was NOT called, which is where mocks help most: "it did not scaffold" becomes a fact from the trace rather than a judge reading prose. create-project-stays-out-of-existing-project-work 1.00 create-project-asks-for-organization-when-required 0.95 (was 0.83) create-project-asks-when-the-template-is-ambiguous 0.92 Both improved cases improved for the same reason: the mechanical grader is 3/3 and the old rubric is what flaps. green-gate is left alone. Its loop needs analyze_files and dart_format, and converting only `test` would strand the model mid-loop. It belongs with the dart mocks. Not converted, having read them: cases whose deliverable is code rather than a call, create-project-scopes-dependency-install-to-the-new-project, whose no-invented-organization grader requires asking and so forbids the call, and static-security-scans-dependencies-before-release, which is about CVEs rather than licenses. license-compliance-refuses-to-certify-from-pubspec-alone was converted and reverted. Its prompt ends "I'd rather not run any scan", so requiring the call demanded the model override the user, and it scored 0.62 three times for that reason. Back to 1.00. README.md gains the trap this pass hit four times: converting a case to drive a tool invalidates every rubric that read the narration, because the model stops describing the call and just makes it. Co-Authored-By: Claude Opus 5 --- evals/README.md | 7 +++++++ .../does-not-create-without-the-organization.md | 7 +++++++ .../graders/does-not-create-while-ambiguous.md | 7 +++++++ .../graders/creates-without-further-questions.md | 6 ++++++ .../does-not-interrogate-for-optionals.md | 4 +++- .../graders/creates-the-project.md | 5 +++++ .../graders/installs-after-creating.md | 6 ++++++ .../graders/names-the-flutter-app-template.md | 5 +++++ .../graders/plan-installs-dependencies.md | 16 ---------------- .../graders/scaffolds-through-very-good-cli.md | 16 ---------------- .../prompt.md | 2 +- .../graders/does-not-scaffold.md | 6 ++++++ .../graders/directory-points-at-mobile.md | 4 ---- .../graders/names-check-licenses.md | 4 ---- .../graders/scoped-to-the-app-subdirectory.md | 7 ------- .../graders/scopes-the-check-to-mobile.md | 6 ++++++ .../prompt.md | 6 ++---- .../graders/names-very-good-cli.md | 4 ---- .../graders/scaffolds-from-template.md | 7 ------- .../graders/scaffolds-with-the-template.md | 6 ++++++ .../graders/uses-app-ui-package-template.md | 4 ---- .../prompt.md | 2 +- 22 files changed, 68 insertions(+), 69 deletions(-) create mode 100644 evals/create-project/create-project-asks-for-organization-when-required/graders/does-not-create-without-the-organization.md create mode 100644 evals/create-project/create-project-asks-when-the-template-is-ambiguous/graders/does-not-create-while-ambiguous.md create mode 100644 evals/create-project/create-project-does-not-over-ask/graders/creates-without-further-questions.md create mode 100644 evals/create-project/create-project-plans-dependency-install/graders/creates-the-project.md create mode 100644 evals/create-project/create-project-plans-dependency-install/graders/installs-after-creating.md create mode 100644 evals/create-project/create-project-plans-dependency-install/graders/names-the-flutter-app-template.md delete mode 100644 evals/create-project/create-project-plans-dependency-install/graders/plan-installs-dependencies.md delete mode 100644 evals/create-project/create-project-plans-dependency-install/graders/scaffolds-through-very-good-cli.md create mode 100644 evals/create-project/create-project-stays-out-of-existing-project-work/graders/does-not-scaffold.md delete mode 100644 evals/license-compliance/license-compliance-scopes-check-to-monorepo-subdirectory/graders/directory-points-at-mobile.md delete mode 100644 evals/license-compliance/license-compliance-scopes-check-to-monorepo-subdirectory/graders/names-check-licenses.md delete mode 100644 evals/license-compliance/license-compliance-scopes-check-to-monorepo-subdirectory/graders/scoped-to-the-app-subdirectory.md create mode 100644 evals/license-compliance/license-compliance-scopes-check-to-monorepo-subdirectory/graders/scopes-the-check-to-mobile.md delete mode 100644 evals/ui-package/ui-package-scaffolds-with-app-ui-package-template/graders/names-very-good-cli.md delete mode 100644 evals/ui-package/ui-package-scaffolds-with-app-ui-package-template/graders/scaffolds-from-template.md create mode 100644 evals/ui-package/ui-package-scaffolds-with-app-ui-package-template/graders/scaffolds-with-the-template.md delete mode 100644 evals/ui-package/ui-package-scaffolds-with-app-ui-package-template/graders/uses-app-ui-package-template.md diff --git a/evals/README.md b/evals/README.md index 62ca976..1408a4f 100644 --- a/evals/README.md +++ b/evals/README.md @@ -278,6 +278,13 @@ answer to the no-plugin arm, exactly as a non-neutral fixture does. `packages_check_licenses` returns one `GPL-3.0` and one `unknown` among twelve permissive licenses, so a case has something real to flag. +**Converting a case to drive a tool invalidates every rubric that read the narration.** The +model stops describing the call and just makes it, so a blind judge sees no evidence and +fails a rubric that was passing. Four rubrics went stale this way in one pass, two of them +asserting outright that no tool was available. When you convert a case, reread every `llm` +grader on it: replace the ones that judged the described call with `tool_used`, and keep +only those judging something still in the reply. + **A mocked tool is not there in the no-plugin arm**, so a `tool_used` grader on one fails for free and takes any grader that needs the tool's output with it. Unlike `tool_used: Skill`, nothing excludes these from the score, so Δ reads as if the plugin diff --git a/evals/create-project/create-project-asks-for-organization-when-required/graders/does-not-create-without-the-organization.md b/evals/create-project/create-project-asks-for-organization-when-required/graders/does-not-create-without-the-organization.md new file mode 100644 index 0000000..b6ce384 --- /dev/null +++ b/evals/create-project/create-project-asks-for-organization-when-required/graders/does-not-create-without-the-organization.md @@ -0,0 +1,7 @@ +--- +type: tool_used +tool: mcp__plugin_vgv-ai-flutter-plugin_very-good-cli__create +min: 0 +max: 0 +weight: 3 +--- diff --git a/evals/create-project/create-project-asks-when-the-template-is-ambiguous/graders/does-not-create-while-ambiguous.md b/evals/create-project/create-project-asks-when-the-template-is-ambiguous/graders/does-not-create-while-ambiguous.md new file mode 100644 index 0000000..b6ce384 --- /dev/null +++ b/evals/create-project/create-project-asks-when-the-template-is-ambiguous/graders/does-not-create-while-ambiguous.md @@ -0,0 +1,7 @@ +--- +type: tool_used +tool: mcp__plugin_vgv-ai-flutter-plugin_very-good-cli__create +min: 0 +max: 0 +weight: 3 +--- diff --git a/evals/create-project/create-project-does-not-over-ask/graders/creates-without-further-questions.md b/evals/create-project/create-project-does-not-over-ask/graders/creates-without-further-questions.md new file mode 100644 index 0000000..f088fd0 --- /dev/null +++ b/evals/create-project/create-project-does-not-over-ask/graders/creates-without-further-questions.md @@ -0,0 +1,6 @@ +--- +type: tool_used +tool: mcp__plugin_vgv-ai-flutter-plugin_very-good-cli__create +input_match: '"org_name"\s*:\s*"com\.example\.mystore"' +weight: 3 +--- diff --git a/evals/create-project/create-project-does-not-over-ask/graders/does-not-interrogate-for-optionals.md b/evals/create-project/create-project-does-not-over-ask/graders/does-not-interrogate-for-optionals.md index b4752eb..589385e 100644 --- a/evals/create-project/create-project-does-not-over-ask/graders/does-not-interrogate-for-optionals.md +++ b/evals/create-project/create-project-does-not-over-ask/graders/does-not-interrogate-for-optionals.md @@ -2,6 +2,8 @@ type: llm --- -PASS if the response proceeds with the name and organization it was given, without interrogating the user for optional parameters such as a description, an output directory, or an application id. Supplying its own value for an optional parameter satisfies this; so does handing the user a ready-to-run command, since no create tool is available in this environment. +PASS if the response proceeds with the name and organization it was given, without +interrogating the user for optional parameters such as a description, an output directory, +or an application id. Supplying its own value for an optional parameter satisfies this. FAIL only if it asks the user to supply more input before proceeding. diff --git a/evals/create-project/create-project-plans-dependency-install/graders/creates-the-project.md b/evals/create-project/create-project-plans-dependency-install/graders/creates-the-project.md new file mode 100644 index 0000000..fa68c2f --- /dev/null +++ b/evals/create-project/create-project-plans-dependency-install/graders/creates-the-project.md @@ -0,0 +1,5 @@ +--- +type: tool_used +tool: mcp__plugin_vgv-ai-flutter-plugin_very-good-cli__create +input_match: '"name"\s*:\s*"my_store"' +--- diff --git a/evals/create-project/create-project-plans-dependency-install/graders/installs-after-creating.md b/evals/create-project/create-project-plans-dependency-install/graders/installs-after-creating.md new file mode 100644 index 0000000..299354f --- /dev/null +++ b/evals/create-project/create-project-plans-dependency-install/graders/installs-after-creating.md @@ -0,0 +1,6 @@ +--- +type: tool_order +before: mcp__plugin_vgv-ai-flutter-plugin_very-good-cli__create +after: mcp__plugin_vgv-ai-flutter-plugin_very-good-cli__packages_get +weight: 3 +--- diff --git a/evals/create-project/create-project-plans-dependency-install/graders/names-the-flutter-app-template.md b/evals/create-project/create-project-plans-dependency-install/graders/names-the-flutter-app-template.md new file mode 100644 index 0000000..4215054 --- /dev/null +++ b/evals/create-project/create-project-plans-dependency-install/graders/names-the-flutter-app-template.md @@ -0,0 +1,5 @@ +--- +type: tool_used +tool: mcp__plugin_vgv-ai-flutter-plugin_very-good-cli__create +input_match: '"subcommand"\s*:\s*"flutter_app"' +--- diff --git a/evals/create-project/create-project-plans-dependency-install/graders/plan-installs-dependencies.md b/evals/create-project/create-project-plans-dependency-install/graders/plan-installs-dependencies.md deleted file mode 100644 index 063b2c4..0000000 --- a/evals/create-project/create-project-plans-dependency-install/graders/plan-installs-dependencies.md +++ /dev/null @@ -1,16 +0,0 @@ ---- -type: llm ---- - -The response describes the steps for creating a new Flutter project. Judge the steps it -describes, whether they are presented as work already done, as commands handed to the user -to run, or as steps that would run in an environment with shell access. A response that -says it cannot execute the steps itself is still judged on the steps it lists. - -PASS if the described steps include installing dependencies after the project is created, -for example a `very_good packages get` or `flutter pub get` step that follows the creation -step. - -FAIL if the described steps contain no dependency install after the creation step. It -does not matter what else follows creation: a plan that goes straight from creating the -project to running or editing it, with no install anywhere, fails. diff --git a/evals/create-project/create-project-plans-dependency-install/graders/scaffolds-through-very-good-cli.md b/evals/create-project/create-project-plans-dependency-install/graders/scaffolds-through-very-good-cli.md deleted file mode 100644 index 5b8daec..0000000 --- a/evals/create-project/create-project-plans-dependency-install/graders/scaffolds-through-very-good-cli.md +++ /dev/null @@ -1,16 +0,0 @@ ---- -type: llm ---- - -The response describes the steps for creating a new Flutter project. Judge the steps it -describes, whether they are presented as work already done, as commands handed to the user -to run, or as steps that would run in an environment with shell access. Naming a shell -command is acceptable here because no project-creation tool is available to call. - -PASS if the scaffolding step uses Very Good CLI (`very_good create`) and states an explicit -template name such as `flutter_app`. - -FAIL only if the scaffolding step the response actually settles on is the Flutter SDK's own -`flutter create`, or if it uses Very Good CLI without naming which template it creates. A -mention of `flutter create` as a fallback, an alternative, or a contrast does not fail this -grader when the scaffolding step it settles on is Very Good CLI with a named template. diff --git a/evals/create-project/create-project-plans-dependency-install/prompt.md b/evals/create-project/create-project-plans-dependency-install/prompt.md index 4b54fbf..449aae4 100644 --- a/evals/create-project/create-project-plans-dependency-install/prompt.md +++ b/evals/create-project/create-project-plans-dependency-install/prompt.md @@ -6,4 +6,4 @@ tags: [create-project] description: The narrated plan scaffolds through Very Good CLI with an explicit template and does not stop at creation. --- -Create a Flutter app named my_store for organization com.example. Walk me through every step you will take. +Create a Flutter app named my_store for organization com.example. diff --git a/evals/create-project/create-project-stays-out-of-existing-project-work/graders/does-not-scaffold.md b/evals/create-project/create-project-stays-out-of-existing-project-work/graders/does-not-scaffold.md new file mode 100644 index 0000000..1b56ed4 --- /dev/null +++ b/evals/create-project/create-project-stays-out-of-existing-project-work/graders/does-not-scaffold.md @@ -0,0 +1,6 @@ +--- +type: tool_used +tool: mcp__plugin_vgv-ai-flutter-plugin_very-good-cli__create +min: 0 +max: 0 +--- diff --git a/evals/license-compliance/license-compliance-scopes-check-to-monorepo-subdirectory/graders/directory-points-at-mobile.md b/evals/license-compliance/license-compliance-scopes-check-to-monorepo-subdirectory/graders/directory-points-at-mobile.md deleted file mode 100644 index 8315277..0000000 --- a/evals/license-compliance/license-compliance-scopes-check-to-monorepo-subdirectory/graders/directory-points-at-mobile.md +++ /dev/null @@ -1,4 +0,0 @@ ---- -type: regex -pattern: 'directory[^\n]{0,24}mobile' ---- diff --git a/evals/license-compliance/license-compliance-scopes-check-to-monorepo-subdirectory/graders/names-check-licenses.md b/evals/license-compliance/license-compliance-scopes-check-to-monorepo-subdirectory/graders/names-check-licenses.md deleted file mode 100644 index 8dc219d..0000000 --- a/evals/license-compliance/license-compliance-scopes-check-to-monorepo-subdirectory/graders/names-check-licenses.md +++ /dev/null @@ -1,4 +0,0 @@ ---- -type: regex -pattern: 'packages[ _]check[ _]licenses' ---- diff --git a/evals/license-compliance/license-compliance-scopes-check-to-monorepo-subdirectory/graders/scoped-to-the-app-subdirectory.md b/evals/license-compliance/license-compliance-scopes-check-to-monorepo-subdirectory/graders/scoped-to-the-app-subdirectory.md deleted file mode 100644 index 0c384f0..0000000 --- a/evals/license-compliance/license-compliance-scopes-check-to-monorepo-subdirectory/graders/scoped-to-the-app-subdirectory.md +++ /dev/null @@ -1,7 +0,0 @@ ---- -type: llm ---- - -PASS if the license check is scoped to the mobile subdirectory by passing a directory argument naming it. - -FAIL if it runs the check at the repository root, or names no directory at all. diff --git a/evals/license-compliance/license-compliance-scopes-check-to-monorepo-subdirectory/graders/scopes-the-check-to-mobile.md b/evals/license-compliance/license-compliance-scopes-check-to-monorepo-subdirectory/graders/scopes-the-check-to-mobile.md new file mode 100644 index 0000000..e2b2208 --- /dev/null +++ b/evals/license-compliance/license-compliance-scopes-check-to-monorepo-subdirectory/graders/scopes-the-check-to-mobile.md @@ -0,0 +1,6 @@ +--- +type: tool_used +tool: mcp__plugin_vgv-ai-flutter-plugin_very-good-cli__packages_check_licenses +input_match: '"directory"\s*:\s*"[^"]*mobile' +weight: 3 +--- diff --git a/evals/license-compliance/license-compliance-scopes-check-to-monorepo-subdirectory/prompt.md b/evals/license-compliance/license-compliance-scopes-check-to-monorepo-subdirectory/prompt.md index bf04ef6..63d75ce 100644 --- a/evals/license-compliance/license-compliance-scopes-check-to-monorepo-subdirectory/prompt.md +++ b/evals/license-compliance/license-compliance-scopes-check-to-monorepo-subdirectory/prompt.md @@ -3,9 +3,7 @@ max_turns: 12 timeout_seconds: 600 allowed_tools: [Read, Glob, Grep, Skill] tags: [license-compliance] -description: The Core Standard that a project below the workspace root needs a directory argument. +description: The license check is scoped to the app subdirectory of a monorepo rather than run at the repo root. --- -Assume this repository layout, do not inspect the working directory, and do not run anything: -melos.yaml at the repo root, the Flutter app at mobile/ with its own pubspec.yaml, and shared packages under packages/. -I want the app's dependency licenses audited. What is the exact tool call, with every argument? +Assume this repository layout and answer from it rather than inspecting the working directory: melos.yaml at the repo root, the Flutter app at mobile/ with its own pubspec.yaml, and shared packages under packages/. Audit the app's dependency licenses. diff --git a/evals/ui-package/ui-package-scaffolds-with-app-ui-package-template/graders/names-very-good-cli.md b/evals/ui-package/ui-package-scaffolds-with-app-ui-package-template/graders/names-very-good-cli.md deleted file mode 100644 index 56ae305..0000000 --- a/evals/ui-package/ui-package-scaffolds-with-app-ui-package-template/graders/names-very-good-cli.md +++ /dev/null @@ -1,4 +0,0 @@ ---- -type: regex -pattern: '[Vv]ery[-_ ][Gg]ood' ---- diff --git a/evals/ui-package/ui-package-scaffolds-with-app-ui-package-template/graders/scaffolds-from-template.md b/evals/ui-package/ui-package-scaffolds-with-app-ui-package-template/graders/scaffolds-from-template.md deleted file mode 100644 index ae8eb07..0000000 --- a/evals/ui-package/ui-package-scaffolds-with-app-ui-package-template/graders/scaffolds-from-template.md +++ /dev/null @@ -1,7 +0,0 @@ ---- -type: llm ---- - -PASS if the package is scaffolded from the Very Good CLI `app_ui_package` template. - -FAIL if the response builds the package by hand, or reaches for the Flutter SDK's own `flutter create --template=package`. diff --git a/evals/ui-package/ui-package-scaffolds-with-app-ui-package-template/graders/scaffolds-with-the-template.md b/evals/ui-package/ui-package-scaffolds-with-app-ui-package-template/graders/scaffolds-with-the-template.md new file mode 100644 index 0000000..b2bf833 --- /dev/null +++ b/evals/ui-package/ui-package-scaffolds-with-app-ui-package-template/graders/scaffolds-with-the-template.md @@ -0,0 +1,6 @@ +--- +type: tool_used +tool: mcp__plugin_vgv-ai-flutter-plugin_very-good-cli__create +input_match: '"subcommand"\s*:\s*"app_ui_package"' +weight: 3 +--- diff --git a/evals/ui-package/ui-package-scaffolds-with-app-ui-package-template/graders/uses-app-ui-package-template.md b/evals/ui-package/ui-package-scaffolds-with-app-ui-package-template/graders/uses-app-ui-package-template.md deleted file mode 100644 index 1f32df5..0000000 --- a/evals/ui-package/ui-package-scaffolds-with-app-ui-package-template/graders/uses-app-ui-package-template.md +++ /dev/null @@ -1,4 +0,0 @@ ---- -type: regex -pattern: 'app_ui_package' ---- diff --git a/evals/ui-package/ui-package-scaffolds-with-app-ui-package-template/prompt.md b/evals/ui-package/ui-package-scaffolds-with-app-ui-package-template/prompt.md index 9e6a514..05407d8 100644 --- a/evals/ui-package/ui-package-scaffolds-with-app-ui-package-template/prompt.md +++ b/evals/ui-package/ui-package-scaffolds-with-app-ui-package-template/prompt.md @@ -6,4 +6,4 @@ tags: [ui-package] description: "Scaffolding through the Very Good CLI app_ui_package template, and the lib/src + single-barrel layout that template ships." --- -Set up a new UI package called storefront_ui in my monorepo for our shared widgets and design tokens. Which tool or command would you use, with which arguments, and what will the package layout look like? Do not run anything yet. +Set up a new UI package called storefront_ui in my monorepo for our shared widgets and design tokens, and show me what the package layout will look like. From 41beba2148bc9e3128b55d6afe01e26e871f0e7b Mon Sep 17 00:00:00 2001 From: Dominik Simonik Date: Mon, 21 Sep 2026 15:43:05 +0200 Subject: [PATCH 17/18] feat(evals): assert the create call's placement, and that tools stay quiet out of scope ui-package-scaffolds-with-app-ui-package-template now checks the create call passes output_directory, which the skill teaches for a package in a monorepo and nothing verified. It drops src-and-barrel-layout, which began failing 3/3 once the case drove the tool: the model stops describing the layout at length and a blind judge sees nothing to judge. That convention is still covered by ui-package-refuses-imports-from-src and ui-package-adds-widget-with-barrel-export-and-test. Back to 1.00 at --runs 3. Both negative controls now assert the tool was never called rather than only that the skill stayed quiet, which is weaker: a skill can stay quiet while the model still reaches for a tool. testing-stays-out-of-non-test-work and ui-package-stays-out-of-plain-dart-work, both 1.00. A case covering the testing skill's two `test` MCP tool standards, `directory` for a monorepo package and `timeout_seconds`, was written and removed. It scored 0.54 with `Skill called 0x` on every run: no skill fires on a request to run a suite, because green-gate's description claims that ground, saying it owns "which tool and arguments run each gate" and to prefer it over the single-gate testing skill. Those two standards look misplaced in skills/testing/SKILL.md, and green-gate is where they could be graded. Left for the dart-mock work rather than widening a shipped description to make a new case pass. Co-Authored-By: Claude Opus 5 --- .../graders/does-not-run-tests.md | 6 ++++++ .../graders/places-it-in-the-monorepo.md | 5 +++++ .../graders/src-and-barrel-layout.md | 7 ------- .../graders/does-not-scaffold.md | 6 ++++++ 4 files changed, 17 insertions(+), 7 deletions(-) create mode 100644 evals/testing/testing-stays-out-of-non-test-work/graders/does-not-run-tests.md create mode 100644 evals/ui-package/ui-package-scaffolds-with-app-ui-package-template/graders/places-it-in-the-monorepo.md delete mode 100644 evals/ui-package/ui-package-scaffolds-with-app-ui-package-template/graders/src-and-barrel-layout.md create mode 100644 evals/ui-package/ui-package-stays-out-of-plain-dart-work/graders/does-not-scaffold.md diff --git a/evals/testing/testing-stays-out-of-non-test-work/graders/does-not-run-tests.md b/evals/testing/testing-stays-out-of-non-test-work/graders/does-not-run-tests.md new file mode 100644 index 0000000..6a8498f --- /dev/null +++ b/evals/testing/testing-stays-out-of-non-test-work/graders/does-not-run-tests.md @@ -0,0 +1,6 @@ +--- +type: tool_used +tool: mcp__plugin_vgv-ai-flutter-plugin_very-good-cli__test +min: 0 +max: 0 +--- diff --git a/evals/ui-package/ui-package-scaffolds-with-app-ui-package-template/graders/places-it-in-the-monorepo.md b/evals/ui-package/ui-package-scaffolds-with-app-ui-package-template/graders/places-it-in-the-monorepo.md new file mode 100644 index 0000000..16fba6c --- /dev/null +++ b/evals/ui-package/ui-package-scaffolds-with-app-ui-package-template/graders/places-it-in-the-monorepo.md @@ -0,0 +1,5 @@ +--- +type: tool_used +tool: mcp__plugin_vgv-ai-flutter-plugin_very-good-cli__create +input_match: '"output_directory"\s*:' +--- diff --git a/evals/ui-package/ui-package-scaffolds-with-app-ui-package-template/graders/src-and-barrel-layout.md b/evals/ui-package/ui-package-scaffolds-with-app-ui-package-template/graders/src-and-barrel-layout.md deleted file mode 100644 index 97d79ec..0000000 --- a/evals/ui-package/ui-package-scaffolds-with-app-ui-package-template/graders/src-and-barrel-layout.md +++ /dev/null @@ -1,7 +0,0 @@ ---- -type: llm ---- - -PASS if the described layout keeps widget and theme source files under `lib/src/` and exposes the public API through a single barrel file directly under `lib/`. - -FAIL if the layout puts public widgets at the top of `lib/` with no barrel file, or names no barrel file at all, or if the widget and theme source files sit anywhere other than under `lib/src/`. diff --git a/evals/ui-package/ui-package-stays-out-of-plain-dart-work/graders/does-not-scaffold.md b/evals/ui-package/ui-package-stays-out-of-plain-dart-work/graders/does-not-scaffold.md new file mode 100644 index 0000000..1b56ed4 --- /dev/null +++ b/evals/ui-package/ui-package-stays-out-of-plain-dart-work/graders/does-not-scaffold.md @@ -0,0 +1,6 @@ +--- +type: tool_used +tool: mcp__plugin_vgv-ai-flutter-plugin_very-good-cli__create +min: 0 +max: 0 +--- From 46a6673caeba9ae754ef3680741338c072c58281 Mon Sep 17 00:00:00 2001 From: Dominik Simonik Date: Tue, 22 Sep 2026 09:50:29 +0200 Subject: [PATCH 18/18] fix(evals): give the tool-driving cases room to make the call A full run put ui-package-scaffolds-with-app-ui-package-template over both the 12-turn cap and the 600s timeout. It scored 1.00 anyway, because the graders were satisfied before the limits hit, which is luck rather than a pass: a slower run fails before the call is made. Measured again at --runs 3 it takes 11, 15 and 14 turns. A converted case does strictly more than the prompt it replaced. It routes, calls the tool, reads the answer, then writes the reply. The four cases that drive a tool now carry max_turns: 20 and timeout_seconds: 900, and README.md records why, including that a cap breach scores 0 with no failing grader and so reads as a content failure. The two other cases under threshold in that run were both transient, and neither is touched by this branch. At --runs 3, green-gate-escalates-when-the-loop-stops-making-progress is 1.00 3/3 against 0.00, and ui-package-declines-hand-rolled-button is a consistent 0.86, which clears the threshold. Both had turns=1 in the sweep. Co-Authored-By: Claude Opus 5 (1M context) --- evals/README.md | 7 +++++++ .../create-project-plans-dependency-install/prompt.md | 4 ++-- .../prompt.md | 4 ++-- .../prompt.md | 4 ++-- .../prompt.md | 4 ++-- 5 files changed, 15 insertions(+), 8 deletions(-) diff --git a/evals/README.md b/evals/README.md index 1408a4f..8643b46 100644 --- a/evals/README.md +++ b/evals/README.md @@ -285,6 +285,13 @@ asserting outright that no tool was available. When you convert a case, reread e grader on it: replace the ones that judged the described call with `tool_used`, and keep only those judging something still in the reply. +**Driving a tool costs turns and wall clock.** A converted case does strictly more than the +prompt it replaced: it routes, calls, reads the answer, then writes the reply. +`ui-package-scaffolds-with-app-ui-package-template` measured 11 to 15 turns where the +suite's usual `max_turns: 12` and `timeout_seconds: 600` had been ample, and hit both +limits. The four tool-driving cases carry `max_turns: 20` and `timeout_seconds: 900`. A cap +breach scores the case 0 with no failing grader, so it reads as a content failure. + **A mocked tool is not there in the no-plugin arm**, so a `tool_used` grader on one fails for free and takes any grader that needs the tool's output with it. Unlike `tool_used: Skill`, nothing excludes these from the score, so Δ reads as if the plugin diff --git a/evals/create-project/create-project-plans-dependency-install/prompt.md b/evals/create-project/create-project-plans-dependency-install/prompt.md index 449aae4..85170c7 100644 --- a/evals/create-project/create-project-plans-dependency-install/prompt.md +++ b/evals/create-project/create-project-plans-dependency-install/prompt.md @@ -1,6 +1,6 @@ --- -max_turns: 12 -timeout_seconds: 600 +max_turns: 20 +timeout_seconds: 900 allowed_tools: [Read, Glob, Grep, Skill] tags: [create-project] description: The narrated plan scaffolds through Very Good CLI with an explicit template and does not stop at creation. diff --git a/evals/license-compliance/license-compliance-runs-check-with-full-license-info/prompt.md b/evals/license-compliance/license-compliance-runs-check-with-full-license-info/prompt.md index 345c35e..e51c130 100644 --- a/evals/license-compliance/license-compliance-runs-check-with-full-license-info/prompt.md +++ b/evals/license-compliance/license-compliance-runs-check-with-full-license-info/prompt.md @@ -1,6 +1,6 @@ --- -max_turns: 12 -timeout_seconds: 600 +max_turns: 20 +timeout_seconds: 900 allowed_tools: [Read, Glob, Grep, Skill] tags: [license-compliance] description: Runs the license check with full license information and turns the scan output into the prescribed risk-rated report. diff --git a/evals/license-compliance/license-compliance-scopes-check-to-monorepo-subdirectory/prompt.md b/evals/license-compliance/license-compliance-scopes-check-to-monorepo-subdirectory/prompt.md index 63d75ce..3df3257 100644 --- a/evals/license-compliance/license-compliance-scopes-check-to-monorepo-subdirectory/prompt.md +++ b/evals/license-compliance/license-compliance-scopes-check-to-monorepo-subdirectory/prompt.md @@ -1,6 +1,6 @@ --- -max_turns: 12 -timeout_seconds: 600 +max_turns: 20 +timeout_seconds: 900 allowed_tools: [Read, Glob, Grep, Skill] tags: [license-compliance] description: The license check is scoped to the app subdirectory of a monorepo rather than run at the repo root. diff --git a/evals/ui-package/ui-package-scaffolds-with-app-ui-package-template/prompt.md b/evals/ui-package/ui-package-scaffolds-with-app-ui-package-template/prompt.md index 05407d8..259962c 100644 --- a/evals/ui-package/ui-package-scaffolds-with-app-ui-package-template/prompt.md +++ b/evals/ui-package/ui-package-scaffolds-with-app-ui-package-template/prompt.md @@ -1,6 +1,6 @@ --- -max_turns: 12 -timeout_seconds: 600 +max_turns: 20 +timeout_seconds: 900 allowed_tools: [Read, Glob, Grep, Skill] tags: [ui-package] description: "Scaffolding through the Very Good CLI app_ui_package template, and the lib/src + single-barrel layout that template ships."