diff --git a/.github/workflows/evals.yaml b/.github/workflows/evals.yaml index 645abd6..4bdfeb5 100644 --- a/.github/workflows/evals.yaml +++ b/.github/workflows/evals.yaml @@ -82,12 +82,14 @@ jobs: echo "$CHANGED" | { grep -oE '^skills/[^/]+/' || true; } | cut -d/ -f2 echo "$CHANGED" | { grep -oE '^evals/[^_][^/]*/' || true; } | cut -d/ -f2 } | sort -u | while read -r s; do - [ -n "$s" ] && [ -d "evals/$s" ] && echo "$s" + if [ -n "$s" ] && compgen -G "evals/$s/*/case.yaml" >/dev/null; then + echo "$s" + fi done ) - if echo "$CHANGED" | grep -qE '^evals/_fixture/'; then - echo "fixture changed — widening to all skills" + if echo "$CHANGED" | grep -qE '^evals/(_fixture|mocks)/'; then + echo "shared eval input changed — widening to all skills" SKILLS=$(all_skills) fi fi diff --git a/AGENTS.md b/AGENTS.md index 860d2e5..2afbd9f 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -17,7 +17,6 @@ docs/ # Gitignored, local only plan/ # Planning and design documents evals/ # `claude plugin eval` suite — all 15 skills, 100 cases README.md # Case format, grader reference, how to add a case - BASELINE.md # Last full two-arm measurement: per-skill Δ and routing _fixture/ fixture.sh # The only copy of the neutral Flutter skeleton; every case symlinks here / # One directory per skill, named for the skill it covers diff --git a/config/cspell.json b/config/cspell.json index cf6cadb..99c4f7b 100644 --- a/config/cspell.json +++ b/config/cspell.json @@ -13,6 +13,7 @@ "Bienvenido", "bypassable", "Codex", + "compgen", "CSPRNG", "cupertino", "Cupertino", diff --git a/evals/BASELINE.md b/evals/BASELINE.md deleted file mode 100644 index b9ab8f1..0000000 --- a/evals/BASELINE.md +++ /dev/null @@ -1,110 +0,0 @@ -# Measured baseline - -Observed numbers from `claude plugin eval`, not estimates. Re-measure after any change to -a skill description, a skill body, a prompt, or the fixture, and date new numbers rather -than editing these. - -```bash -claude plugin eval . --trust-plugin --scaffold \ - --ablation with-without --runs 1 --threshold 0.8 \ - --model claude-sonnet-5 --judge-model claude-sonnet-5 \ - --no-publish --max-cost-usd 45 -j 4 -``` - -`--runs 1` is deliberate for a two-arm sweep: one no-plugin pass is enough to disqualify a -grader, which is what the sweep is for. Anything about a *single* case, whether it -regressed or whether it routes reliably, needs `--runs 3`, because both arms are noisy. - -## Full two-arm run, 2026-09-17, Claude Code 2.1.270 - -100 cases, 200 agent runs, 0 run errors, `partial: false`, $24.93, 30 minutes at `-j 4`. - -| | value | -| ------------------------------- | ---------: | -| Mean Δ, all 100 cases | **+0.65** | -| Mean Δ where the skill routed | **+0.76** | -| Routing misses (positive cases) | **0/84** | -| Cases at or above threshold 0.8 | **97/100** | -| Cases scoring a perfect 1.00 | **85/100** | -| Positive cases with Δ <= 0 | **0/84** | -| Suite score | **0.971** | - -**Routing is solved for now.** Every one of the 84 positive cases activated its skill. -The previous sweep, under a Haiku judge and before the refusal-shaped description fixes, -measured 5 misses out of 83, and the one before that measured 14. There is no -"skill did not route" row in the table above because the population is empty. - -## Reading the two numbers that look bad - -**15 cases show Δ <= 0.** All 15 are negative controls, and that is the design: a model -with no plugin passes a "must not invoke the skill" check for free, so its without-arm -score is already near perfect and there is no lift to measure. Among the 84 positive -cases, nothing scored Δ <= 0. Exclude negative controls whenever you compare arms. - -**121 content graders pass with no plugin loaded.** 28 of those sit on the six positive -cases that clear less than Δ 0.50, which is where a free grader actually costs signal: - -- `accessibility-declines-gesture-detector-tap-target` -- `bloc-tests-with-bloc-test-and-mocktail` -- `internationalization-uses-directional-insets-for-rtl` -- `layered-architecture-transforms-models-in-the-repository` -- `license-compliance-refuses-to-clear-missing-licenses` -- `testing-uses-pump-app-in-widget-tests` - -A grader is only worth keeping if it can fail in the no-plugin arm. Those six cases are -the shortlist for the next grader pass. The remaining 93 free graders sit on cases that -discriminate through their other graders, which is untidy rather than disqualifying. - -## Per-skill mean Δ - -| skill | mean Δ | -| -------------------------- | -----: | -| ui-package | +0.79 | -| material-theming | +0.73 | -| very-good-analysis-upgrade | +0.71 | -| animations | +0.69 | -| dart-flutter-sdk-upgrade | +0.69 | -| testing | +0.66 | -| green-gate | +0.65 | -| internationalization | +0.63 | -| layered-architecture | +0.63 | -| bloc | +0.63 | -| license-compliance | +0.61 | -| navigation | +0.58 | -| accessibility | +0.58 | -| static-security | +0.56 | -| create-project | +0.56 | - -The spread is narrow and every skill clears +0.55, so no skill is currently carrying the -suite or dragging on it. `create-project` sits last partly because its `SKILL.md` pins -`model: haiku`, so its with-arm answers on a weaker model than its baseline. See the -caveats. - -## Cases below threshold - -| case | score | Δ | -| ----------------------------------------------------- | ----: | ----: | -| `static-security-refuses-platform-channel-biometrics` | 0.50 | +0.50 | -| `green-gate-budgets-per-package-across-a-monorepo` | 0.62 | +0.62 | -| `create-project-does-not-over-ask` | 0.75 | +0.50 | - -All three routed, and all three carry a healthy Δ, so the skill is firing and helping. -They are graded strictly rather than broken. Confirm any of them with `--runs 3` before -editing, because a single reading of a case is not a measurement. - -## Caveats - -- **Single run per arm.** Treat any one case's Δ as a shortlist entry, not a verdict. -- **Sonnet judges.** The suite moved off the Haiku judge after it marked two correct - answers wrong on a 100-case run, both clean passes under Sonnet. Sonnet is not - infallible either: read the judge's votes in the report before editing a skill. -- **Not comparable with anything before 2026-09-17.** Earlier numbers in this file's - history were measured under a Haiku judge and before the harness-insulation fixes to - three prompts, so they understate the suite by an unmeasured amount. -- **`create-project` pins `model: haiku`** in its `SKILL.md`. Routing is decided before - the switch, so it never explains a routing miss, but every answer after the skill fires - ran on Haiku while the no-plugin arm ran on Sonnet. Its Δ is understated by an - unmeasured amount. -- **Tool-driven skills are graded on narration.** The MCP servers are not mocked, so the - six skills that drive tools are measured on the calls they describe, not the calls they - make. diff --git a/evals/README.md b/evals/README.md index b9f30f4..8643b46 100644 --- a/evals/README.md +++ b/evals/README.md @@ -9,8 +9,7 @@ claude plugin eval . --scaffold --tag bloc # one skill claude plugin eval . --scaffold --ablation none # with-plugin arm only, half the cost ``` -- Claude Code **>= 2.1.269**. Earlier builds answer `plugin eval is currently in early - access`. +- Claude Code **>= 2.1.269**. - No API key locally. Runs authenticate the same way your normal session does. - `--scaffold` is **not optional**. Without it the fixture never runs and cases fail for unrelated reasons. @@ -24,13 +23,11 @@ claude plugin eval . --scaffold --ablation none # with-plugin arm only, ha | with | The plugin loaded | What the plugin produces | | without | No plugin loaded at all | What the bare model produces | -`Δ` is the with-arm score minus the without-arm score. A grader that passes in both arms -is measuring the model, not the skill. +`Δ` is the with-arm score minus the without-arm score. A grader that passes in both arms is +measuring the model, not the skill. -Each run gets a throwaway home directory, workspace, and Claude Code configuration. Your -settings, `CLAUDE.md`, MCP servers, other plugins, memory, and skills are all absent, and -the eval directory is unreadable from inside a run, so a case cannot read its own graders -or its siblings. +Each run gets a throwaway home directory, workspace, and Claude Code configuration, and the +eval directory is unreadable from inside a run. Exclude negative controls when you compare arms. A model with no plugin passes a "must not invoke the skill" check for free. @@ -62,12 +59,11 @@ tags: [bloc] description: Sealed hierarchies, Equatable, and the pinned state names for a LoginBloc. --- -Write a LoginBloc for email and password authentication with submit and logout events, -and success and failure states. Output Dart code only. +Write a LoginBloc for email and password authentication with submit and logout events. +Output Dart code only. ``` -Every case repeats that frontmatter verbatim. `claude plugin eval` has no shared-defaults -mechanism, so it is forced rather than duplication worth removing. +Every case repeats that frontmatter verbatim; there is no shared-defaults mechanism. `case.yaml`: @@ -101,13 +97,15 @@ threshold is `1.0`, so **pass `--threshold 0.8` or every imperfect case exits 1* | `llm` | `criteria`, `focus` | Frontmatter is just `type: llm`; the file body is the rubric | | `baseline` | `baseline_file`, `criteria` | Unused here | -There are **no custom-code graders**. A check that needs to execute something has no home. +There are **no custom-code graders**. + +A grader reads `last_message` unless its `target` says otherwise. The other targets are +`trace`, `files`, `{ source: file, path: }`, and `mock_calls`, which is every call +made to a [mocked MCP tool](#mocking-the-mcp-servers). ### Routing graders -Ninety-nine of the hundred cases carry one. -`create-project-infers-dart-package-for-api-client` is the exception, and routing is -covered by that skill's other six cases. The shape is always: +Ninety-nine of the hundred cases carry one: ```markdown --- @@ -120,14 +118,13 @@ arm: both ``` **`arm: both` is mandatory.** Without it a two-arm run drops every `tool_used: Skill` -grader from the score and reports it as an indicator, so a case whose skill never fires -still scores `1.00`. The cost is that the without-arm is then penalized for not having the -plugin, which inflates `Δ`. Read the two arm scores rather than `Δ` alone. +grader from the score, so a case whose skill never fires still scores `1.00`. It also +penalizes the without-arm, which inflates `Δ`, so read the two arm scores rather than `Δ`. -**`weight: 3` is mandatory** and nothing applies it for you. It is what sinks a case on a -routing miss alone. It does not protect the content graders: against `n` content graders -of weight 1, a single content miss scores `(n + 2) / (n + 3)`, which clears 0.8 for every -`n >= 2`. When one grader carries a case's whole point, weight that grader too. +**`weight: 3` is mandatory** and nothing applies it for you. It sinks a case on a routing +miss alone. It does not protect the content graders: against `n` content graders of weight +1, a single content miss scores `(n + 2) / (n + 3)`, which clears 0.8 for every `n >= 2`. +When one grader carries a case's whole point, weight that grader too. Negative controls use the same grader with `min: 0` and `max: 0`. @@ -142,21 +139,46 @@ Negative controls use the same grader with `min: 0` and `max: 0`. the conventions the case exists to measure. - **Prompts must be self-contained.** Paste in any class a prompt refers to, and name pasted text as authoritative when it describes state not on disk. -- **The plugin's own SessionStart hook fires inside the run.** `warn-missing-mcp.sh` - injects "Very Good CLI is not installed" into every with-plugin run, because the sandbox - has no `very_good` on PATH. Tool-driven cases answer with that blocker instead of the - question. Say in the prompt that the CLI is installed and that a startup notice saying - otherwise should be ignored. +- **Check both arms answered.** A no-plugin arm that asks a clarifying question instead of + doing the work makes every grader look discriminating. That is a prompt that is not + self-contained, not a result. +- **The plugin's SessionStart hook can fire inside a run.** `warn-missing-mcp.sh` injects + "Very Good CLI is not installed" whenever `check_vgv_cli` returns `not_installed`, and a + tool-driven case then answers with that blocker instead of the question. The + `unverifiable` status keeps it quiet. Two prompts still say the session cannot reach a + real toolchain, which is a different problem and true regardless. Beyond that: write prompts as a user would send them, name no skill in a prompt, grade -mechanically where you can, include the cases where the skill must say no, keep a negative -control's rubric to the absence of the skill's vocabulary, and check a grader fails in the -without-arm before trusting it. +mechanically where you can, include the cases where the skill must say no, and keep a +negative control's rubric to the absence of the skill's vocabulary. + +### Graders that cannot fail + +A grader that passes in the no-plugin arm measures the model, not the skill. Find them by +running the case two-arm and comparing the two `graders` lists. + +**One two-arm run cannot disqualify a grader.** Use `--runs 3` on both arms, and prefer +reasons that do not depend on a score at all. + +**Adding beats deleting.** A free grader still fails if the skill later regresses, so it is +a regression test even when it earns no Δ. Keep every free grader that pins a Core Standard +or an anti-pattern. Delete only for a reason that holds without a score: + +1. **A genuine duplicate** of a grader beside it. +2. **It restates the prompt.** `class WeatherRepository` passes whenever the model read the + question. +3. **It is table stakes for the format**, such as `testWidgets` in a widget test. + +Everything else gets added to. Read both arms' output side by side, find what only the +plugin produced, grade that, and weight it so the case turns on it. + +If nothing discriminates even then, the skill may teach nothing the model does not already +do, and the fix is the skill rather than the case. ### Writing an `llm` rubric -Frontmatter is only `type: llm`. The body is the rubric, written as concrete PASS and FAIL -conditions with an example of each where the wording is open to reading: +Frontmatter is only `type: llm`. The body is the rubric, as concrete PASS and FAIL +conditions with an example of each: ```markdown --- @@ -169,9 +191,8 @@ verb, for example LoginSubmitted. FAIL if any event name is imperative, such as SubmitLogin. ``` -Match the skill's own vocabulary. A rubric asking for "rounds" against a skill that teaches -"iteration count" fails on correct answers. Keep `llm` graders for short output; for -anything long a `regex` reads the whole thing the same way every time. +Match the skill's own vocabulary. Keep `llm` graders for short output; for anything long a +`regex` reads the whole thing the same way every time. --- @@ -180,22 +201,103 @@ anything long a `regex` reads the whole thing the same way every time. Every run starts in an empty workspace. `_fixture/fixture.sh` recreates the neutral Flutter skeleton, hooked up through `context.scaffold_script`, and runs **only with `--scaffold`**. -`context.scaffold_script` will not take a path that leaves the case directory: - -```text -path "../../_fixture/fixture.sh" escapes the case directory (`..` or an absolute path) -— it must name something inside it -``` - -It does resolve a symlink inside the case directory, so each case's `fixture.sh` is a -symlink to the one script. A new case needs its own: +`context.scaffold_script` will not take a path that leaves the case directory, but it does +resolve a symlink inside it, so each case's `fixture.sh` is a symlink to the one script. A +new case needs its own: ```bash ln -s ../../_fixture/fixture.sh evals///fixture.sh ``` **On Windows**, a checkout without `core.symlinks=true` turns each link into a text file -holding the path, and runs then fail at scaffold time. CI runs on Linux and is unaffected. +and runs fail at scaffold time. CI runs on Linux and is unaffected. + +--- + +## Mocking the MCP servers + +`evals/mocks//.md` registers a stand-in under the server's own name from +`.mcp.json`. A server with no mock directory is not started and its tools are absent, which +every run reports on a `mocked:` line. Real servers start only with `--allow-real-servers` +or `--mocks off`, neither of which is used here. + +This applies to eval runs only. A real session still reaches the real `very_good mcp` +server. + +```text +evals/mocks/very-good-cli/ +├── _tools.json # the real tools/list result +├── create.md +├── packages_check_licenses.md +├── packages_get.md +└── test.md +``` + +Frontmatter is optional and the body is the tool result. Three of the four here have no +frontmatter, which is the same as `type: fixed`. + +| Key | Default | Purpose | +| ------------ | ------- | --------------------------------------------------------------- | +| `type` | `fixed` | `agent` instead plays the server through a judge-model call | +| `expect` | unset | Per-input guard: a type name, a literal, a list, or a `/regex/` | +| `error` | `false` | `fixed` only. Return the body as a tool error | +| `abort_when` | unset | `agent` only | + +`{{input.}}` substitutes a call argument into the body, and +`{{file:fixtures/}}` inserts a file from a `fixtures/` directory beside the mock. +`_server.md` answers several tools from one `agent` mock; a `.md` for the same tool +wins. A case's own `mocks/` directory overrides the suite's file by file. + +Four things that are not obvious: + +- **A mocked tool needs no `--allow-tools` grant and no `allowed_tools` entry.** Name it in + `allowed_tools` and you get `not granted (missing --allow-tools grant, or a malformed + entry)`, which means the name is unknown, not that a grant is missing. +- **Use the plugin-namespaced name** in a `tool_used` grader: + `mcp__plugin_vgv-ai-flutter-plugin_very-good-cli__`. The bare + `mcp__very-good-cli__` form the skills use for a real session is not valid here. +- **`expect` treats a missing field as a violation**, and a violation aborts the run at + score 0 with no failing grader to read. Guard only what the real schema requires: here + that is `create`, so only `create.md` carries an `expect`. Grade argument *choices* with + `tool_used` or a `regex` against `mock_calls`. +- **`_tools.json` drives the schema the model sees.** It is the `tools/list` result, the + object with the `tools` array. Regenerate it after a Very Good CLI release by speaking + MCP to `very_good mcp` over stdio; a stale one teaches a schema the CLI no longer has. + +The mocks are reachable only when `check_vgv_cli` returns `unverifiable`. In a run +`very_good` resolves on PATH but `very_good --version` answers nothing, because it is a +shim that execs `dart` under a throwaway `$HOME`. Read as `not_installed`, the PreToolUse +hook denies the call and the model gets the hook's text as the tool result. + +Every mock here is `type: fixed`. An `agent` mock answers through the judge model, so it +costs money, varies run to run, and needs a recording adopted from +`results//mock-recordings/` into `.replay/` before CI repeats. + +Keep mock bodies as raw tool output. A mock that names what a skill teaches hands the +answer to the no-plugin arm, exactly as a non-neutral fixture does. +`packages_check_licenses` returns one `GPL-3.0` and one `unknown` among twelve permissive +licenses, so a case has something real to flag. + +**Converting a case to drive a tool invalidates every rubric that read the narration.** The +model stops describing the call and just makes it, so a blind judge sees no evidence and +fails a rubric that was passing. Four rubrics went stale this way in one pass, two of them +asserting outright that no tool was available. When you convert a case, reread every `llm` +grader on it: replace the ones that judged the described call with `tool_used`, and keep +only those judging something still in the reply. + +**Driving a tool costs turns and wall clock.** A converted case does strictly more than the +prompt it replaced: it routes, calls, reads the answer, then writes the reply. +`ui-package-scaffolds-with-app-ui-package-template` measured 11 to 15 turns where the +suite's usual `max_turns: 12` and `timeout_seconds: 600` had been ample, and hit both +limits. The four tool-driving cases carry `max_turns: 20` and `timeout_seconds: 900`. A cap +breach scores the case 0 with no failing grader, so it reads as a content failure. + +**A mocked tool is not there in the no-plugin arm**, so a `tool_used` grader on one fails +for free and takes any grader that needs the tool's output with it. Unlike +`tool_used: Skill`, nothing excludes these from the score, so Δ reads as if the plugin +supplied the content when it mostly supplied the tool. Read the with-arm score. +`license-compliance-runs-check-with-full-license-info` is the worked example: 1.00 with the +plugin and 0.00 without, on three runs each. --- @@ -211,29 +313,30 @@ $E --runs 3 --threshold 0.8 # is a red case real? ``` Read the two arm scores, not the total. The without-arm is supposed to score badly. -[BASELINE.md](BASELINE.md) records the last full two-arm measurement. **One run is not a measurement**, and these are not a merge gate. -- `--runs 3` before believing a red case. Cases drift between 2/3 and 3/3 on their own. +- `--runs 3` before believing a red case. - A case that hit its turn cap or timed out scores 0 with **no failing grader**, which reads exactly like a content failure. Check the `NOTES` column, or `cases[].arms.with[].error` in the JSON. - A usage limit hit mid-suite makes every later run fail the same way without marking the document `partial`. Check the errors. -- The suite judges with `claude-sonnet-5`, not the default small model, which marked two - correct answers wrong on a 100-case run. Sonnet is not infallible either: when a case - routes but scores badly, read the judge's votes in the report before editing the skill. - -Measured on full runs: **$13.14** for 100 cases in one arm, **$24.93** for both arms, at -roughly **$0.12 per run**. Per-case cost varies several-fold. At `--runs 3` a two-arm sweep -is six runs per case, so budget around **$75**. `--ablation none` halves it, -`--max-cost-usd` bounds it, and `-j` up to 8 shortens wall clock. +- The suite judges with `claude-sonnet-5`. When a case routes but scores badly, read the + judge's votes in the report before editing the skill. +- **The no-plugin arm swings between runs.** One two-arm run cannot disqualify a grader: + `accessibility-declines-gesture-detector-tap-target` failed all three content graders + without the plugin on one run and passed all three on the next. +- **`create-project` pins `model: haiku`.** Its with-arm answers on a weaker model than its + baseline, so its Δ reads low. Routing is decided before the switch, so the pin never + explains a routing miss. + +Costs: **$13** for 100 cases in one arm, **$25** for both, roughly **$0.12 per run**. At +`--runs 3` a two-arm sweep is six runs per case, so budget around **$75**. `-j` up to 8 +shortens wall clock. Every run writes `results//` with `aggregate-result.json` and a self-contained -`report.html`. The report is where you find out *why* a run scored low: failed graders are -expanded, and an `llm` grader shows the judge's votes and the text it judged. `results/` is -gitignored. +`report.html`, which is where you find out *why* a run scored low. `results/` is gitignored. --- @@ -241,38 +344,37 @@ gitignored. `.github/workflows/evals.yaml` runs after a merge to `main`, never on a pull request, scoped by `--tag` to the changed skills, with-plugin arm only, and `continue-on-error`. -A regression is therefore reported after it lands, and a case that has stopped -discriminating goes unnoticed until you re-check with `include_baseline`. - -- Changing `_fixture/` widens the scope to all 15 skills. -- CI needs `ANTHROPIC_API_KEY`, having no Claude Code session, and `--trust-plugin`, - because a run with no terminal cannot answer the trust prompt. -- `--ablation none` is the only mode where a `tool_used: Skill` grader is scored by - default. This suite sets `arm: both` so routing scores in either mode, but the two modes - weight the baseline differently. Compare runs from one mode at a time. -- The job has a one-hour ceiling. 100 cases in one arm at `--runs 1` measured roughly 35 - minutes at `-j 4`. A full two-arm run at `--runs 3` is 600 runs and does not fit. + +- Changing `_fixture/` or `mocks/` widens the scope to all 15 skills. +- A directory under `evals/` only becomes a `--tag` if it holds `*/case.yaml`. `mocks/` and + `results/` sit there without being skills. +- The scope job's `find` uses `-exec dirname {} \;` rather than `-printf '%h\n'`, which is + GNU-only and fails on macOS. +- CI needs `ANTHROPIC_API_KEY`, having no Claude Code session, and `--trust-plugin`. +- `--ablation none` and `--ablation with-without` weight the baseline differently. Compare + runs from one mode at a time. +- The job has a one-hour ceiling. 100 cases in one arm measured roughly 35 minutes at + `-j 4`. A two-arm run at `--runs 3` is 600 runs and does not fit. --- ## What this does not cover -- **Dart syntax.** The previous harness parsed every fenced `dart` block through a custom - JavaScript assertion. Native evals have no custom-code graders, so 30 cases across 11 - skills lost that check. The response text is not in `aggregate-result.json` either, only - a `tracePath` into a sandbox deleted unless `--keep-temp` is passed. `report.html` does - show the judged text. +- **Dart syntax.** Native evals have no custom-code graders, so 30 cases across 11 skills + lost the fenced-block parse the previous harness did. The response text is not in + `aggregate-result.json`, only a `tracePath` into a sandbox deleted unless `--keep-temp` + is passed. `report.html` does show the judged text. - **Judge calibration.** Most graders are `llm` with no human-labelled gold set. -- **Tool execution.** The six tool-driven skills are graded on the decisions they narrate. - `claude plugin eval` can mock MCP servers under `evals/mocks//.md`, which - would grade them on the calls they actually make. Nothing here uses it yet. +- **Tool execution, mostly.** One case drives a mocked tool, + `license-compliance-runs-check-with-full-license-info`. The other tool-driven skills are + still graded on the calls they narrate. - **Stable routing.** Whether a skill activates is nondeterministic, which is why routing is a `tool_used` grader rather than inferred from content. -- **Prose in a `SKILL.md`.** Deliberate. An earlier version asserted a hundred `contains` +- **Prose in a `SKILL.md`.** Deliberate: an earlier version asserted a hundred `contains` patterns against skill bodies, so a copy-edit failed the gate. -- **The skills' own surfaces.** `skills_lint` catches a link pointing at a missing file and - a `name` that does not match its directory. Nothing catches a link resolving to the - *wrong* file, or an `allowed-tools` name that does not exist. Four invariants are - convention alone: `create-project` must not declare `Bash`, `green-gate` must declare it - for parsing `coverage/lcov.info`, `.mcp.json` must keep `--enable dart_format`, and +- **The skills' own surfaces.** `skills_lint` catches a link to a missing file and a `name` + that does not match its directory. Nothing catches a link resolving to the *wrong* file, + or an `allowed-tools` name that does not exist. Four invariants are convention alone: + `create-project` must not declare `Bash`, `green-gate` must declare it for parsing + `coverage/lcov.info`, `.mcp.json` must keep `--enable dart_format`, and `flutter-reviewer` must declare no write tools. diff --git a/evals/bloc/bloc-tests-with-bloc-test-and-mocktail/graders/build-returns-a-fresh-bloc.md b/evals/bloc/bloc-tests-with-bloc-test-and-mocktail/graders/build-returns-a-fresh-bloc.md new file mode 100644 index 0000000..96fb44e --- /dev/null +++ b/evals/bloc/bloc-tests-with-bloc-test-and-mocktail/graders/build-returns-a-fresh-bloc.md @@ -0,0 +1,12 @@ +--- +type: llm +weight: 3 +--- + +Judge only the `build:` callbacks of the `blocTest` calls. + +PASS if every `build:` constructs the bloc inside the callback, for example +`build: () => LoginBloc(authRepository: authRepository)` or `build: LoginBloc.new`. + +FAIL if any `build:` returns a bloc that was constructed outside the callback, for example +one assigned to a variable in `setUp` and referenced as `build: () => loginBloc`. diff --git a/evals/bloc/bloc-writes-sealed-events-and-states/prompt.md b/evals/bloc/bloc-writes-sealed-events-and-states/prompt.md index b759d8d..464c945 100644 --- a/evals/bloc/bloc-writes-sealed-events-and-states/prompt.md +++ b/evals/bloc/bloc-writes-sealed-events-and-states/prompt.md @@ -3,7 +3,7 @@ max_turns: 12 timeout_seconds: 600 allowed_tools: [Read, Glob, Grep, Skill] tags: [bloc] -description: "The house shape of a bloc: sealed event and state hierarchies with `final class` subclasses, Equatable with props, past-tense event names, and the state names SKILL.md pins for a login flow." +description: "The house shape of a bloc: states carrying different data select the subclass approach, so sealed hierarchies with `final class` subclasses, Equatable with props, past-tense event names, and the state names SKILL.md pins for a login flow." --- -Write a LoginBloc for email and password authentication with submit and logout events, and success and failure states. Output Dart code only. +Write a LoginBloc for email and password authentication, with a submit event and a logout event. While the request is in flight the button shows a spinner. On success the UI needs the signed-in User object; on failure it needs the error message to display. Nothing else is on screen before the first attempt. Output Dart code only. diff --git a/evals/create-project/create-project-asks-for-organization-when-required/graders/does-not-create-without-the-organization.md b/evals/create-project/create-project-asks-for-organization-when-required/graders/does-not-create-without-the-organization.md new file mode 100644 index 0000000..b6ce384 --- /dev/null +++ b/evals/create-project/create-project-asks-for-organization-when-required/graders/does-not-create-without-the-organization.md @@ -0,0 +1,7 @@ +--- +type: tool_used +tool: mcp__plugin_vgv-ai-flutter-plugin_very-good-cli__create +min: 0 +max: 0 +weight: 3 +--- diff --git a/evals/create-project/create-project-asks-when-the-template-is-ambiguous/graders/does-not-create-while-ambiguous.md b/evals/create-project/create-project-asks-when-the-template-is-ambiguous/graders/does-not-create-while-ambiguous.md new file mode 100644 index 0000000..b6ce384 --- /dev/null +++ b/evals/create-project/create-project-asks-when-the-template-is-ambiguous/graders/does-not-create-while-ambiguous.md @@ -0,0 +1,7 @@ +--- +type: tool_used +tool: mcp__plugin_vgv-ai-flutter-plugin_very-good-cli__create +min: 0 +max: 0 +weight: 3 +--- diff --git a/evals/create-project/create-project-does-not-over-ask/graders/creates-without-further-questions.md b/evals/create-project/create-project-does-not-over-ask/graders/creates-without-further-questions.md new file mode 100644 index 0000000..f088fd0 --- /dev/null +++ b/evals/create-project/create-project-does-not-over-ask/graders/creates-without-further-questions.md @@ -0,0 +1,6 @@ +--- +type: tool_used +tool: mcp__plugin_vgv-ai-flutter-plugin_very-good-cli__create +input_match: '"org_name"\s*:\s*"com\.example\.mystore"' +weight: 3 +--- diff --git a/evals/create-project/create-project-does-not-over-ask/graders/does-not-interrogate-for-optionals.md b/evals/create-project/create-project-does-not-over-ask/graders/does-not-interrogate-for-optionals.md index b4752eb..589385e 100644 --- a/evals/create-project/create-project-does-not-over-ask/graders/does-not-interrogate-for-optionals.md +++ b/evals/create-project/create-project-does-not-over-ask/graders/does-not-interrogate-for-optionals.md @@ -2,6 +2,8 @@ type: llm --- -PASS if the response proceeds with the name and organization it was given, without interrogating the user for optional parameters such as a description, an output directory, or an application id. Supplying its own value for an optional parameter satisfies this; so does handing the user a ready-to-run command, since no create tool is available in this environment. +PASS if the response proceeds with the name and organization it was given, without +interrogating the user for optional parameters such as a description, an output directory, +or an application id. Supplying its own value for an optional parameter satisfies this. FAIL only if it asks the user to supply more input before proceeding. diff --git a/evals/create-project/create-project-plans-dependency-install/graders/creates-the-project.md b/evals/create-project/create-project-plans-dependency-install/graders/creates-the-project.md new file mode 100644 index 0000000..fa68c2f --- /dev/null +++ b/evals/create-project/create-project-plans-dependency-install/graders/creates-the-project.md @@ -0,0 +1,5 @@ +--- +type: tool_used +tool: mcp__plugin_vgv-ai-flutter-plugin_very-good-cli__create +input_match: '"name"\s*:\s*"my_store"' +--- diff --git a/evals/create-project/create-project-plans-dependency-install/graders/installs-after-creating.md b/evals/create-project/create-project-plans-dependency-install/graders/installs-after-creating.md new file mode 100644 index 0000000..299354f --- /dev/null +++ b/evals/create-project/create-project-plans-dependency-install/graders/installs-after-creating.md @@ -0,0 +1,6 @@ +--- +type: tool_order +before: mcp__plugin_vgv-ai-flutter-plugin_very-good-cli__create +after: mcp__plugin_vgv-ai-flutter-plugin_very-good-cli__packages_get +weight: 3 +--- diff --git a/evals/create-project/create-project-plans-dependency-install/graders/names-the-flutter-app-template.md b/evals/create-project/create-project-plans-dependency-install/graders/names-the-flutter-app-template.md new file mode 100644 index 0000000..4215054 --- /dev/null +++ b/evals/create-project/create-project-plans-dependency-install/graders/names-the-flutter-app-template.md @@ -0,0 +1,5 @@ +--- +type: tool_used +tool: mcp__plugin_vgv-ai-flutter-plugin_very-good-cli__create +input_match: '"subcommand"\s*:\s*"flutter_app"' +--- diff --git a/evals/create-project/create-project-plans-dependency-install/graders/plan-installs-dependencies.md b/evals/create-project/create-project-plans-dependency-install/graders/plan-installs-dependencies.md deleted file mode 100644 index 063b2c4..0000000 --- a/evals/create-project/create-project-plans-dependency-install/graders/plan-installs-dependencies.md +++ /dev/null @@ -1,16 +0,0 @@ ---- -type: llm ---- - -The response describes the steps for creating a new Flutter project. Judge the steps it -describes, whether they are presented as work already done, as commands handed to the user -to run, or as steps that would run in an environment with shell access. A response that -says it cannot execute the steps itself is still judged on the steps it lists. - -PASS if the described steps include installing dependencies after the project is created, -for example a `very_good packages get` or `flutter pub get` step that follows the creation -step. - -FAIL if the described steps contain no dependency install after the creation step. It -does not matter what else follows creation: a plan that goes straight from creating the -project to running or editing it, with no install anywhere, fails. diff --git a/evals/create-project/create-project-plans-dependency-install/graders/scaffolds-through-very-good-cli.md b/evals/create-project/create-project-plans-dependency-install/graders/scaffolds-through-very-good-cli.md deleted file mode 100644 index 5b8daec..0000000 --- a/evals/create-project/create-project-plans-dependency-install/graders/scaffolds-through-very-good-cli.md +++ /dev/null @@ -1,16 +0,0 @@ ---- -type: llm ---- - -The response describes the steps for creating a new Flutter project. Judge the steps it -describes, whether they are presented as work already done, as commands handed to the user -to run, or as steps that would run in an environment with shell access. Naming a shell -command is acceptable here because no project-creation tool is available to call. - -PASS if the scaffolding step uses Very Good CLI (`very_good create`) and states an explicit -template name such as `flutter_app`. - -FAIL only if the scaffolding step the response actually settles on is the Flutter SDK's own -`flutter create`, or if it uses Very Good CLI without naming which template it creates. A -mention of `flutter create` as a fallback, an alternative, or a contrast does not fail this -grader when the scaffolding step it settles on is Very Good CLI with a named template. diff --git a/evals/create-project/create-project-plans-dependency-install/prompt.md b/evals/create-project/create-project-plans-dependency-install/prompt.md index 4b54fbf..85170c7 100644 --- a/evals/create-project/create-project-plans-dependency-install/prompt.md +++ b/evals/create-project/create-project-plans-dependency-install/prompt.md @@ -1,9 +1,9 @@ --- -max_turns: 12 -timeout_seconds: 600 +max_turns: 20 +timeout_seconds: 900 allowed_tools: [Read, Glob, Grep, Skill] tags: [create-project] description: The narrated plan scaffolds through Very Good CLI with an explicit template and does not stop at creation. --- -Create a Flutter app named my_store for organization com.example. Walk me through every step you will take. +Create a Flutter app named my_store for organization com.example. diff --git a/evals/create-project/create-project-scopes-dependency-install-to-the-new-project/prompt.md b/evals/create-project/create-project-scopes-dependency-install-to-the-new-project/prompt.md index ed831e2..9f190d2 100644 --- a/evals/create-project/create-project-scopes-dependency-install-to-the-new-project/prompt.md +++ b/evals/create-project/create-project-scopes-dependency-install-to-the-new-project/prompt.md @@ -6,4 +6,4 @@ tags: [create-project] description: The create-then-install order, with the install scoped to the created project via directory, plus the organization prompt. --- -I want to scaffold a new Flutter app named storefront, placed at apps/storefront inside my existing monorepo, and then install its dependencies. Which tools would you call, in order, and what arguments would you pass to each? Tell me anything you still need from me. Do not run anything yet. Very Good CLI is installed and on PATH, so treat its tools as available and ignore any startup notice saying otherwise. +I want to scaffold a new Flutter app named storefront, placed at apps/storefront inside my existing monorepo, and then install its dependencies. Which tools would you call, in order, and what arguments would you pass to each? Tell me anything you still need from me. Do not run anything yet. diff --git a/evals/create-project/create-project-stays-out-of-existing-project-work/graders/does-not-scaffold.md b/evals/create-project/create-project-stays-out-of-existing-project-work/graders/does-not-scaffold.md new file mode 100644 index 0000000..1b56ed4 --- /dev/null +++ b/evals/create-project/create-project-stays-out-of-existing-project-work/graders/does-not-scaffold.md @@ -0,0 +1,6 @@ +--- +type: tool_used +tool: mcp__plugin_vgv-ai-flutter-plugin_very-good-cli__create +min: 0 +max: 0 +--- diff --git a/evals/internationalization/internationalization-uses-directional-insets-for-rtl/graders/no-match-text-direction-on-icon.md b/evals/internationalization/internationalization-uses-directional-insets-for-rtl/graders/no-match-text-direction-on-icon.md index 87499aa..4bad5e9 100644 --- a/evals/internationalization/internationalization-uses-directional-insets-for-rtl/graders/no-match-text-direction-on-icon.md +++ b/evals/internationalization/internationalization-uses-directional-insets-for-rtl/graders/no-match-text-direction-on-icon.md @@ -2,4 +2,5 @@ type: regex pattern: 'Icon\(\s*Icons\.[A-Za-z0-9_]+,\s*matchTextDirection' match: not_contains +weight: 3 --- diff --git a/evals/layered-architecture/layered-architecture-transforms-models-in-the-repository/graders/declares-weather-repository.md b/evals/layered-architecture/layered-architecture-transforms-models-in-the-repository/graders/declares-weather-repository.md deleted file mode 100644 index 6aaaf9f..0000000 --- a/evals/layered-architecture/layered-architecture-transforms-models-in-the-repository/graders/declares-weather-repository.md +++ /dev/null @@ -1,4 +0,0 @@ ---- -type: regex -pattern: 'class WeatherRepository' ---- diff --git a/evals/layered-architecture/layered-architecture-wires-repositories-in-bootstrap/prompt.md b/evals/layered-architecture/layered-architecture-wires-repositories-in-bootstrap/prompt.md index 8fb0700..70c39cf 100644 --- a/evals/layered-architecture/layered-architecture-wires-repositories-in-bootstrap/prompt.md +++ b/evals/layered-architecture/layered-architecture-wires-repositories-in-bootstrap/prompt.md @@ -6,4 +6,4 @@ tags: [layered-architecture] description: "Bootstrap wiring: clients and repositories constructed in a main entrypoint, passed into App, provided to the tree, and depended on by path rather than by version." --- -Wire my AuthRepository and UserRepository into the app so every feature can reach them. Show the Dart code and the pubspec changes. +My monorepo has two local packages that nothing is wired to yet: `packages/auth_repository`, exposing `AuthRepository`, and `packages/user_repository`, exposing `UserRepository`. They are not in your working directory, so answer from this description rather than inspecting what is on disk. Wire them into the app so every feature can reach them. Show the Dart code and the pubspec changes. diff --git a/evals/license-compliance/license-compliance-refuses-to-clear-missing-licenses/graders/names-all-rights-reserved.md b/evals/license-compliance/license-compliance-refuses-to-clear-missing-licenses/graders/names-all-rights-reserved.md deleted file mode 100644 index f12bf84..0000000 --- a/evals/license-compliance/license-compliance-refuses-to-clear-missing-licenses/graders/names-all-rights-reserved.md +++ /dev/null @@ -1,5 +0,0 @@ ---- -type: regex -pattern: 'all rights reserved' -flags: i ---- diff --git a/evals/license-compliance/license-compliance-refuses-to-clear-missing-licenses/graders/rates-both-packages-high-risk.md b/evals/license-compliance/license-compliance-refuses-to-clear-missing-licenses/graders/rates-both-packages-high-risk.md new file mode 100644 index 0000000..03746b4 --- /dev/null +++ b/evals/license-compliance/license-compliance-refuses-to-clear-missing-licenses/graders/rates-both-packages-high-risk.md @@ -0,0 +1,11 @@ +--- +type: llm +weight: 3 +--- + +PASS if both unlicensed packages are given an explicit risk level and that level is the +highest the response uses, for example "High risk", "High", or "critical", whether in a +table column or in prose. + +FAIL if the response attaches no risk level to them, or rates them anything below its own +top level. diff --git a/evals/license-compliance/license-compliance-runs-check-with-full-license-info/graders/asks-for-full-license-info.md b/evals/license-compliance/license-compliance-runs-check-with-full-license-info/graders/asks-for-full-license-info.md new file mode 100644 index 0000000..1720779 --- /dev/null +++ b/evals/license-compliance/license-compliance-runs-check-with-full-license-info/graders/asks-for-full-license-info.md @@ -0,0 +1,5 @@ +--- +type: tool_used +tool: mcp__plugin_vgv-ai-flutter-plugin_very-good-cli__packages_check_licenses +input_match: '"licenses"\s*:\s*true' +--- diff --git a/evals/license-compliance/license-compliance-runs-check-with-full-license-info/graders/calls-the-license-check.md b/evals/license-compliance/license-compliance-runs-check-with-full-license-info/graders/calls-the-license-check.md new file mode 100644 index 0000000..0581340 --- /dev/null +++ b/evals/license-compliance/license-compliance-runs-check-with-full-license-info/graders/calls-the-license-check.md @@ -0,0 +1,5 @@ +--- +type: tool_used +tool: mcp__plugin_vgv-ai-flutter-plugin_very-good-cli__packages_check_licenses +weight: 3 +--- diff --git a/evals/license-compliance/license-compliance-runs-check-with-full-license-info/graders/describes-the-compliance-report.md b/evals/license-compliance/license-compliance-runs-check-with-full-license-info/graders/describes-the-compliance-report.md index e3c52b2..90ab7a1 100644 --- a/evals/license-compliance/license-compliance-runs-check-with-full-license-info/graders/describes-the-compliance-report.md +++ b/evals/license-compliance/license-compliance-runs-check-with-full-license-info/graders/describes-the-compliance-report.md @@ -2,6 +2,10 @@ type: llm --- -PASS if the response describes the written report the audit will produce, and that description has both of these: flagged dependencies held in their own section or table, separate from the compliant ones, and a risk level attached to every flagged dependency (a `Risk` column, or wording such as high/medium/low risk per flagged package). A blank template, a skeleton with placeholder rows, or a plain statement of what the report will contain all count. The response is describing a report it has not run yet, so it does not have to name real packages or fill in real risk levels. +PASS if the report keeps flagged dependencies in their own section or table, separate from +the compliant ones, and attaches a risk level to every flagged dependency, whether as a +`Risk` column or as wording such as high or medium risk per package. -FAIL only if one of those two is missing: the response describes no report at all (for example it promises only to "list the licenses" or to return a pass/fail verdict), the report it describes keeps flagged and compliant dependencies in one undifferentiated list, or the report it describes attaches no risk level to flagged dependencies. +FAIL if flagged and compliant dependencies sit in one undifferentiated list, if the answer +is only a pass/fail verdict or a count, or if no risk level is attached to the flagged +dependencies. diff --git a/evals/license-compliance/license-compliance-runs-check-with-full-license-info/graders/names-check-licenses.md b/evals/license-compliance/license-compliance-runs-check-with-full-license-info/graders/names-check-licenses.md deleted file mode 100644 index 8dc219d..0000000 --- a/evals/license-compliance/license-compliance-runs-check-with-full-license-info/graders/names-check-licenses.md +++ /dev/null @@ -1,4 +0,0 @@ ---- -type: regex -pattern: 'packages[ _]check[ _]licenses' ---- diff --git a/evals/license-compliance/license-compliance-runs-check-with-full-license-info/graders/reports-the-flagged-packages.md b/evals/license-compliance/license-compliance-runs-check-with-full-license-info/graders/reports-the-flagged-packages.md new file mode 100644 index 0000000..9e23be2 --- /dev/null +++ b/evals/license-compliance/license-compliance-runs-check-with-full-license-info/graders/reports-the-flagged-packages.md @@ -0,0 +1,12 @@ +--- +type: llm +--- + +The scan flagged two dependencies: `glyph_atlas` under GPL-3.0 and `refresh_indicator_x` +with no license detected. + +PASS if the report names both of them as needing attention and attaches a risk level to +each, whether in a `Risk` column or in prose. + +FAIL if either is missing, if either is presented as compliant, or if no risk level is +attached to them. diff --git a/evals/license-compliance/license-compliance-runs-check-with-full-license-info/graders/requests-full-license-info.md b/evals/license-compliance/license-compliance-runs-check-with-full-license-info/graders/requests-full-license-info.md deleted file mode 100644 index 808e8ee..0000000 --- a/evals/license-compliance/license-compliance-runs-check-with-full-license-info/graders/requests-full-license-info.md +++ /dev/null @@ -1,4 +0,0 @@ ---- -type: regex -pattern: 'licenses[^\n]{0,24}\btrue\b' ---- diff --git a/evals/license-compliance/license-compliance-runs-check-with-full-license-info/graders/retrieves-every-dependency-license.md b/evals/license-compliance/license-compliance-runs-check-with-full-license-info/graders/retrieves-every-dependency-license.md deleted file mode 100644 index f813fc0..0000000 --- a/evals/license-compliance/license-compliance-runs-check-with-full-license-info/graders/retrieves-every-dependency-license.md +++ /dev/null @@ -1,7 +0,0 @@ ---- -type: llm ---- - -PASS if the response says it will retrieve the full license information for every dependency. - -FAIL if it offers only a pass/fail verdict or a count of packages. diff --git a/evals/license-compliance/license-compliance-runs-check-with-full-license-info/prompt.md b/evals/license-compliance/license-compliance-runs-check-with-full-license-info/prompt.md index 2815e9d..e51c130 100644 --- a/evals/license-compliance/license-compliance-runs-check-with-full-license-info/prompt.md +++ b/evals/license-compliance/license-compliance-runs-check-with-full-license-info/prompt.md @@ -1,9 +1,9 @@ --- -max_turns: 12 -timeout_seconds: 600 +max_turns: 20 +timeout_seconds: 900 allowed_tools: [Read, Glob, Grep, Skill] tags: [license-compliance] -description: The audit is the license check with full license information, and the deliverable is the prescribed risk-rated report. +description: Runs the license check with full license information and turns the scan output into the prescribed risk-rated report. --- -We ship this Flutter app next week and I need a license audit of its dependencies first. You can't run anything in this session, so tell me exactly what you would call, with what arguments, and what the finished audit will contain. +We ship this Flutter app next week and I need a license audit of its dependencies first. Run the audit and give me the finished report. diff --git a/evals/license-compliance/license-compliance-scopes-check-to-monorepo-subdirectory/graders/directory-points-at-mobile.md b/evals/license-compliance/license-compliance-scopes-check-to-monorepo-subdirectory/graders/directory-points-at-mobile.md deleted file mode 100644 index 8315277..0000000 --- a/evals/license-compliance/license-compliance-scopes-check-to-monorepo-subdirectory/graders/directory-points-at-mobile.md +++ /dev/null @@ -1,4 +0,0 @@ ---- -type: regex -pattern: 'directory[^\n]{0,24}mobile' ---- diff --git a/evals/license-compliance/license-compliance-scopes-check-to-monorepo-subdirectory/graders/names-check-licenses.md b/evals/license-compliance/license-compliance-scopes-check-to-monorepo-subdirectory/graders/names-check-licenses.md deleted file mode 100644 index 8dc219d..0000000 --- a/evals/license-compliance/license-compliance-scopes-check-to-monorepo-subdirectory/graders/names-check-licenses.md +++ /dev/null @@ -1,4 +0,0 @@ ---- -type: regex -pattern: 'packages[ _]check[ _]licenses' ---- diff --git a/evals/license-compliance/license-compliance-scopes-check-to-monorepo-subdirectory/graders/scoped-to-the-app-subdirectory.md b/evals/license-compliance/license-compliance-scopes-check-to-monorepo-subdirectory/graders/scoped-to-the-app-subdirectory.md deleted file mode 100644 index 0c384f0..0000000 --- a/evals/license-compliance/license-compliance-scopes-check-to-monorepo-subdirectory/graders/scoped-to-the-app-subdirectory.md +++ /dev/null @@ -1,7 +0,0 @@ ---- -type: llm ---- - -PASS if the license check is scoped to the mobile subdirectory by passing a directory argument naming it. - -FAIL if it runs the check at the repository root, or names no directory at all. diff --git a/evals/license-compliance/license-compliance-scopes-check-to-monorepo-subdirectory/graders/scopes-the-check-to-mobile.md b/evals/license-compliance/license-compliance-scopes-check-to-monorepo-subdirectory/graders/scopes-the-check-to-mobile.md new file mode 100644 index 0000000..e2b2208 --- /dev/null +++ b/evals/license-compliance/license-compliance-scopes-check-to-monorepo-subdirectory/graders/scopes-the-check-to-mobile.md @@ -0,0 +1,6 @@ +--- +type: tool_used +tool: mcp__plugin_vgv-ai-flutter-plugin_very-good-cli__packages_check_licenses +input_match: '"directory"\s*:\s*"[^"]*mobile' +weight: 3 +--- diff --git a/evals/license-compliance/license-compliance-scopes-check-to-monorepo-subdirectory/prompt.md b/evals/license-compliance/license-compliance-scopes-check-to-monorepo-subdirectory/prompt.md index bf04ef6..3df3257 100644 --- a/evals/license-compliance/license-compliance-scopes-check-to-monorepo-subdirectory/prompt.md +++ b/evals/license-compliance/license-compliance-scopes-check-to-monorepo-subdirectory/prompt.md @@ -1,11 +1,9 @@ --- -max_turns: 12 -timeout_seconds: 600 +max_turns: 20 +timeout_seconds: 900 allowed_tools: [Read, Glob, Grep, Skill] tags: [license-compliance] -description: The Core Standard that a project below the workspace root needs a directory argument. +description: The license check is scoped to the app subdirectory of a monorepo rather than run at the repo root. --- -Assume this repository layout, do not inspect the working directory, and do not run anything: -melos.yaml at the repo root, the Flutter app at mobile/ with its own pubspec.yaml, and shared packages under packages/. -I want the app's dependency licenses audited. What is the exact tool call, with every argument? +Assume this repository layout and answer from it rather than inspecting the working directory: melos.yaml at the repo root, the Flutter app at mobile/ with its own pubspec.yaml, and shared packages under packages/. Audit the app's dependency licenses. diff --git a/evals/mocks/very-good-cli/_tools.json b/evals/mocks/very-good-cli/_tools.json new file mode 100644 index 0000000..032c6bd --- /dev/null +++ b/evals/mocks/very-good-cli/_tools.json @@ -0,0 +1,205 @@ +{ + "tools": [ + { + "name": "create", + "description": "Create a very good Dart or Flutter project in seconds based on the provided template. Each template has a corresponding sub-command.\n ", + "inputSchema": { + "type": "object", + "properties": { + "subcommand": { + "type": "string", + "description": "The available subcommands to provide an specific template, are:\napp_ui_package - Generate a Very Good App UI package.\ndart_cli - Generate a Very Good Dart CLI application.\ndart_package - Generate a Very Good Dart package.\ndocs_site - Generate a Very Good documentation site.\nflame_game - Generate a Very Good Flame game.\nflutter_app - Generate a Very Good Flutter application.\nflutter_package - Generate a Very Good Flutter package.\nflutter_plugin - Generate a Very Good Flutter plugin.\n", + "enum": [ + "app_ui_package", + "flame_game", + "flutter_app", + "flutter_package", + "flutter_plugin", + "dart_cli", + "dart_package", + "docs_site" + ] + }, + "name": { + "type": "string", + "description": "Project name" + }, + "description": { + "type": "string", + "description": "The description for this new project.\n(defaults to \"A Very Good Project created by Very Good CLI.\")" + }, + "org_name": { + "type": "string", + "description": "The organization for this new project.\n(defaults to \"com.example.verygoodcore\")" + }, + "output_directory": { + "type": "string", + "description": "The desired output directory when creating a new project." + }, + "application_id": { + "type": "string", + "description": "The bundle identifier on iOS or application id on Android. (defaults to .)" + }, + "platforms": { + "type": "string", + "description": "Comma-separated platforms. 'Example: \"android,ios,web\".\nThe values for platforms are: android, ios, web, macos, linux, and windows.\nIf is omitted, then all platforms are enabled by default.\nOnly available for subcommands: flutter_plugin with all values) and flame_game (only android and ios)\n " + }, + "publishable": { + "type": "boolean", + "description": "Whether package is intended for publishing (flutter_package, dart_package only)" + }, + "workspace": { + "type": "boolean", + "description": "Whether the generated project should resolve its dependencies from a parent Pub workspace." + }, + "executable-name": { + "type": "string", + "description": "CLI custom executable name (dart_cli only)" + }, + "template": { + "type": "string", + "description": "The template used to generate this new project.\nThe values are:\ncore - Generate a Very Good Flutter application.\nIf is omitted, then core will be selected.\n" + } + }, + "required": [ + "subcommand", + "name" + ] + } + }, + { + "name": "test", + "description": "Run tests in a Dart or Flutter project.", + "inputSchema": { + "type": "object", + "properties": { + "directory": { + "type": "string", + "description": "The package to test (defaults to current directory). Can be absolute or relative path to project root. A package path belongs here, not in \"paths\"." + }, + "paths": { + "type": "array", + "description": "Test files or directories to run, relative to the package selected by \"directory\" (e.g. ['test/src/foo_test.dart', 'test/widgets']). These are targets inside that one package: to test a different package, set \"directory\" to it rather than putting its path here. When omitted, the whole suite runs. Cannot be combined with \"recursive\". Note that targeting specific paths disables the test optimization step.", + "items": { + "type": "string" + } + }, + "dart": { + "type": "boolean", + "description": "Whether to run Dart tests. If not specified, Flutter tests will be run if a Flutter project is detected." + }, + "coverage": { + "type": "boolean", + "description": "Whether to collect coverage information." + }, + "recursive": { + "type": "boolean", + "description": "Run tests recursively for all nested packages. Cannot be combined with \"paths\", because each package runs in its own working directory and a path is only meaningful within one of them." + }, + "optimization": { + "type": "boolean", + "description": "Whether to apply optimizations for test performance.\nAutomatically disabled when --platform is specified.\nAdd the `skip_very_good_optimization` tag to specific test files to disable them individually.\n(defaults to on)" + }, + "concurrency": { + "type": "string", + "description": "The number of concurrent test suites run.\nAutomatically set to 1 when --platform is specified.\n(defaults to \"4\")" + }, + "tags": { + "type": "string", + "description": "Run only tests associated with the specified tags." + }, + "exclude_coverage": { + "type": "string", + "description": "A glob which will be used to exclude files that match from the coverage (e.g. '**/*.g.dart')." + }, + "exclude_tags": { + "type": "string", + "description": "Run only tests that do not have the specified tags." + }, + "min_coverage": { + "type": "string", + "description": "Whether to enforce a minimum coverage percentage." + }, + "test_randomize_ordering_seed": { + "type": "string", + "description": "The seed to randomize the execution order of test cases within test files." + }, + "update_goldens": { + "type": "boolean", + "description": "Whether \"matchesGoldenFile()\" calls within your test methods should update the golden files." + }, + "force_ansi": { + "type": "boolean", + "description": "Whether to force ansi output. If not specified, it will maintain the default behavior based on stdout and stderr." + }, + "dart-define": { + "type": "string", + "description": "Additional key-value pairs that will be available as constants from the String.fromEnvironment, bool.fromEnvironment, int.fromEnvironment, and double.fromEnvironment constructors.\nMultiple defines can be passed by repeating \"--dart-define\" multiple times.\n(e.g., foo=bar)\n" + }, + "dart-define-from-file": { + "type": "string", + "description": "The path of a .json or .env file containing key-value pairs that will be available as environment variables. These can be accessed using the String.fromEnvironment, bool.fromEnvironment, and int.fromEnvironment constructors.\nMultiple defines can be passed by repeating \"--dart-define-from-file\" multiple times. Entries from \"--dart-define\" with identical keys take precedence over entries from these files." + }, + "platform": { + "type": "string", + "description": "The platform to run tests on.\nThe available values are: chrome, vm, android, ios.\nOnly one value can be selected.\n " + }, + "run_skipped": { + "type": "boolean", + "description": "Run skipped tests instead of skipping them. Only applies to Dart tests (dart: true)." + }, + "check_ignore": { + "type": "boolean", + "description": "Whether to check for and respect coverage ignore comments (e.g. // coverage:ignore-line). Only applies to Dart tests (dart: true)." + }, + "show_uncovered": { + "type": "boolean", + "description": "Whether to list uncovered lines when coverage is below 100%. Implicitly enables coverage collection when used alone. Useful for identifying which lines still need tests after a min_coverage failure." + }, + "timeout_seconds": { + "type": "integer", + "description": "Maximum seconds to wait for the test run before killing the Flutter test process. Flutter tests can hang indefinitely when pumpAndSettle() is called without a timeout argument. When omitted, no timeout is applied." + } + } + } + }, + { + "name": "packages_get", + "description": " Install or update a Dart/Flutter package dependencies.\n Use after creating a project or modifying pubspec.yaml.\n Supports recursive installation and package exclusion.", + "inputSchema": { + "type": "object", + "properties": { + "directory": { + "type": "string", + "description": "Target directory path (defaults to current directory). Can be absolute or relative path to project root." + }, + "recursive": { + "type": "boolean", + "description": "Install dependencies for all nested packages recursively. Useful for monorepos or projects with multiple packages." + }, + "ignore": { + "type": "string", + "description": "Comma-separated list of package names to skip. Example: \"package1,package2\". Useful to avoid processing problematic packages." + } + } + } + }, + { + "name": "packages_check_licenses", + "description": " Verify package licenses for compliance and validation in a Dart or Flutter project.\n Identifies license types (MIT, BSD, Apache, etc.) for all\n dependencies. Use to ensure license compatibility.", + "inputSchema": { + "type": "object", + "properties": { + "directory": { + "type": "string", + "description": "Target directory path (defaults to current directory). Path to the project root containing pubspec.yaml." + }, + "licenses": { + "type": "boolean", + "description": "Verify all package licenses (defaults to true). Currently the only supported check type. Reports license types for all dependencies." + } + } + } + } + ] +} diff --git a/evals/mocks/very-good-cli/create.md b/evals/mocks/very-good-cli/create.md new file mode 100644 index 0000000..6ac1b3b --- /dev/null +++ b/evals/mocks/very-good-cli/create.md @@ -0,0 +1,8 @@ +--- +expect: + subcommand: [app_ui_package, flame_game, flutter_app, flutter_package, flutter_plugin, dart_cli, dart_package, docs_site] + name: string +--- + +✓ Generated 87 file(s) from the {{input.subcommand}} template. +Created {{input.name}}. diff --git a/evals/mocks/very-good-cli/packages_check_licenses.md b/evals/mocks/very-good-cli/packages_check_licenses.md new file mode 100644 index 0000000..bc1514a --- /dev/null +++ b/evals/mocks/very-good-cli/packages_check_licenses.md @@ -0,0 +1,20 @@ +Retrieved 14 licenses from 14 packages. + +| package | version | license | +| ------------------------ | ------- | ------------ | +| async | 2.13.0 | BSD-3-Clause | +| characters | 1.4.1 | BSD-3-Clause | +| collection | 1.19.1 | BSD-3-Clause | +| cupertino_icons | 1.0.8 | MIT | +| glyph_atlas | 0.7.1 | GPL-3.0 | +| http | 1.5.0 | BSD-3-Clause | +| material_color_utilities | 0.13.0 | Apache-2.0 | +| meta | 1.17.0 | BSD-3-Clause | +| path | 1.9.1 | BSD-3-Clause | +| refresh_indicator_x | 2.4.0 | unknown | +| source_span | 1.10.1 | BSD-3-Clause | +| string_scanner | 1.4.1 | BSD-3-Clause | +| term_glyph | 1.2.2 | BSD-3-Clause | +| vector_math | 2.2.0 | BSD-3-Clause | + +License types: BSD-3-Clause (10), MIT (1), Apache-2.0 (1), GPL-3.0 (1), unknown (1). diff --git a/evals/mocks/very-good-cli/packages_get.md b/evals/mocks/very-good-cli/packages_get.md new file mode 100644 index 0000000..05add4f --- /dev/null +++ b/evals/mocks/very-good-cli/packages_get.md @@ -0,0 +1,2 @@ +Running "flutter pub get"... +✓ Retrieved dependencies (2.1s) diff --git a/evals/mocks/very-good-cli/test.md b/evals/mocks/very-good-cli/test.md new file mode 100644 index 0000000..64868ed --- /dev/null +++ b/evals/mocks/very-good-cli/test.md @@ -0,0 +1,8 @@ +Running "flutter test"... +00:04 +38: All tests passed! + +Collected coverage to coverage/lcov.info. +lines......: 94.2% (356 of 378 lines) + +lib/src/login/login_bloc.dart 81.0% 17 of 21 lines +lib/src/settings/settings_page.dart 72.7% 8 of 11 lines diff --git a/evals/testing/testing-stays-out-of-non-test-work/graders/does-not-run-tests.md b/evals/testing/testing-stays-out-of-non-test-work/graders/does-not-run-tests.md new file mode 100644 index 0000000..6a8498f --- /dev/null +++ b/evals/testing/testing-stays-out-of-non-test-work/graders/does-not-run-tests.md @@ -0,0 +1,6 @@ +--- +type: tool_used +tool: mcp__plugin_vgv-ai-flutter-plugin_very-good-cli__test +min: 0 +max: 0 +--- diff --git a/evals/testing/testing-uses-pump-app-in-widget-tests/graders/uses-pump-app.md b/evals/testing/testing-uses-pump-app-in-widget-tests/graders/uses-pump-app.md deleted file mode 100644 index 4a94320..0000000 --- a/evals/testing/testing-uses-pump-app-in-widget-tests/graders/uses-pump-app.md +++ /dev/null @@ -1,4 +0,0 @@ ---- -type: regex -pattern: 'pumpApp' ---- diff --git a/evals/testing/testing-uses-pump-app-in-widget-tests/graders/uses-test-widgets.md b/evals/testing/testing-uses-pump-app-in-widget-tests/graders/uses-test-widgets.md deleted file mode 100644 index 9f07bb0..0000000 --- a/evals/testing/testing-uses-pump-app-in-widget-tests/graders/uses-test-widgets.md +++ /dev/null @@ -1,4 +0,0 @@ ---- -type: regex -pattern: 'testWidgets' ---- diff --git a/evals/testing/testing-uses-pump-app-in-widget-tests/graders/wraps-with-pump-app.md b/evals/testing/testing-uses-pump-app-in-widget-tests/graders/wraps-with-pump-app.md index ce78fff..8be4adb 100644 --- a/evals/testing/testing-uses-pump-app-in-widget-tests/graders/wraps-with-pump-app.md +++ b/evals/testing/testing-uses-pump-app-in-widget-tests/graders/wraps-with-pump-app.md @@ -1,7 +1,11 @@ --- type: llm +weight: 3 --- -PASS if the test wraps the widget with a pumpApp helper. +PASS if the test puts the widget on screen through a `pumpApp` helper on the tester, for +example `await tester.pumpApp(const LoginView());`. -FAIL if it calls pumpWidget(MaterialApp(...)) inline in the test body. +FAIL if it calls `pumpWidget` directly in the test body, including +`await tester.pumpWidget(MaterialApp(home: LoginView()));`, or defines its own inline +wrapper instead of using the helper. diff --git a/evals/testing/testing-uses-pump-app-in-widget-tests/prompt.md b/evals/testing/testing-uses-pump-app-in-widget-tests/prompt.md index 61ee1c5..fbb4eb6 100644 --- a/evals/testing/testing-uses-pump-app-in-widget-tests/prompt.md +++ b/evals/testing/testing-uses-pump-app-in-widget-tests/prompt.md @@ -6,4 +6,23 @@ tags: [testing] description: "Widget tests wrap through the shared `pumpApp` helper and drive the view with a MockBloc or MockCubit rather than a real bloc." --- -Write a widget test for a LoginView that shows an error SnackBar when the bloc emits LoginFailure, using our existing widget-test helpers. Output Dart code only. +Write a widget test for `LoginView`. It should assert that a SnackBar carrying the error +message appears when the bloc emits a failure state. Our project already has the usual +shared widget-test helpers available under `test/helpers/`. Output Dart code only. + +These are the types, and this description is authoritative over anything on disk: + +```dart +class LoginView extends StatelessWidget { const LoginView({super.key}); } + +class LoginBloc extends Bloc { ... } + +sealed class LoginState extends Equatable { const LoginState(); } + +final class LoginInitial extends LoginState {} + +final class LoginFailure extends LoginState { + const LoginFailure(this.errorMessage); + final String errorMessage; +} +``` diff --git a/evals/ui-package/ui-package-scaffolds-with-app-ui-package-template/graders/names-very-good-cli.md b/evals/ui-package/ui-package-scaffolds-with-app-ui-package-template/graders/names-very-good-cli.md deleted file mode 100644 index 56ae305..0000000 --- a/evals/ui-package/ui-package-scaffolds-with-app-ui-package-template/graders/names-very-good-cli.md +++ /dev/null @@ -1,4 +0,0 @@ ---- -type: regex -pattern: '[Vv]ery[-_ ][Gg]ood' ---- diff --git a/evals/ui-package/ui-package-scaffolds-with-app-ui-package-template/graders/places-it-in-the-monorepo.md b/evals/ui-package/ui-package-scaffolds-with-app-ui-package-template/graders/places-it-in-the-monorepo.md new file mode 100644 index 0000000..16fba6c --- /dev/null +++ b/evals/ui-package/ui-package-scaffolds-with-app-ui-package-template/graders/places-it-in-the-monorepo.md @@ -0,0 +1,5 @@ +--- +type: tool_used +tool: mcp__plugin_vgv-ai-flutter-plugin_very-good-cli__create +input_match: '"output_directory"\s*:' +--- diff --git a/evals/ui-package/ui-package-scaffolds-with-app-ui-package-template/graders/scaffolds-from-template.md b/evals/ui-package/ui-package-scaffolds-with-app-ui-package-template/graders/scaffolds-from-template.md deleted file mode 100644 index ae8eb07..0000000 --- a/evals/ui-package/ui-package-scaffolds-with-app-ui-package-template/graders/scaffolds-from-template.md +++ /dev/null @@ -1,7 +0,0 @@ ---- -type: llm ---- - -PASS if the package is scaffolded from the Very Good CLI `app_ui_package` template. - -FAIL if the response builds the package by hand, or reaches for the Flutter SDK's own `flutter create --template=package`. diff --git a/evals/ui-package/ui-package-scaffolds-with-app-ui-package-template/graders/scaffolds-with-the-template.md b/evals/ui-package/ui-package-scaffolds-with-app-ui-package-template/graders/scaffolds-with-the-template.md new file mode 100644 index 0000000..b2bf833 --- /dev/null +++ b/evals/ui-package/ui-package-scaffolds-with-app-ui-package-template/graders/scaffolds-with-the-template.md @@ -0,0 +1,6 @@ +--- +type: tool_used +tool: mcp__plugin_vgv-ai-flutter-plugin_very-good-cli__create +input_match: '"subcommand"\s*:\s*"app_ui_package"' +weight: 3 +--- diff --git a/evals/ui-package/ui-package-scaffolds-with-app-ui-package-template/graders/src-and-barrel-layout.md b/evals/ui-package/ui-package-scaffolds-with-app-ui-package-template/graders/src-and-barrel-layout.md deleted file mode 100644 index 97d79ec..0000000 --- a/evals/ui-package/ui-package-scaffolds-with-app-ui-package-template/graders/src-and-barrel-layout.md +++ /dev/null @@ -1,7 +0,0 @@ ---- -type: llm ---- - -PASS if the described layout keeps widget and theme source files under `lib/src/` and exposes the public API through a single barrel file directly under `lib/`. - -FAIL if the layout puts public widgets at the top of `lib/` with no barrel file, or names no barrel file at all, or if the widget and theme source files sit anywhere other than under `lib/src/`. diff --git a/evals/ui-package/ui-package-scaffolds-with-app-ui-package-template/graders/uses-app-ui-package-template.md b/evals/ui-package/ui-package-scaffolds-with-app-ui-package-template/graders/uses-app-ui-package-template.md deleted file mode 100644 index 1f32df5..0000000 --- a/evals/ui-package/ui-package-scaffolds-with-app-ui-package-template/graders/uses-app-ui-package-template.md +++ /dev/null @@ -1,4 +0,0 @@ ---- -type: regex -pattern: 'app_ui_package' ---- diff --git a/evals/ui-package/ui-package-scaffolds-with-app-ui-package-template/prompt.md b/evals/ui-package/ui-package-scaffolds-with-app-ui-package-template/prompt.md index 9e6a514..259962c 100644 --- a/evals/ui-package/ui-package-scaffolds-with-app-ui-package-template/prompt.md +++ b/evals/ui-package/ui-package-scaffolds-with-app-ui-package-template/prompt.md @@ -1,9 +1,9 @@ --- -max_turns: 12 -timeout_seconds: 600 +max_turns: 20 +timeout_seconds: 900 allowed_tools: [Read, Glob, Grep, Skill] tags: [ui-package] description: "Scaffolding through the Very Good CLI app_ui_package template, and the lib/src + single-barrel layout that template ships." --- -Set up a new UI package called storefront_ui in my monorepo for our shared widgets and design tokens. Which tool or command would you use, with which arguments, and what will the package layout look like? Do not run anything yet. +Set up a new UI package called storefront_ui in my monorepo for our shared widgets and design tokens, and show me what the package layout will look like. diff --git a/evals/ui-package/ui-package-stays-out-of-plain-dart-work/graders/does-not-scaffold.md b/evals/ui-package/ui-package-stays-out-of-plain-dart-work/graders/does-not-scaffold.md new file mode 100644 index 0000000..1b56ed4 --- /dev/null +++ b/evals/ui-package/ui-package-stays-out-of-plain-dart-work/graders/does-not-scaffold.md @@ -0,0 +1,6 @@ +--- +type: tool_used +tool: mcp__plugin_vgv-ai-flutter-plugin_very-good-cli__create +min: 0 +max: 0 +--- diff --git a/skills/bloc/SKILL.md b/skills/bloc/SKILL.md index ea8c21b..ff45cb3 100644 --- a/skills/bloc/SKILL.md +++ b/skills/bloc/SKILL.md @@ -27,6 +27,7 @@ Apply these standards to ALL Bloc/Cubit work: - **Use `blocTest()` from `package:bloc_test`** for all Bloc and Cubit tests — never raw `test()` with manual stream assertions - **Use `package:mocktail` for mocking** — never `package:mockito` +- **`build:` constructs the bloc under test** — return a new instance from the callback, never one built in `setUp` and shared, so every `blocTest` starts from a clean bloc and its lifecycle stays with `blocTest` - **No bloc-to-bloc direct dependencies** — blocs communicate through the UI or shared repositories - **Page/View separation** — Page provides the Bloc/Cubit via `BlocProvider`, View consumes via `BlocBuilder`/`BlocListener` - **Sealed classes for events and multi-state types** — enables exhaustive pattern matching with Dart 3 `switch`