Add experimental Claude Code support and reproducible native evaluations - #1
Merged
Merged
Conversation
Packaging, relative links, declared versions and agreement between the README result tables and the published artifacts were checked by hand. Run them as one script and in GitHub Actions, together with the study self-tests and report regeneration that need no container, model or credentials. Each installable skill now declares its own version in a VERSION file, so a copied package states what it is, and CHANGELOG.md reconstructs the per-skill history from the commits and dated reports each version shipped with. The repository tags name core releases only; the adoption skill moves separately. These are mechanical repository checks. Every entry records what was actually executed for a version, keeping static validation separate from agent evaluation as CONTRIBUTING.md requires. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The maintenance rule and marker went to AGENTS.md by name. A client may not read that file: Claude Code loads a project's AGENTS.md only when no CLAUDE.md exists in the working directory or above it, so a rule written beside a CLAUDE.md is never loaded. Adoption then reports success while the ongoing mechanism it establishes does nothing. Identify the instruction file in effect instead of assuming a name, and when a project keeps several, keep the rule in one of them and reachable from the others by their own reference or import rather than duplicating guidance. State the rule without a client-specific invocation prefix, so the wording written into a project stays valid wherever the skill is installed. Add two instruction-file routing cases with ten static grader controls, covering a CLAUDE.md-only project and a project holding both files. Only those controls ran: no model session has been executed for adoption v0.1.3, and the README, changelog and case notes say so. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…text Five completed studies report no skill-specific advantage. Three properties of the current designs would produce that tie from any skill: the shared task text already states much of the skill's guidance to both arms, the fixture evidence labels its own approval and observation status, and small explicit inputs let every reader answer from raw sources. More trials under these conditions would cost more and tie again. Record the proposed changes rather than editing any frozen fixture or published result: a neutral baseline request, evidence that does not classify itself, a real retrieval cost, and discovery treated as its own factor, which is the one place existing evidence already suggests an effect. The protocol is proposed and not owner-approved; it adds no result. Update project context for the static checks, Claude Code packaging, adoption v0.1.3 and the delegated contribution authority, keeping observed, approved and proposed statements separate. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Every published study ran on Codex, and the v0.1.3 follow-up records native Claude Code behavior as unestablished. This harness checks the four questions that review raised on the other client: skill discovery and invocation, rule placement in Claude-only and mixed instruction projects, repeat passes that preserve files and the original adoption date, and audit-only requests leaving the project byte-identical. Fixtures come from the shared adoption cases, so both clients exercise the same projects, protected content and immutable files, with each immutable file stated explicitly in the request as the initial v0.1.3 review required. Three things differ by necessity and are recorded in the protocol: slash-command or unnamed invocation, installation into the project's own .claude/skills, and an audit that reports in its final message because this client has no /output directory. Installed packages are committed with the fixture and hash-checked after every session, so a reference to an installed file validates while tampering fails. Invocation is measured from the session's own streamed tool calls, following the routing study's exposure method. Mechanical checks and invocation are recorded separately from semantics, which must be supplied as explicit reviews. Executed here: 34 stream, fixture and grader assertions, now in CI, plus a discarded dry run that confirmed the failure path records execution errors rather than silently passing. No model session ran: the available Claude Code CLI could not authenticate, so the six sessions remain outstanding and this establishes no behavior on that client. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The smoke-test harness raises choices that belong to the evaluation design rather than to one session: where the skills are installed, where an audit reports, whether unnamed skill selection gates acceptance, whether this isolation is sufficient without a container, which model to pin, and how the client's AGENTS.md version gate affects the split-file wording. Each is recorded with a recommendation so the collaborator settling them has the reasons, not only the conclusion. The report also states what was executed and what was not: every static check and its count, and six model sessions that did not run because the available Claude Code CLI could not authenticate. It accepts both findings from the v0.1.3 review as fixture and verifier defects, and keeps the open absolute-path portability concern visible rather than treating the project-scoped install as a fix in the skill. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
jaredchu
marked this pull request as ready for review
September 28, 2026 07:30
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Adoption assumed
AGENTS.md, which could leave the maintenance rule in a file Claude Code did not load. Route adoption to the active instruction file, document native installation, and tighten evidence/authority guidance in core 0.1.2 and adoption 0.1.4.Add CI checks for package boundaries, links, versions and published reports, plus native evaluation runners with frozen inputs, retained failures and separate semantic review. The latest evaluation observes initial instruction loading and guards a restricted set of file tools, including a fresh read-only session after adoption.
Validation:
Recommend experimental merge with the limits documented. The guard is evaluation-only, excludes shell execution and is not an OS sandbox; installing the skills does not install confinement. No general reliability or accuracy advantage is claimed.
Evidence: guarded follow-up, paired comparison.