Skip to content

Add experimental Claude Code support and reproducible native evaluations - #1

Merged
jaredchu merged 33 commits into
mainfrom
claude/claude-support-and-static-checks
Sep 28, 2026
Merged

jaredchu merged 33 commits into
mainfrom
claude/claude-support-and-static-checks

Conversation

@jaredchu

Copy link
Copy Markdown
Owner

Adoption assumed AGENTS.md, which could leave the maintenance rule in a file Claude Code did not load. Route adoption to the active instruction file, document native installation, and tighten evidence/authority guidance in core 0.1.2 and adoption 0.1.4.

Add CI checks for package boundaries, links, versions and published reports, plus native evaluation runners with frozen inputs, retained failures and separate semantic review. The latest evaluation observes initial instruction loading and guards a restricted set of file tools, including a fresh read-only session after adoption.

Validation:

  • 25 native-runner/guard tests and 34 smoke assertions pass locally; repository checks pass.
  • Four native sessions pass the guarded loading/boundary integration gates. Three semantic reviews pass; the CSV response retains a known string-value overclaim.
  • The earlier eight-session paired comparison reproduces CSV overclaims with and without the skills. Both conditions correct stale documentation.
  • Report regeneration, original stream/event hashes and unchanged package hashes are verified. Earlier attempts retain their original scores.

Recommend experimental merge with the limits documented. The guard is evaluation-only, excludes shell execution and is not an OS sandbox; installing the skills does not install confinement. No general reliability or accuracy advantage is claimed.

Evidence: guarded follow-up, paired comparison.

jaredchu and others added 30 commits September 28, 2026 10:18
Packaging, relative links, declared versions and agreement between the README
result tables and the published artifacts were checked by hand. Run them as one
script and in GitHub Actions, together with the study self-tests and report
regeneration that need no container, model or credentials.

Each installable skill now declares its own version in a VERSION file, so a
copied package states what it is, and CHANGELOG.md reconstructs the per-skill
history from the commits and dated reports each version shipped with. The
repository tags name core releases only; the adoption skill moves separately.

These are mechanical repository checks. Every entry records what was actually
executed for a version, keeping static validation separate from agent
evaluation as CONTRIBUTING.md requires.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The maintenance rule and marker went to AGENTS.md by name. A client may not read
that file: Claude Code loads a project's AGENTS.md only when no CLAUDE.md exists
in the working directory or above it, so a rule written beside a CLAUDE.md is
never loaded. Adoption then reports success while the ongoing mechanism it
establishes does nothing.

Identify the instruction file in effect instead of assuming a name, and when a
project keeps several, keep the rule in one of them and reachable from the others
by their own reference or import rather than duplicating guidance. State the rule
without a client-specific invocation prefix, so the wording written into a project
stays valid wherever the skill is installed.

Add two instruction-file routing cases with ten static grader controls, covering
a CLAUDE.md-only project and a project holding both files. Only those controls
ran: no model session has been executed for adoption v0.1.3, and the README,
changelog and case notes say so.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…text

Five completed studies report no skill-specific advantage. Three properties of
the current designs would produce that tie from any skill: the shared task text
already states much of the skill's guidance to both arms, the fixture evidence
labels its own approval and observation status, and small explicit inputs let
every reader answer from raw sources. More trials under these conditions would
cost more and tie again.

Record the proposed changes rather than editing any frozen fixture or published
result: a neutral baseline request, evidence that does not classify itself, a real
retrieval cost, and discovery treated as its own factor, which is the one place
existing evidence already suggests an effect. The protocol is proposed and not
owner-approved; it adds no result.

Update project context for the static checks, Claude Code packaging, adoption
v0.1.3 and the delegated contribution authority, keeping observed, approved and
proposed statements separate.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Every published study ran on Codex, and the v0.1.3 follow-up records native
Claude Code behavior as unestablished. This harness checks the four questions
that review raised on the other client: skill discovery and invocation, rule
placement in Claude-only and mixed instruction projects, repeat passes that
preserve files and the original adoption date, and audit-only requests leaving
the project byte-identical.

Fixtures come from the shared adoption cases, so both clients exercise the same
projects, protected content and immutable files, with each immutable file stated
explicitly in the request as the initial v0.1.3 review required. Three things
differ by necessity and are recorded in the protocol: slash-command or unnamed
invocation, installation into the project's own .claude/skills, and an audit that
reports in its final message because this client has no /output directory.
Installed packages are committed with the fixture and hash-checked after every
session, so a reference to an installed file validates while tampering fails.

Invocation is measured from the session's own streamed tool calls, following the
routing study's exposure method. Mechanical checks and invocation are recorded
separately from semantics, which must be supplied as explicit reviews.

Executed here: 34 stream, fixture and grader assertions, now in CI, plus a
discarded dry run that confirmed the failure path records execution errors rather
than silently passing. No model session ran: the available Claude Code CLI could
not authenticate, so the six sessions remain outstanding and this establishes no
behavior on that client.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The smoke-test harness raises choices that belong to the evaluation design rather
than to one session: where the skills are installed, where an audit reports, whether
unnamed skill selection gates acceptance, whether this isolation is sufficient
without a container, which model to pin, and how the client's AGENTS.md version gate
affects the split-file wording. Each is recorded with a recommendation so the
collaborator settling them has the reasons, not only the conclusion.

The report also states what was executed and what was not: every static check and
its count, and six model sessions that did not run because the available Claude Code
CLI could not authenticate. It accepts both findings from the v0.1.3 review as
fixture and verifier defects, and keeps the open absolute-path portability concern
visible rather than treating the project-scoped install as a fix in the skill.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@jaredchu
jaredchu marked this pull request as ready for review September 28, 2026 07:30
@jaredchu
jaredchu merged commit 998ff77 into main Sep 28, 2026
3 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant