Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
33 commits
Select commit Hold shift + click to select a range
8f30089
Add static repository checks, CI and per-skill version records
jaredchu Sep 28, 2026
894dcec
Route the adoption rule to the instruction file the client loads
jaredchu Sep 28, 2026
0e88914
Propose a protocol that can distinguish the methods, and maintain con…
jaredchu Sep 28, 2026
6597389
Fix package and version validation and clarify Claude loading
jaredchu Sep 28, 2026
17e8dd7
Evaluate adoption v0.1.3 and retain split-file validation failure
jaredchu Sep 28, 2026
84a2cab
Fix routing evaluation constraints and validate installed references
jaredchu Sep 28, 2026
6f1a9ac
Add a native Claude Code smoke test harness, not yet executed
jaredchu Sep 28, 2026
2fc7b69
Record the Claude Code handoff, open decisions and what was not run
jaredchu Sep 28, 2026
bedc176
Fix frozen Claude smoke protocols and whole-tree validation
jaredchu Sep 28, 2026
336a414
Freeze repaired native Claude Code smoke protocol
jaredchu Sep 28, 2026
1f9fc84
Retain authentication-blocked native Claude Code attempt
jaredchu Sep 28, 2026
9540074
Freeze native Claude Code retry after refreshed authentication
jaredchu Sep 28, 2026
27ea769
Isolate Claude evaluation stdin and retain contaminated retry
jaredchu Sep 28, 2026
e676300
Freeze native Claude evaluation with isolated stdin
jaredchu Sep 28, 2026
b473c65
Handle Claude permission events and allow read-only hash verification
jaredchu Sep 28, 2026
b8e40b1
Freeze targeted Claude audit and discovery follow-up
jaredchu Sep 28, 2026
65b39b5
Publish native Claude results and hold merge for factual failures
jaredchu Sep 28, 2026
aad4426
Require scoped evidence and verified instruction loading in skills
jaredchu Sep 28, 2026
866e382
Freeze strengthened native regression for core 0.1.2 and adoption 0.1.4
jaredchu Sep 28, 2026
dce1a47
Validate directory links and refine grounded native documentation
jaredchu Sep 28, 2026
1024fa5
Freeze final native regression with directory and claim-scope repairs
jaredchu Sep 28, 2026
6772889
Preserve decision attribution and qualifications in final summaries
jaredchu Sep 28, 2026
a8481f1
Freeze targeted native discovery decision-attribution regression
jaredchu Sep 28, 2026
269e260
Review context claims against the resulting project state
jaredchu Sep 28, 2026
cdb4cee
Freeze final-state consistency discovery regression
jaredchu Sep 28, 2026
03092c5
Record strengthened native evaluations and remaining discovery failures
jaredchu Sep 28, 2026
a512762
Add frozen paired native comparison with ordinary maintenance baseline
jaredchu Sep 28, 2026
c962009
Freeze eight-session native paired evaluation before execution
jaredchu Sep 28, 2026
153b4ef
Record paired native outcomes and limits on loading verification
jaredchu Sep 28, 2026
e8a041d
Observe native instruction loading and guard evaluation file tools
jaredchu Sep 28, 2026
d07f40d
Freeze guarded loading and boundary follow-up before native runs
jaredchu Sep 28, 2026
0b57a85
Record passing guarded integration checks and experimental merge reco…
jaredchu Sep 28, 2026
ce7b0ba
Link the experimental Claude Code merge review
jaredchu Sep 28, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
72 changes: 72 additions & 0 deletions .github/workflows/checks.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,72 @@
name: Static checks

# Repository checks only: packaging, links, published-table agreement and the
# study self-tests that need no container, model or credentials. These establish
# nothing about agent behavior; model evaluations stay manual and are reported
# separately, as CONTRIBUTING.md requires.
on:
push:
branches: ['**']
pull_request:
workflow_dispatch:

permissions:
contents: read

jobs:
checks:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-python@v5
with:
python-version: '3.12'

- name: Packaging, links and published tables
run: python3 evals/checks/static_checks.py

- name: Static checker regression tests
run: python3 -m unittest discover -s evals/checks -p 'test_*.py'

- name: Core suite self-test
run: python3 selftest.py
working-directory: evals/suite

- name: Verifier regression tests
run: python3 -m unittest discover -s evals/suite -p 'test_*.py'

- name: Quality study self-test
run: python3 study.py selftest
working-directory: evals/quality

- name: Claude Code smoke-test selftest
run: |
python3 evals/claude-code/smoke.py selftest
python3 -m unittest discover -s evals/claude-code -p 'test_*.py'

- name: Adoption grader controls
run: |
python3 evals/adoption/build.py "$RUNNER_TEMP/adoption-suite"
python3 evals/adoption/marker.py "$RUNNER_TEMP/adoption-marker-suite"
python3 evals/adoption/instructions.py "$RUNNER_TEMP/adoption-instructions-suite"

- name: Published reports regenerate from published evidence
run: |
set -eu
copy() { rm -rf "$RUNNER_TEMP/$2"; cp -r "evals/results/$1" "$RUNNER_TEMP/$2"; }
copy 2026-09-27-quality quality
(cd evals/quality && python3 report.py "$RUNNER_TEMP/quality" > /dev/null)
diff "$RUNNER_TEMP/quality/table.md" evals/results/2026-09-27-quality/table.md
copy 2026-09-27-history history
(cd evals/history && python3 report.py "$RUNNER_TEMP/history" > /dev/null)
diff "$RUNNER_TEMP/history/table.md" evals/results/2026-09-27-history/table.md
(cd evals/routing && python3 report.py \
"$GITHUB_WORKSPACE/evals/results/2026-09-27-routing/trials.json" \
"$GITHUB_WORKSPACE/evals/results/2026-09-27-routing/reviews.json" \
"$GITHUB_WORKSPACE/evals/results/2026-09-27-routing/exposure.json" \
"$RUNNER_TEMP/routing" > /dev/null)
diff "$RUNNER_TEMP/routing/table.md" evals/results/2026-09-27-routing/table.md
copy 2026-09-28-native-guarded guarded
python3 evals/claude-code/guarded_report.py "$RUNNER_TEMP/guarded" > /dev/null
diff "$RUNNER_TEMP/guarded/table.md" evals/results/2026-09-28-native-guarded/table.md
diff "$RUNNER_TEMP/guarded/summary.json" evals/results/2026-09-28-native-guarded/summary.json
5 changes: 4 additions & 1 deletion AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -12,7 +12,10 @@
- Keep guidance applicable to varied project layouts. Do not add a universal
rule for an isolated example without a demonstrated need.
- Validate skill metadata, packaged reference links and affected evaluation
scenarios. Report static checks separately from actual agent evaluations.
scenarios with `python3 evals/checks/static_checks.py`; CI runs it on every push.
Report static checks separately from actual agent evaluations.
- Record each skill change in `CHANGELOG.md` with the version it lands in and what
was actually tested. Keep each skill's `VERSION` file in step with it.
- Use synthetic examples. Do not copy private project context into this repo.
- Update current project context when behavior, scope or status changes.
- Publication requires authorization from the task; these instructions grant none.
Expand Down
170 changes: 170 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,170 @@
# Changelog

Each installable skill under `skills/` carries its own version in a `VERSION`
file and moves independently. The repository tags `v0.1.0` and `v0.1.1` name core
skill releases only; later entries identify the package they change. Versions are
reconstructed here from the commits and dated reports they were published with,
so the history is auditable without reading every evaluation document.

Every entry states what was actually tested. A static check or an authored example
is not a behavioral evaluation, and adopting a version is an implementation choice,
not a demonstrated accuracy gain.

## Repository checks and installation guidance — unreleased

- Enforce package boundaries for skill links and nested references, allowing the
adoption package's declared dependency on its sibling core. Reject repository-only,
absolute and escaping symlink targets. Exclude gitignored `.local/` study output
from repository link checks.
- Compare `VERSION` files against the current README status and latest changelog
entry for each skill, so historical mentions cannot hide a stale declaration.
- Correct Claude Code installation guidance for version requirements, local
instruction files and configurable loading behavior; keep that detail canonical
in the README.
- Twelve checker regression tests passed. These are static checks, not model
evaluations. These repository-only changes do not alter installable skills;
separate skill changes are recorded below.

## Native loading observer and file-tool guard — 2026-09-28

- Add evaluation-only InstructionsLoaded observation and a PreToolUse guard for
six file/skill tools. Record loading hashes and correlate tool calls with guard
decisions; exclude shell/delegation/network tools. No packaged skill changed.
- Seven new guard/observer controls join the eighteen existing native tests;
twenty-five credential-free tests pass. CI also regenerates the new result summary.
- [Four actual native sessions](evals/results/2026-09-28-native-guarded/README.md)
pass all loading/boundary gates, including a blocked outside write and fresh
read-only loading of the newly installed rule. Three semantic reviews pass;
the CSV response retains the known universal string-value error.
- Recommend experimental merge with factual limits. This clears the restricted
integration gates only, not unrestricted-client reliability. The guard is not
installed with the skills and is not an OS sandbox. Earlier scores are unchanged.

## Native paired evaluation — 2026-09-28

- Add a bounded ordinary-versus-skills comparison with matched requests, fresh
cases, two attempts per condition and frozen inputs/criteria. Five new runner
tests and the existing thirteen native smoke tests pass; these are static checks.
- [Eight actual native sessions](evals/results/2026-09-28-native-paired/README.md)
use the unchanged core 0.1.2 / adoption 0.1.4. CSV overclaims appear in both
conditions; stale-state reconciliation passes in both. No skill-specific cause
of the earlier errors is established.
- Retain unverified loading confirmations, scratch-directory scope violations and
optional command denials separately. The hold remains; conservative acceptance
totals are not an accuracy ranking. No new skill revision or extra retry follows.

## context-docs 0.1.2 — unreleased

- Tighten the existing evidence check: verify the exact claim and retain the
source's scope in both documents and final reports. Distinguish installed
package versions from reference versions, and local configuration from runtime
behavior or publication history. Omit unverified incidental details.
- Clarify that creating a minimal README link does not require inventing project
intent when a project has no README or index.
- The second unreleased candidate also avoids unverified guarantees about every
input and keeps final reports focused instead of restating unchanged context.
- A targeted refinement attributes agent-chosen methods/layouts to their actual
decision maker and preserves document qualifications in final summaries.
- Final review reconciles statements with the resulting project, including facts
made stale by the agent's own edits.
- Addresses the retained [native discovery failures](evals/results/2026-09-28-claude-code-followup/README.md).
Four separately frozen native runs retain all candidates: the
[first six-session run](evals/results/2026-09-28-native-v014/README.md) passes
one semantic session, the [second](evals/results/2026-09-28-native-v014-followup/README.md)
passes five, and two targeted discovery runs still fail. The
[latest result](evals/results/2026-09-28-native-state-followup/README.md) retains
an unsupported universal CSV claim and a stale final README statement.
Hold the merge; this is author-reviewed development evidence, not a reliability
claim. Exact commits distinguish candidates with the same unreleased versions.

## adopt-context-docs 0.1.4 — unreleased

- Require evidence for automatic instruction loading rather than inferring it
from file presence or an explicit file read. Keep existing workflow adoption
distinct from verified loading, including during audits.
- Align the Codex default prompt with the loaded instruction file instead of
hardcoding AGENTS.md. This metadata change is statically validated.
- Addresses the unsupported loading claim in the
[native audit](evals/results/2026-09-28-claude-code-followup/README.md).
The [second native candidate](evals/results/2026-09-28-native-v014-followup/README.md)
passes routing, repeat and audit semantic criteria, including qualified loading.
Discovery remains factually unreliable in the
[latest targeted run](evals/results/2026-09-28-native-state-followup/README.md).
Permission denials remain separate execution failures; the branch stays on hold.

## adopt-context-docs 0.1.3 — unreleased

- Route the ongoing maintenance rule and adoption marker to the instruction file
the client actually loads, instead of assuming `AGENTS.md`. Both `AGENTS.md` and
`CLAUDE.md` are common, and a client may load only one of them; a rule written to
an unloaded file silently does nothing. When a project keeps several, the rule
belongs in the one in effect and stays reachable from the others by reference or
import rather than being duplicated.
- State the maintenance rule without a client-specific invocation prefix, so the
wording written into a project is valid wherever the skill is installed.
- Add [instruction-file routing cases](evals/adoption/instructions.py) with ten
static grader controls. Before model evaluation, remove their inherited AGENTS.md
destination and restate the client condition in the fresh repeat session.
- [September 28 merge review](evals/results/2026-09-28-adoption-v013/README.md):
nine model sessions, eight accepted; all three repeats and the audit preserved
bytes. The split-file case passed routing rubrics but failed frozen immutable-file
and external-link checks. Fixture limitations are retained with that failure;
the review recommends holding merge for a targeted follow-up. No native Claude
Code evaluation was run, and the skill contents were not tuned after this result.
- [Separately frozen routing follow-up](evals/results/2026-09-28-adoption-v013-followup/README.md):
four sessions passed after exposing immutable-file constraints in requests and
checking declared installed references by hash in links, code spans and prose.
Both next-day repeats preserved bytes and dates. Nine new verifier regression
tests and existing static checks passed. The reviewer now considers the branch
ready to merge within its experimental scope. The original failure is retained;
neither skill changed, and native Claude Code behavior remains untested.

## adopt-context-docs 0.1.2 — 2026-09-27 (`eed23ed`)

- Prefer a `Context maintenance` heading, then the marker, then instructions when
creating a new maintenance section; preserve existing equivalent layouts.
- Metadata, package links and existing static controls were checked. No new model
evaluation was run for this version.

## adopt-context-docs 0.1.1 — 2026-09-27 (`9a5211a`)

- When equivalent context and maintenance guidance already exist, add only the
missing marker and preserve the guidance's wording.
- [Marker regression](evals/results/2026-09-27-adoption-marker/README.md): ten
sessions across initial and revised conditions. The initial run's guidance
rewrite failure and date-test ambiguity are retained, not erased. The final
revision was tested on the two affected unmarked cases only.

## adopt-context-docs 0.1.0 — 2026-09-27 (`d0febcf`)

- First adoption wrapper: applies the core method and merges a maintenance rule
into the project's instructions. Requires `context-docs` as a sibling package.
- [Adoption and repeat regression](evals/results/2026-09-27-adoption/README.md):
two cases, four sessions. Both repeat passes left project files unchanged.

## context-docs 0.1.1 — 2026-09-27 (`6c96d2a`, tag `v0.1.1`)

- Shorter entry point, 645 to 417 words, with the frozen preservation and concision
rule met.
- [Concision comparison](docs/evaluation-2026-09-27-concise.md): 12 trials, both
the original and the candidate passed 6/6. This was a prompt refinement test; it
established no better knowledge retention or reader accuracy. The candidate was
slower and emitted more output tokens.
- Later studies exercised this version without changing it: the
[quality and fresh-reader study](docs/evaluation-2026-09-27-quality.md), the
[public-source handoff pilot](docs/evaluation-2026-09-27-public.md), the
[routing follow-up](docs/evaluation-2026-09-27-routing.md) and the
[decision-history study](docs/evaluation-2026-09-27-history.md). None of them
found an accuracy advantage over competent ordinary maintenance.

## context-docs 0.1.0 — 2026-09-26 (`7517367`, tag `v0.1.0`)

- First published skill, [standard](skills/context-docs/references/standard.md)
0.1.0, templates and evaluation scenarios. Instructions only: no runtime
dependency, background service or scheduler.
- [Local-project pilot](docs/evaluation-2026-09-26.md): eight runs. All 30
controlled checks met, and a specific workflow contradiction was missed that the
ordinary cleanup baseline resolved. That miss stands as part of the record.
- [48-trial public suite](docs/evaluation-2026-09-27.md): both arms passed 24/24
under mechanical checks and author-reviewed criteria. No correctness advantage
was observed. The skill used more time and tokens.
11 changes: 8 additions & 3 deletions CONTRIBUTING.md
Original file line number Diff line number Diff line change
Expand Up @@ -22,10 +22,15 @@ For a skill change:
controls before model runs. Record the agent/model, inputs, outcome,
verification limits and any lost information. Report semantic review separately
from mechanical checks, and retain failed trials.
3. Check relative links and that the skill works when copied without the rest of
this repository. Use your client's skill validator if available.
3. Run `python3 evals/checks/static_checks.py` for packaging, links, declared
versions and agreement between README tables and published results. GitHub
Actions runs it with the container-free study self-tests on every push. Also
check that the skill works when copied without the rest of this repository, and
use your client's skill validator if available.
4. Describe what changed, why, and what was actually tested in the pull request.

An authored example or static metadata check is not an independent behavioral
evaluation. Keep those claims separate. Contributions are provided under the
evaluation. Keep those claims separate, including in the
[changelog](CHANGELOG.md): record the version a change lands in and what was
actually executed for it. A version whose only evidence is static checks says so. Contributions are provided under the
repository's MIT license.
Loading
Loading