Preserve project knowledge and maintain accurate, consistent documentation across AI sessions.
Context Docs is an open-source convention and reusable agent skill for maintaining Markdown project knowledge. It adapts to existing documentation, preserves decisions and evidence, and keeps current context from becoming a session diary.
Status: experimental, core skill v0.1.2; adoption skill v0.1.4. The package contains instructions only. It runs when an agent uses it; there is no background service, automatic scheduler, cloud account or runtime dependency. Git remains available for history and review.
With both skills installed, adopt a new or existing project with:
$adopt-context-docs
This applies the core method and adds or merges a maintenance rule in the
instruction file the project's agent actually loads, such as AGENTS.md or
CLAUDE.md. After checking the setup, it records a small marker beside that rule:
Method: context-docs
Adopted: YYYY-MM-DD
Entry point: path/to/context.md
New maintenance sections preferably use a Context maintenance heading, then the
marker, then maintenance instructions. Existing equivalent layouts are preserved.
The path reflects the project's actual entry point. Repeat runs preserve the
original adoption date and reuse equivalent guidance. When retrofitting a marker,
an undocumented original adoption date stays unknown. Audit-only runs do not
write markers. The marker records workflow adoption; it is not an accuracy or
freshness certificate, and its absence alone does not prove non-adoption.
For routine audits and maintenance, use $context-docs directly:
Use $context-docs to audit this project's context documents. Report the gaps
without editing files.
Use $context-docs to update project context from the work we just completed.
Preserve the existing layout, approved decisions, and unresolved blockers.
Use $context-docs to establish a minimal context entry point for this project.
Use existing docs and code as evidence; mark anything you cannot verify.
Adoption v0.1.0 passed two adoption-and-repeat regression cases (four model sessions). Both repeat passes left project files unchanged. These are small author-reviewed checks, not evidence of general reliability. The v0.1.1 marker regression retains an initial guidance-rewrite failure and date-test ambiguity. Targeted follow-ups passed marker-only adoption, read-only auditing and explicit-date preservation. Ten sessions across initial and revised conditions are reported separately; the final revision was tested on the two affected unmarked cases. Adoption v0.1.2 adds only the preferred layout for new sections. Metadata, package links and existing static controls were checked; no new model evaluation was run. Adoption v0.1.3 routes the rule and marker to the instruction file a client loads and drops client-specific invocation wording. Its instruction-file cases pass ten static grader controls. The v0.1.3 merge review ran nine model sessions: eight met frozen acceptance, with a split-file failure involving an immutable heading and an installed-skill link. Routing rubrics passed, but the original result is not a full regression pass. A separately frozen follow-up passed four routing/repeat sessions after making file constraints explicit and validating installed references consistently. The skills are unchanged; the initial failure remains recorded. Per-version history is in the changelog.
The workflow reads existing context, checks relevant evidence, updates canonical sections, consolidates duplication, and reviews the resulting diff and links. An audit stays read-only. Maintenance produces ordinary, reviewable file edits.
Make the context entry point discoverable: link it from the project README or documentation index. For a fresh session, ask the agent to start there and read the linked context before working. A context file's presence alone does not ensure that an agent will read it. This routing is already part of the skill's initialization workflow.
Ask the built-in installer:
Use $skill-installer to install both skills from jaredchu/context-docs:
- skills/context-docs
- skills/adopt-context-docs
Install them together in ~/.agents/skills/ for use across my projects.
If context-docs is already installed, preserve it and install only the missing
adoption skill; report any version mismatch before upgrading.
The adoption skill requires context-docs as a sibling folder. It is a small
setup wrapper, not a standalone replacement for the core skill. The core skill
can be installed and used by itself.
Alternatively, clone this repository and copy both complete folders from skills/
into ~/.agents/skills/. For a repository-scoped installation, copy both into
.agents/skills/ at the project root instead. Use one scope to avoid duplicate
skill entries. Check existing folders and review changes before upgrading; retain
references, assets, metadata and licenses. Codex normally detects changes
automatically; restart if they do not appear.
See the official skill documentation.
Copy both folders into a skills directory Claude Code reads: ~/.claude/skills/
for every project, or .claude/skills/ inside one repository. Use one scope.
mkdir -p ~/.claude/skills
cp -r skills/context-docs skills/adopt-context-docs ~/.claude/skills/Keep both folders side by side: the adoption skill reads its sibling core skill.
Invoke a skill as /context-docs or /adopt-context-docs, or describe the task
and let Claude select it. The packaged agents/openai.yaml is Codex metadata and
is ignored here.
Direct AGENTS.md loading requires Claude Code v2.1.277 or later with its built-in
agents-md plugin enabled. By default, it loads AGENTS.md only when no
CLAUDE.md, .claude/CLAUDE.md or CLAUDE.local.md exists in the working directory
or its ancestors; user-level ~/.claude/CLAUDE.md and managed instructions do not
disable that fallback. The Project instructions setting can instead load both
file types, only Claude files, or only managed instructions at launch. Older or
otherwise unsupported sessions can import AGENTS.md from CLAUDE.md.
Adoption v0.1.3 targets the instructions actually loaded in the session. Confirm
the result: the rule should be in a loaded file, or included through that file's
supported import. Check /context and the Project instructions setting rather
than inferring loading solely from filenames. See the
Claude Code skill and
memory documentation.
Native support remains experimental. The guarded follow-up passes four loading/boundary checks, including a denied outside-project write and a fresh read-only session loading the newly added maintenance rule. This clears those integration gates and supports an experimental merge recommendation. The guard is evaluation tooling with a restricted tool set; installing the skills does not install a sandbox or enforce write confinement.
Factual limitations remain: the CSV response still overclaims value types. The eight-session paired study reproduced that error with and without the skills, while both conditions corrected stale documentation. These author-reviewed runs on Claude Code 2.1.234 with Claude Opus 5 establish neither a skill-specific accuracy advantage nor general reliability. Earlier failures and scores remain recorded; file checks, execution and semantic review are reported separately.
Other agents can use the same instructions when they support SKILL.md folders,
or read the standard directly.
Behavior on clients other than Codex and Claude Code has not been evaluated.
| Concern | Convention |
|---|---|
| Structure | Shared document roles, mapped to existing files |
| Truth | Distinguish observations, approved decisions, proposals and unknowns |
| Maintenance | Refresh current state after meaningful work; preserve evidence |
| Growth | Consolidate repetition, link to details, retain useful history |
| Portability | Keep canonical content in readable Markdown with source links |
A project with context.md and one with docs/project-context.md can follow the
same method. No forced directory migration or universal document-size limit.
- Core maintenance skill
- Project adoption shortcut
- Context standard
- Optional context template
- Optional decision template
- Before-and-after example
- Behavioral evaluation scenarios
- Static repository checks
- Per-version changelog
- Quality and fresh-reader study
- Larger public-source handoff pilot
- README routing and verified reading
- Decision histories across four maintenance passes
- Public Harbor suite and results
- Earlier local-project pilot
- This project's own context
The primary goal is durable project knowledge: retain useful content, keep claims accurate, and make decisions and procedures consistent across documents and sessions. Shorter files or faster editing do not establish that goal.
The quality-first study compared two synthetic projects with scattered, conflicting or absent context. It ran 12 three-stage maintenance trials, followed by 36 fresh-reader sessions: 72 model sessions and 216 reader answers, plus reference/no-op controls.
| Final-artifact measure | Unmaintained | Ordinary maintenance | Context Docs v0.1.1 |
|---|---|---|---|
| Required knowledge recorded | 12/54 | 54/54 | 54/54 |
| Correct, supported reader answers | 72/72 | 72/72 | 72/72 |
| Correct in both reader sessions | 36/36 | 36/36 | 36/36 |
| Target-claim accuracy | 37.5% | 100.0% | 100.0% |
| Missing required items | 22 | 0 | 0 |
| Incorrect required items | 20 | 0 | 0 |
| Critical document findings | 11 | 0 | 0 |
| Conflicting required items | 0 | 0 | 0 |
Both maintenance methods improved the stored record. Context Docs did not outperform ordinary maintenance on these quality measures. Readers could inspect the same small raw evidence set in every arm and all answered correctly, including with unmaintained docs. This reader ceiling establishes no accuracy advantage or general equivalence. Both maintenance arms retained all required items through an evidence update and made no edits in the final no-change pass.
Coverage counts required knowledge captured accurately in maintained Markdown; raw-file survival alone does not count. Target-claim accuracy excludes missing items, which are shown separately. Incorrect current claims and unresolved operative conflicts are separate categories. The two projects contain only 514–656 total input words, are author-created and author-reviewed, and have one maintenance trial per scenario/arm. This is a reproducible development study, not an independent benchmark. See the method, per-project results and limits, execution commands, and snapshots and scored evidence.
A larger public-source pilot used pinned Flask and Click releases containing 143,604 and 102,647 source/test/doc words. It completed four handoffs and 12 fresh readers: 16 sessions, 72 answers. Existing upstream documentation stayed intact; this complements the weak-docs comparison above.
| Reader outcome | Upstream docs/code | Ordinary handoff | Context Docs v0.1.1 |
|---|---|---|---|
| Flask | 12/12 | 11/12 | 12/12 |
| Click | 12/12 | 12/12 | 12/12 |
| Total | 24/24 | 23/24 | 24/24 |
| Correct in both phrasings | 12/12 | 11/12 | 12/12 |
Both handoff methods covered 24/24 required items. The one ordinary-reader miss omitted a required cleanup guard. No skill-specific accuracy advantage is established: recorded traces show that none of the eight readers given handoffs opened their content; they answered from upstream sources. The ordinary handoff already contained the omitted guard. The protocol also prohibited adding a link from the upstream README, limiting discovery. This tests handoff availability, not the effect of confirmed handoff consumption.
Sources were externally authored, but questions and reviews were not independent. There was one maintenance attempt per method/project, and both projects belong to the same ecosystem. See the discovery finding, exact miss and limits and all outputs and judgments.
A routing follow-up then reused those handoffs unchanged, added equivalent orientation links to temporary READMEs, and compared README-first discovery with explicitly directed reading. It ran 12 fresh sessions and 72 answers, using one previously tested phrasing.
| Reader mode and outcome | Upstream orientation | Ordinary handoff | Context Docs handoff |
|---|---|---|---|
| README-first discovery: supported answers | 12/12 | 12/12 | 12/12 |
| README-first discovery: full orientation text exposed | 1/2 | 2/2 | 2/2 |
| Directed reading: supported answers | 12/12 | 12/12 | 12/12 |
| Directed reading: full orientation text exposed | 2/2 | 2/2 | 2/2 |
All eight handoff readers received the complete document before answer-writing, verified from model-visible tool output. All 12 readers saw the complete README. Accuracy remains tied: the original sources also support every answer. This checks reading under the new routing/instructions; it does not establish a skill advantage, unguided discovery or equivalence. The earlier omitted cleanup guard is present in the repeated answers. See the unchanged criteria, exposure method and limitations and reproducible evidence.
A decision-history follow-up replayed two externally authored Python policy histories through four maintenance passes: 28 model sessions and 96 reader answers. It tests evolving requirements, authority, rationale and authentic source contradictions; both histories come from one project and the questions/review remain author-created.
| Outcome | Unmaintained | Ordinary maintenance | Context Docs v0.1.1 |
|---|---|---|---|
| Required knowledge in final context | 0/16 | 16/16 | 16/16 |
| Required knowledge across all stages | 0/63 | 63/63 | 63/63 |
| Correct, supported reader answers | 32/32 | 32/32 | 32/32 |
| Correct in both reader sessions | 16/16 | 16/16 | 16/16 |
| Readers exposed to all Markdown | 4/4 | 4/4 | 4/4 |
| Critical document findings across stages | 0 | 0 | 0 |
Both methods retained every assessed item and made no Markdown changes in the final no-new-evidence pass. Reader accuracy tied, including readers of raw history alone; those original PEPs remain accessible in every arm. Zero in the unmaintained column measures absent context summaries, not absent source facts. One skill session briefly wrote a backup outside the allowed directories, then corrected it; the project-file grader did not catch that. See the scope finding and limits and all artifacts and explicit judgments.
Existing studies provide maintenance regression coverage:
- 48 trials, v0.1.0 versus ordinary instructions: both passed 24/24 under mechanical checks and author-reviewed semantic criteria. No correctness advantage was observed on these small fixtures. All audits preserved project bytes and final no-change passes preserved Markdown.
- 12 trials, v0.1.0 versus concise v0.1.1: both passed 6/6. The candidate met its frozen preservation/concision rule and was adopted. This was a prompt refinement test, not evidence of better knowledge retention or reader answers.
- The earlier local-project pilot includes a workflow contradiction the skill missed. Passing synthetic cases does not erase that miss.
| Task | Ordinary instructions | With Context Docs |
|---|---|---|
| stale-config | 3/3 | 3/3 |
| workflow-conflict | 3/3 | 3/3 |
| decision-status | 3/3 | 3/3 |
| uncommitted-work | 3/3 | 3/3 |
| links-and-fences | 3/3 | 3/3 |
| audit-only | 3/3 | 3/3 |
| initialize | 3/3 | 3/3 |
| repeated-maintenance | 3/3 | 3/3 |
| Total trials | 24/24 | 24/24 |
The 48-trial table is generated from recorded trials and explicit reviews. Its fixtures contain only 19–87 Markdown words; the follow-up fixtures contain 572–676. Both studies use author-created cases and author review, not independent held-out projects. Global instructions were excluded and recorded inputs checked.
Full results retain all observations, including document growth, slower runs and resource overhead. These are secondary diagnostics, not the product's success criteria: 48-trial report, v0.1.1 comparison, reproduction commands, and machine-readable results.
The skill guides an agent; it cannot guarantee factual correctness, conflict-free edits or decision preservation. Review consequential changes. It neither captures every conversation nor grants permission to publish documents or alter systems.
Next evaluations should prioritize independently contributed histories from other projects and condition-blind review, with verified context reading. Current results do not establish a skill-specific correctness advantage. Cloud retrieval can be an optional integration later; no backend is required or included.
Contributions are welcome: see CONTRIBUTING.md. Licensed under MIT.