Skip to content

Repository files navigation

Version License Claude Code Plugin 3 Execution Modes Patterns

Harness in 15 seconds: one request becomes an agent team
15-second motion reel · Watch the full 2-minute interactive walkthrough (Korean) →

Harness v2 — The Team-Architecture Factory for Claude Code

English | 한국어

Harness is a team-architecture factory for Claude Code. One sentence — "build a harness for this project" · "하네스 구성해줘" — and the plugin turns your domain description into an agent team and the skills they use.

What's new in v2

v2 is a ground-up rebuild for the current Claude Code multi-agent runtime:

  • Three native execution modes. v1 knew two modes built on the experimental TeamCreate API, which no longer exists. v2 targets what actually ships today:
    1. Workflow orchestration — deterministic scripts (pipeline() / parallel() / schemas / budgets) for fan-outs, verification loops, and large-scale runs
    2. Persistent agent collaboration — named agents + SendMessage + shared task lists, with context retained across turns
    3. Sub-agent delegation — lightweight one-shot parallel dispatch
  • Workflow-native quality patterns. Adversarial verification, judge panels, loop-until-dry, multi-modal sweeps, completeness critics — codified so generated harnesses filter out plausible-but-wrong output.
  • No experimental flags. The CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1 dependency is gone entirely.
  • Sane model policy. v1 pinned every agent to model: "opus". v2 selects a tier per agent — fable / opus / sonnet — based on the task's complexity, duration, autonomy, and latency needs, and forbids unjustified blanket pins.
  • /harness:evolve actually ships. The evolution mechanism v1 only documented is now a real skill: it captures the delta between your initial and current harness, generalizes feedback, and feeds it back into agents/skills/orchestrators.
  • v1 migration built in. The factory detects v1 artifacts (TeamCreate, TeamDelete, experimental flags) and offers a mechanical migration path.

Core features

  • Agent team design — six architecture patterns (Pipeline, Fan-out/Fan-in, Expert Pool, Producer-Reviewer, Supervisor, Hierarchical Delegation), each mapped to its best v2 execution mode
  • Skill generation — context-efficient skills via Progressive Disclosure
  • Orchestration — data-passing protocols (structured schemas, files, messages, tasks), error handling, resume support
  • Verification — trigger evals, dry runs, with-skill vs. without-skill A/B testing (optionally as a workflow itself)
  • Evolution — /harness:evolve turns usage feedback into measurable next-generation improvements

Workflow

Phase 0: Audit existing harness (new / extend / maintain — v1 artifacts detected here)
Phase 1: Domain analysis (incl. control-flow shape of the work)
Phase 2: Execution mode & team architecture design
Phase 3: Agent definitions (.claude/agents/)
Phase 4: Skill generation (.claude/skills/)
Phase 5: Orchestration & CLAUDE.md pointer
Phase 6: Verification & testing
Phase 7: Maintenance — evolution via /harness:evolve

Install

Via marketplace

/plugin marketplace add revfactory/harness
/plugin install harness@harness-marketplace

As global skills

cp -r skills/harness ~/.claude/skills/harness
cp -r skills/evolve ~/.claude/skills/harness-evolve

No environment variables or experimental flags required.

Usage

하네스 구성해줘
build a harness for this project
design an agent team for <domain>

After using a generated harness:

하네스 회고해줘 / evolve the harness with this feedback

Choosing an execution mode

Mode Primitive When
Workflow orchestration Workflow scripts Control flow is deterministic: enumerable fan-outs, verification loops, large scale, structured outputs
Persistent agents Agent(name:) + SendMessage + tasks Long-lived specialists that keep context; iterative feedback and negotiation
Sub-agent delegation one-shot Agent calls Fire-and-forget parallel work; results only

The factory picks the mode from the shape of the control flow, not from team size — and mixes modes per phase when that fits better.

Generated artifacts

your-project/
├── .claude/
│   ├── agents/          # agent definitions (who)
│   │   ├── analyst.md
│   │   ├── builder.md
│   │   └── qa.md
│   └── skills/          # skills (how) + one orchestrator (who-when-in-what-order)
│       ├── analyze/SKILL.md
│       └── build/SKILL.md
└── CLAUDE.md            # minimal pointer: trigger rule + change history

Migrating from v1

See docs/migration-v1-to-v2.md. Summary: remove TeamCreate/TeamDelete/broadcast/flag references, convert fan-outs to Workflow scripts, rewrite remaining collaboration with named agents + SendMessage, drop blanket model: "opus" pins. The factory automates this when it detects v1 artifacts (Phase 0).

Prior results (v1)

A controlled A/B on 15 software-engineering tasks measured the effect of structured pre-configuration on LLM code-agent output quality: mean quality 49.5 → 79.3 (+60%), 15/15 win rate, −32% output variance (n=15, author-run, see revfactory/claude-code-harness). Treat these as author-measured numbers; run your own pilot for adoption decisions.

License

Apache 2.0

About

A meta-skill that designs domain-specific agent teams, defines specialized agents, and generates the skills they use.

Topics

Resources

Contributing

Stars

9.1k stars

Watchers

59 watching

Forks

Releases

Packages

Contributors