Skip to content

feat(delegation): every caller is taught that a receipt means running — never re-send on silence (abilityai/trinity-enterprise#568) - #3240

Open
webmixgamer wants to merge 5 commits into
devfrom
feature/568-delegation-contract
Open

webmixgamer wants to merge 5 commits into
devfrom
feature/568-delegation-contract

Conversation

@webmixgamer

@webmixgamer webmixgamer commented Oct 5, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Acceptance criteria

  • AC1 — §Agent Collaboration states the contract. replayed is deliberately not named, although AC1 lists it: no dispatch route emits it until ent#566 ships intent-scoped dedupe (only ask_operator answers replayed), and test_every_status_the_contract_names_is_emitted_by_a_dispatch_route admits only statuses a dispatch route really emits. A replay is still covered by "whatever its status" and "an exact repeat is normally answered with the original" — observed live: an identical repeat answered in 28 ms with the first call's execution_id, and one task ran. Recorded in requirements §37.5.
  • AC2 — the same words in chat_with_agent and every dynamic chat_with_<agent> tool (byte parity pinned by a test that parses the TS array).
  • AC3 — fan_out and send_message carry the rule sentence and point at the contract.
  • AC4 — reviewed against the prompt-authoring conventions in platform_prompt_service.py (the tool description is the single source; always-render only for turn discipline no description can carry; capability-fact phrasing; bare tool names; byte-identical VERBOSE). No new ### section; §Agent Collaboration stays shorter than §Operator Communication (pinned).
  • AC5 — verified on localhost after a backend reload, with no agent rebuild: the agent quoted the contract verbatim from its injected prompt.
  • AC6 — after merge (below).

Changes

  • src/backend/services/platform_prompt_service.py — DELEGATION_CONTRACT, spliced into §Agent Collaboration; the contract's tools added to the Codex orientation.
  • src/mcp-server/src/delegation_contract.ts (new) — the shared text, DELEGATION_RULE, readNotResend.
  • src/mcp-server/src/tools/chat.ts, tools/dynamic-agents.ts, tools/messages.ts, client.ts — descriptions and receipt messages.
  • .github/workflows/mcp-server-test.yml — the ESM boot smoke also imports dist/tools/dynamic-agents.js (reached only from index.ts).
  • Docs: requirements §37.5 (+ §37.1 correction, cross-links from core-agent.md and mcp.md), architecture/mcp-server.md and architecture/execution.md, the feature flows that teach delegation, the agent guides' stale "will fail after 60 seconds" sections, user docs.

Review round (6ac2a69, a9217b0)

/review found no critical issues and ten informational ones, six of them claims the code contradicts. All are fixed:

  • The cap history above. The first commit's message (9105e14) still says the advice "never reached a model" — please don't carry that line into the squash message.
  • agent_busy is not "nothing ran": the SUB-003 usage-limit 429 is raised after the turn ran, and the client labels it agent_busy (bug: chat_with_agent reports a usage-limit 429 as agent_busy although the turn already ran, so a retry repeats its side effects #3244). Contract bullet 4 now routes it through the same list_recent_executions check. The new wording is the same length, so the published figures are unchanged.
  • The receipt leads no longer claim liveness, and the backend's sync-/task queued_timeout replay is rewritten too (bug: an async /task keeps replaying its accepted receipt for 24 h after the run fails, so an exact retry returns the failed run #3245).
  • "Every receipt ends with readNotResend" is corrected: fan_out_timeout has its own wording, and two receipts add a sentence after it.
  • The synchronous /chat route is the one path that emits no agent.task.* event. execution.md, task-completion-events.md and the user docs now say so.
  • §37.1 is finished. The guides' 60 s (now 25 s) and "a long call does not fail" overclaims are fixed, and so are seven stale lines that told MCP callers to poll REST, use async=true without parallel=true, or omit agent_name.
  • Tests: the mutations that survived are now caught (below), and a row timestamp that was set at import time is built per test.

Coordination with #3238

#3238 adds a budget test over every published tool description with a shrink-only PENDING entry for chat_with_agent owned by this issue — and an entry that fits the cap fails. Whichever PR lands second deletes chat_with_agent from PENDING in src/mcp-server/src/tool-description-budget.test.ts.

Test plan

Follow-ups filed while building this: #3232, #3233, #3234 (in review as #3238). From the review: #3244 (a usage-limit 429 is reported as agent_busy after the turn ran), #3245 (an async receipt is replayed for 24 h after its run failed).

Debt adjacency, nothing paid down here: chat.ts carries #1074 (removing the deprecated timeout_seconds will also drop the sync-parallel note moved into that parameter's description), messages.ts carries #3087, client.ts carries #1029.

After merge

  • Set status-in-dev on abilityai/trinity-enterprise#568 by hand — the status workflow does not promote a cross-tracker issue
  • AC6 — ask trinity-pm, which owns the canon, to make canon/protocols/playbook-call.md cite §Agent Collaboration's delegation contract instead of restating it
  • Deploy: rebuild the mcp-server image (tool descriptions are baked in); the prompt reaches agents with the backend deploy, no agent rebuild

Fixes abilityai/trinity-enterprise#568

🤖 Generated with Claude Code

webmixgamer and others added 3 commits October 5, 2026 14:24
… — never re-send on silence (Abilityai/trinity-enterprise#568)

The receipts already existed on every dispatch route (#914, #2661, #2670);
what was missing was the caller knowing what they mean. In the 2026-09-08
cascade every duplicate dispatch was an agent re-sending after "could not
confirm delivery". The contract is now ONE text, taught wherever a caller
reads:

- the platform prompt, §Agent Collaboration (DELEGATION_CONTRACT spliced into
  PLATFORM_INSTRUCTIONS) — every agent, every turn, no image rebuild;
- verbatim in the chat_with_agent and every dynamic chat_with_<agent>
  description (src/mcp-server/src/delegation_contract.ts) — the copy an
  external MCP client, or an agent at PromptTier.MINIMAL, reads; fan_out and
  send_message repeat its rule sentence and point at it;
- the receipt message itself: every MCP-authored receipt ends with "Do not
  re-send: read the outcome with get_execution_result(agent_name=…,
  execution_id=…)", chat() reuses queuedTimeoutReceipt, and runAgentChat
  rewrites the backend's REST-only "Poll GET …" line on accepted/queued.

Claude Code cuts every MCP tool description at 2,048 characters
(CLAUDE_CODE_MAX_MCP_DESCRIPTION_LENGTH). The old chat_with_agent description
was 2,424, so its async and list_recent_executions advice never reached a
model. It is now lead + modes + contract (~1.9 KB); per-mode detail moved into
parameter descriptions, which are not cut.

The text names only what the platform produces today: no `replayed` (ent#566
has not shipped it), the agent.task.* wake scoped to parallel runs (a
sequential /chat turn emits none), `retryable: false` → "do what its message
says" (gate refusals say "try again shortly"). The unverified
"concurrent-duplicate guard will kill mid-execution" claim is gone from the
description, the receipts and requirements §37.1 — no such guard exists.

Tests: tests/unit/test_ent568_delegation_contract.py (byte parity with the TS
copy via a strict JSON parse + guard-the-guard cases; every tool, argument,
status and event it names proven real; budget; AC4) and
src/mcp-server/src/delegation-contract.test.ts (the descriptions a real
tools/list publishes, under the cap; receipt messages on chat/task timeout,
409 replay, async accepted/queued and the pull-routed path). 15 wiring
mutations each go red. CI boot smoke also imports dist/tools/dynamic-agents.js.

Docs: requirements §37.5, architecture/mcp-server.md, six feature flows, the
agent guides' stale "will fail after 60 seconds" sections, user-docs.
Follow-ups filed: #3232, #3233, #3234.

Implements Abilityai/trinity-enterprise#568.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… out (Abilityai/trinity-enterprise#568)

AC1 lists `replayed` among the receipts. No dispatch route emits it until
ent#566 ships intent-scoped dedupe (only ask_operator answers `replayed`),
and the guard that checks every named status against a dispatch-route
producer would refuse it. A replay is still covered by "whatever its status"
and "an exact repeat is normally answered with the original" — observed live
on localhost: an identical repeat answered in 28 ms with the first call's
execution_id, and one task ran.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
webmixgamer and others added 2 commits October 5, 2026 17:03
…rding, history corrected (Abilityai/trinity-enterprise#568)

The PR #3240 review found three places where the change said more than
the code does.

- Contract, bullet 4: `agent_busy` is not proof that nothing ran. The
  SUB-003 usage-limit 429 is raised after the turn ran, and the MCP
  client labels it `agent_busy` (#3244). The bullet now sends that case
  through the same `list_recent_executions` check as every other error
  without an `execution_id`, then a word-for-word re-send after
  `retry_after_seconds`.
- Receipt leads: the backend stores the async receipt as the
  idempotency snapshot at dispatch time and replays it for 24 h (#3245),
  so "Accepted: the task is running" can be false on a replay. The leads
  now claim nothing about liveness. A replay of the backend's own
  sync-/task `queued_timeout` snapshot, which tells a REST caller to
  poll GET, is rewritten too; the MCP server's own `queued_timeout`,
  which already carries the read line, is left as it is.
- History: Claude Code's 2,048-character description cap has cut the
  `chat_with_agent` description since #2958 (2026-09-24), which took it
  from 2,035 to 2,424 characters. The receipt advice was delivered
  before that. Requirements, the learnings fragment and the CSO report
  now say so.

Tests: every `chat_with_*` tool must take `parallel` and `async`; the
dynamic-tool factory keeps the backend's line and falls back to the
exact default; each lead is pinned per status for agent, user and
self-task sessions; the backend `queued_timeout` replay is rewritten
and the MCP server's own is not. Ten mutations each turn a test red: a
dropped `async`, swapped or optimistic leads, a skipped rewrite per
session kind, a factory that ignores the backend line, both
`queued_timeout` branches, and TS-only contract drift.

Docs: execution.md names the sync `/chat` exception (it emits no
`agent.task.*` event), requirements §37.5 says why `replayed` is left
out, and follow-ups #3244 and #3245 are linked.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Abilityai/trinity-enterprise#568)

The history now says Claude Code cut the `chat_with_agent` tail before
the model read it, instead of "any model" or "every agent": the 2,048
cap was read out of the Claude Code binary, and agents on the Codex or
Gemini runtimes reach the same tools through MCP clients that were not
measured.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@webmixgamer webmixgamer added the ui PR touches the frontend UI — triggers Playwright e2e tests label Oct 5, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ui PR touches the frontend UI — triggers Playwright e2e tests

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant