feat(delegation): every caller is taught that a receipt means running — never re-send on silence (abilityai/trinity-enterprise#568) - #3240
Open
webmixgamer wants to merge 5 commits into
Open
webmixgamer wants to merge 5 commits into
webmixgamer wants to merge 5 commits into
Conversation
… — never re-send on silence (Abilityai/trinity-enterprise#568) The receipts already existed on every dispatch route (#914, #2661, #2670); what was missing was the caller knowing what they mean. In the 2026-09-08 cascade every duplicate dispatch was an agent re-sending after "could not confirm delivery". The contract is now ONE text, taught wherever a caller reads: - the platform prompt, §Agent Collaboration (DELEGATION_CONTRACT spliced into PLATFORM_INSTRUCTIONS) — every agent, every turn, no image rebuild; - verbatim in the chat_with_agent and every dynamic chat_with_<agent> description (src/mcp-server/src/delegation_contract.ts) — the copy an external MCP client, or an agent at PromptTier.MINIMAL, reads; fan_out and send_message repeat its rule sentence and point at it; - the receipt message itself: every MCP-authored receipt ends with "Do not re-send: read the outcome with get_execution_result(agent_name=…, execution_id=…)", chat() reuses queuedTimeoutReceipt, and runAgentChat rewrites the backend's REST-only "Poll GET …" line on accepted/queued. Claude Code cuts every MCP tool description at 2,048 characters (CLAUDE_CODE_MAX_MCP_DESCRIPTION_LENGTH). The old chat_with_agent description was 2,424, so its async and list_recent_executions advice never reached a model. It is now lead + modes + contract (~1.9 KB); per-mode detail moved into parameter descriptions, which are not cut. The text names only what the platform produces today: no `replayed` (ent#566 has not shipped it), the agent.task.* wake scoped to parallel runs (a sequential /chat turn emits none), `retryable: false` → "do what its message says" (gate refusals say "try again shortly"). The unverified "concurrent-duplicate guard will kill mid-execution" claim is gone from the description, the receipts and requirements §37.1 — no such guard exists. Tests: tests/unit/test_ent568_delegation_contract.py (byte parity with the TS copy via a strict JSON parse + guard-the-guard cases; every tool, argument, status and event it names proven real; budget; AC4) and src/mcp-server/src/delegation-contract.test.ts (the descriptions a real tools/list publishes, under the cap; receipt messages on chat/task timeout, 409 replay, async accepted/queued and the pull-routed path). 15 wiring mutations each go red. CI boot smoke also imports dist/tools/dynamic-agents.js. Docs: requirements §37.5, architecture/mcp-server.md, six feature flows, the agent guides' stale "will fail after 60 seconds" sections, user-docs. Follow-ups filed: #3232, #3233, #3234. Implements Abilityai/trinity-enterprise#568. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… out (Abilityai/trinity-enterprise#568) AC1 lists `replayed` among the receipts. No dispatch route emits it until ent#566 ships intent-scoped dedupe (only ask_operator answers `replayed`), and the guard that checks every named status against a dispatch-route producer would refuse it. A replay is still covered by "whatever its status" and "an exact repeat is normally answered with the original" — observed live on localhost: an identical repeat answered in 28 ms with the first call's execution_id, and one task ran. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This was referenced Oct 5, 2026
…rding, history corrected (Abilityai/trinity-enterprise#568) The PR #3240 review found three places where the change said more than the code does. - Contract, bullet 4: `agent_busy` is not proof that nothing ran. The SUB-003 usage-limit 429 is raised after the turn ran, and the MCP client labels it `agent_busy` (#3244). The bullet now sends that case through the same `list_recent_executions` check as every other error without an `execution_id`, then a word-for-word re-send after `retry_after_seconds`. - Receipt leads: the backend stores the async receipt as the idempotency snapshot at dispatch time and replays it for 24 h (#3245), so "Accepted: the task is running" can be false on a replay. The leads now claim nothing about liveness. A replay of the backend's own sync-/task `queued_timeout` snapshot, which tells a REST caller to poll GET, is rewritten too; the MCP server's own `queued_timeout`, which already carries the read line, is left as it is. - History: Claude Code's 2,048-character description cap has cut the `chat_with_agent` description since #2958 (2026-09-24), which took it from 2,035 to 2,424 characters. The receipt advice was delivered before that. Requirements, the learnings fragment and the CSO report now say so. Tests: every `chat_with_*` tool must take `parallel` and `async`; the dynamic-tool factory keeps the backend's line and falls back to the exact default; each lead is pinned per status for agent, user and self-task sessions; the backend `queued_timeout` replay is rewritten and the MCP server's own is not. Ten mutations each turn a test red: a dropped `async`, swapped or optimistic leads, a skipped rewrite per session kind, a factory that ignores the backend line, both `queued_timeout` branches, and TS-only contract drift. Docs: execution.md names the sync `/chat` exception (it emits no `agent.task.*` event), requirements §37.5 says why `replayed` is left out, and follow-ups #3244 and #3245 are linked. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Abilityai/trinity-enterprise#568) The history now says Claude Code cut the `chat_with_agent` tail before the model read it, instead of "any model" or "every agent": the 2,048 cap was read out of the Claude Code binary, and agents on the Codex or Gemini runtimes reach the same tools through MCP clients that were not measured. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
One text, taught wherever a caller reads it. The delegation contract — a receipt means the work is running; never re-send because a call timed out or its delivery could not be confirmed; read with
get_execution_resultor park withset_reminder; an error without anexecution_id,agent_busyincluded, is checked againstlist_recent_executionsbefore a word-for-word re-send;pending_approvalis not retried;parallel=true, async=truefor long work — lives byte-identically in the platform prompt (§Agent Collaboration: every agent, every turn, no rebuild) and insrc/mcp-server/src/delegation_contract.ts. Thechat_with_agentand every dynamicchat_with_<agent>description carry it verbatim;fan_outandsend_messagerepeat its rule sentence and point at it.Inside Claude Code's 2,048-character description cap. Claude Code shows the model only the first 2,048 characters of an MCP tool description (
CLAUDE_CODE_MAX_MCP_DESCRIPTION_LENGTH). bug(execution): a chat turn continues whatever session was persisted last — after a long scheduled run it inherits that run's context, pays its compaction, and the execution record says nothing #2958 (2026-09-24) took thechat_with_agentdescription from 2,035 to 2,424 characters, so since then Claude Code has cut its tail — thelist_recent_executionsadvice and "preferparallel=true, async=true" — before the model read it. Before bug(execution): a chat turn continues whatever session was persisted last — after a long scheduled run it inherits that run's context, pays its compaction, and the execution record says nothing #2958 the whole text arrived (it was 1,436 characters during the 2026-09-08 cascade); other runtimes' MCP clients were not measured. It now publishes at 1,940 characters (lead + modes + contract); per-mode detail moved into parameter descriptions, which are not cut.The receipt says it too. Every
execution_idreceipt achat_with_*caller gets carries "Do not re-send: read the outcome withget_execution_result(agent_name=…, execution_id=…)":queued_timeoutthatchat()andtask()write when they give up, and the 409 in-flight replay — each adds one sentence after it;accepted/queued, and a replayed sync-/taskqueued_timeout— whose REST-only "Poll GET …" linerunAgentChatreplaces. Those leads claim nothing about liveness ("Accepted by 'worker' as — it may still be running or already done."), because the backend replays its async receipt for 24 h, after the run may have finished or failed (bug: an async /task keeps replaying its accepted receipt for 24 h after the run fails, so an exact retry returns the failed run #3245).fan_out_timeoutkeeps its own bug(mcp): fan_out has no gateway-timeout receipt — the third route of the #914 class, and the one that runs longest #2670 wording, which says to pollget_fan_out_resultinstead of re-sending. The unverified "concurrent-duplicate guard will kill mid-execution" claim is gone — no such guard exists.Cost. Every dynamic
chat_with_<agent>tool carries the full contract, sotools/listgrows by about 1.6 KB per exposed agent. Deliberate: a client may read only the dynamic tool (AC2).Acceptance criteria
replayedis deliberately not named, although AC1 lists it: no dispatch route emits it until ent#566 ships intent-scoped dedupe (onlyask_operatoranswersreplayed), andtest_every_status_the_contract_names_is_emitted_by_a_dispatch_routeadmits only statuses a dispatch route really emits. A replay is still covered by "whatever its status" and "an exact repeat is normally answered with the original" — observed live: an identical repeat answered in 28 ms with the first call'sexecution_id, and one task ran. Recorded in requirements §37.5.chat_with_agentand every dynamicchat_with_<agent>tool (byte parity pinned by a test that parses the TS array).fan_outandsend_messagecarry the rule sentence and point at the contract.platform_prompt_service.py(the tool description is the single source; always-render only for turn discipline no description can carry; capability-fact phrasing; bare tool names; byte-identical VERBOSE). No new###section; §Agent Collaboration stays shorter than §Operator Communication (pinned).Changes
src/backend/services/platform_prompt_service.py—DELEGATION_CONTRACT, spliced into §Agent Collaboration; the contract's tools added to the Codex orientation.src/mcp-server/src/delegation_contract.ts(new) — the shared text,DELEGATION_RULE,readNotResend.src/mcp-server/src/tools/chat.ts,tools/dynamic-agents.ts,tools/messages.ts,client.ts— descriptions and receipt messages..github/workflows/mcp-server-test.yml— the ESM boot smoke also importsdist/tools/dynamic-agents.js(reached only fromindex.ts).core-agent.mdandmcp.md),architecture/mcp-server.mdandarchitecture/execution.md, the feature flows that teach delegation, the agent guides' stale "will fail after 60 seconds" sections, user docs.Review round (6ac2a69, a9217b0)
/reviewfound no critical issues and ten informational ones, six of them claims the code contradicts. All are fixed:agent_busyis not "nothing ran": the SUB-003 usage-limit 429 is raised after the turn ran, and the client labels itagent_busy(bug: chat_with_agent reports a usage-limit 429 as agent_busy although the turn already ran, so a retry repeats its side effects #3244). Contract bullet 4 now routes it through the samelist_recent_executionscheck. The new wording is the same length, so the published figures are unchanged./taskqueued_timeoutreplay is rewritten too (bug: an async /task keeps replaying its accepted receipt for 24 h after the run fails, so an exact retry returns the failed run #3245).readNotResend" is corrected:fan_out_timeouthas its own wording, and two receipts add a sentence after it./chatroute is the one path that emits noagent.task.*event.execution.md,task-completion-events.mdand the user docs now say so.async=truewithoutparallel=true, or omitagent_name.Coordination with #3238
#3238 adds a budget test over every published tool description with a shrink-only
PENDINGentry forchat_with_agentowned by this issue — and an entry that fits the cap fails. Whichever PR lands second deleteschat_with_agentfromPENDINGinsrc/mcp-server/src/tool-description-budget.test.ts.Test plan
cd tests && pytest unit/test_ent568_delegation_contract.py -v— 31 passednode --import tsx --test src/delegation-contract.test.ts(insrc/mcp-server) — 22 passed; full MCP suite 704/704;tscbuild + ESM boot smokeasync, swapped or optimistic leads, a skipped rewrite per session kind, a factory that ignores the backend line, bothqueued_timeoutbranches, TS-only contract drift) — each goes red/verify-local --skip-agent --skip-uniton a05b1fc — build + import smoke, boot, preflight, integration (70 passed, 13 agent/Postgres-gated skips, 2 known false fails). Stage 1's 42 order-dependent failures are identical on a cleanorigin/devworktree. Not re-run after the review round, whose only backend change is the text of one string constant.chat_with_agent1,940 /fan_out1,892 /send_message868 characters; async receipt wording; identical repeat answered with the originalexecution_id. Since then, bullet 4 (same length) and the receipt leads have changed./review— no critical findings, ten informational, all fixed (above);/cso --diff— no findings (docs/security-reports/cso-diff-2026-10-05-ent568-delegation-contract.md)Follow-ups filed while building this: #3232, #3233, #3234 (in review as #3238). From the review: #3244 (a usage-limit 429 is reported as
agent_busyafter the turn ran), #3245 (an async receipt is replayed for 24 h after its run failed).Debt adjacency, nothing paid down here:
chat.tscarries #1074 (removing the deprecatedtimeout_secondswill also drop the sync-parallel note moved into that parameter's description),messages.tscarries #3087,client.tscarries #1029.After merge
status-in-devon abilityai/trinity-enterprise#568 by hand — the status workflow does not promote a cross-tracker issuecanon/protocols/playbook-call.mdcite §Agent Collaboration's delegation contract instead of restating itFixes abilityai/trinity-enterprise#568
🤖 Generated with Claude Code