Skip to content

DO NOT MERGE — merge train: 3257,3258,3259,3260,3267,3217,3219,3250,3253 - #3270

Closed
trinity-ability wants to merge 49 commits into
devfrom
train/20261006-1038
Closed

trinity-ability wants to merge 49 commits into
devfrom
train/20261006-1038

Conversation

@trinity-ability

Copy link
Copy Markdown
Contributor

Integration surface for #3257, #3258, #3259, #3260, #3267, #3217, #3219, #3250, #3253. Never merged; members merge individually once green.

Trinity Agent (trinity) and others added 30 commits October 5, 2026 17:05
…rinity-enterprise#784)

Replaces `landingThread` with one pure rule, `agentLanding`, that every door
which has to RESOLVE a landing calls. ent#523 landed you in the chat you were
most recently active in; most visits to an agent start new work, so resuming
cost two actions every time. The default is now a new, empty chat.

Nothing is minted by the landing: the helper is pure and returns a null
session, and the row is born on the first send (`newThread`, ent#451), so
repeated visits accumulate no empty chats. An unused Main is already kept out
of the chat list by `sidebarThreadsOf` — the AC 4 guard, still green.

`agentLanding` takes `lastOpenSessionId` as the seam ent#621's agent-switch
keys will pass, honoured only when it still names a live, unarchived chat of
this agent in the principal's own thread list — so a stale or forged id falls
back to the default rather than landing somewhere it should not (#3140 class).
No Map is built here; ent#621 adds the state and the handler.

`resolveAgentLanding` keeps its signature and delegates, so the `?agent=` deep
link and the sidebar row cannot drift. `forceNew` stays accepted for existing
`?new=1` links and is now redundant rather than wrong.

Shell wiring (`landOnAgent`, the focus gate) follows in the next commit.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…lityai/trinity-enterprise#784)

Wires the shell to the ent#784 rule and gates the composer focus.

`landOnAgent` is now synchronous. The awaited `ensureMainListed` is gone from
the landing path — the rule needs nothing from the network — and with it the
overtake race that awaited round trip required. The pinned Main is still
minted on the first visit; `watch(activeAgentName)` owns that promise (ent#523)
and no longer blocks the landing.

The landed chat KEEPS `/workspace/a/:name` rather than escaping to bare
`/workspace`, so a reload, a bookmark or a copied link still names the agent;
the first send replaces it with the thread's own URL as today. An idempotency
guard makes the second watcher fire (the thread list arriving) a no-op, so it
cannot remount an unsent chat and throw away what was being typed.

`landOnAgent` now asks `guardLeaveCall` BEFORE touching `activeAgentName`,
which feeds `convKey`. Back/forward and a typed `/workspace/a/:name` reach the
landing without passing a click door, so they could end a live voice call
without a word (the ent#551 class; the click doors were already guarded).

Focus is gated on WHY the fresh composer mounted, not on the pointer alone.
A gesture (New chat, ⌘J, the agent picker, switch-agent) passes `always` and
focuses on any pointer, keeping #2579 AC 2 on Android. A landing passes
`fine-pointer` and focuses only where that cannot summon an on-screen
keyboard — the "unprompted" the issue names. `shouldFocusOnRestore` is renamed
`shouldAutoFocusComposer`: two reasons now share the one rule.

Tests: the focus gate is proven by MOUNTING PortalConversation under a stubbed
`matchMedia` (#2918), asserting `document.activeElement`, with focus shown to
be elsewhere first. Reverting either the rule or the gate turns 10 cases red on
behaviour — session ids and activeElement, not source text.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…y-enterprise#784)

Tiered docs for a feature change: the owning area file, the requirement, and
the feature flow.

- `architecture/workspace.md` — `agentLanding` is the one landing rule and what
  it now answers; the synchronous `landOnAgent`, the kept URL, the guard order,
  and the `focusOnMount` mode.
- `requirements/core-agent.md` — the ent#523 "One page" rule amended rather
  than rewritten: what reversed, and that landing mints no row.
- `feature-flows/workspace-agents-at-the-centre.md` — the rule table, the "a
  tab is not a landing" note (the strip is now the only way back), the two
  doors that RESOLVE vs the gesture doors that ASSERT, and a `## Changed by
  ent#784` section rather than edited history. Names the follow-up: backend
  adoption of an empty Main on `new_thread=True` is not done here.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…ted chat (Abilityai/trinity-enterprise#784)

ent#784's landing stays on `/workspace/a/:name`, so the watcher that resolves
that route can now re-fire with the route unchanged — once per thread-list
refresh — for as long as the person stays there. `landOnAgent`'s idempotency
guard keys on `startingNewChat`, which the first send clears, so a refresh
settling in the window between `onSessionAdopted` and the asynchronous
`router.replace` nulled `pendingSession` and bumped `convGen`, remounting an
empty composer over the thread whose first reply was streaming (and handing the
`/c/:id` watcher a #3140 "chat isn't available" for the chat just created).

Guarded in the watcher instead: a list-only re-fire is a no-op, keyed on "the
route did not change on this fire" AND on having already landed this name. Both
clauses are load-bearing — the name alone swallows the cold deep link, whose
only fire IS the list arriving with the param already in place. Deliberately
NOT keyed on `pendingSession`: back/forward from `/workspace/c/:id` arrives
with a session set and must still land fresh (T4).

New behavioural pin `portalAgentLandingRemount.mount.spec.js` (shallow-mounts
Portal.vue, asserts the conversation's props and instance identity, not source
text): the adopted-session case went red before this change with
`sessionId: null`; the three regression cases — composing, same-agent
back/forward, cold deep link — were green before and stay green.

Also fixes `portalUnavailableTargets.mount.spec.js`, which still asserted
ent#523's landing (`/workspace/a/scout` → `/workspace/c/t1`) and was failing on
this branch; it now proves the page opened by the conversation on screen.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…source-text pins (Abilityai/trinity-enterprise#784)

`landOnAgent` was pinned only by regexes over its own function body
(`workspaceNewChat.spec.js:181-211`), so nothing executed it: an inverted
guard, or the right lines in the wrong order, reads byte-identically to a
regex. The shell is mount-testable (`portalUnavailableTargets.mount.spec.js`),
so these assert STATE instead.

Decision #13 (ent#551 class) — back/forward reaches the landing without passing
a click door, and `activeAgentName` feeds `convKey`, so a live call must be
asked about BEFORE anything is written: mid-call, the dialog opens and the chat,
its props and its instance identity are untouched; confirming then performs the
landing it deferred; cancelling leaves call and chat exactly as they were.

T4 — one navigation to an agent page mounts exactly ONE conversation, however
often the thread list moves underneath it.

Mutation-verified, both directions: with `guardLeaveCall` removed from
`landOnAgent`, the two mid-call cases go red; with the T4 guards removed, the
composing and one-mint cases go red. The source-text pins are kept as
supplements (they still pin "no await / no ensureMainListed / no router.push"
inside the function, which state cannot see).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…inity-enterprise#784)

F1 of the operator's 2026-10-05 reopen. `agentLanding` always answered a new
chat, so a draft typed into an existing chat was not where opening the agent
landed — and the draft mark on the agent's row pointed at words that clicking
it could not get back to. `e2e/workspace-drafts.spec.js:92` and `:116` describe
the wanted behaviour and were red at c8af3d7.

The rule is now the ruling's precedence: a link naming a chat (never reaches
here), then the ent#621 `lastOpenSessionId` seam unchanged as the first arm,
then the agent's chat holding an unsent draft — newest `updatedAt`, with the
unsaved `new:<agent>` chat weighed on the same clock — then a new chat. A
`new:` winner is `sessionId: null`, which is where those words already live; a
thread winner opens that thread and its remount restores the draft.

The candidate set is the set that lights the sidebar's mark: `isDraftedThread`
is shared by `agentsWithDrafts` and the new `draftedLandingFor` rather than
copied, so "the mark means click here to continue" cannot drift — a room and an
archived thread are excluded on both sides. The drafts map is passed IN by both
doors (`landOnAgent`, `resolveAgentLanding`), never read inside the rule, so
one rule serves every door and stays pure.

`composerFocusMode` moves above the branch in `landOnAgent`: a landing on a
drafted chat is the same kind of arrival, and `openThread` does not touch that
ref, so a previous gesture's `always` would otherwise have made the next
landing focus on a touch device (T3(b)). The restore's own caret already shares
`shouldAutoFocusComposer`. The T4 remount guard is untouched: the watcher still
returns on a list-only re-fire, and the drafted-thread arm leaves the agent URL
through `openThread`, so the guard is never reached with a stale answer.

Tests: `portalDraftLanding.spec.js` (each arm, newest-wins both ways, the tie,
room/archived exclusion, and the property that every MARKED row is a landing
candidate) and `portalDraftLanding.mount.spec.js`, which replays both e2e flows
at the shell's seam — red at c8af3d7 on the e2e's own symptom, 2 of 3 failing
with `sessionId: null` where the draft was.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…y-enterprise#784)

F2 of the operator's 2026-10-05 reopen: "opening an agent must not create a new
chat every time; if an empty chat with that agent already exists it is reused,
and at most one empty chat per agent exists at any time."

Arm 4 of the precedence, under the drafts arm and above a new chat.
`agentEmptyChat` is the read side: an unarchived, non-room thread of this agent
with no message sent — ent#523's own "unused Main" test, widened to any row
that fits it. Both fields are required because the two reads disagree: the
cross-agent batch omits `message_count` (so an absent count must not read as
used) while the per-agent read carries it (so a count of 2 must not read as
empty just because `last_message_at` is missing). Several empty rows — legacy
data — resolve Main first, then newest `created_at`, then id, so list order
never decides and two doors reading one list cannot reuse two different rows.

`ensureMainListed` is the write side, and the only place the "at most one" half
can be held: it is a GET that INSERTS (`list_sessions` → `ensure_main_session`),
so for an agent with an empty chat but no Main — legacy, since nothing mints a
non-Main empty row today — visiting it would add a SECOND empty chat and then
land on one of the two. It now returns early for that case. #2579's "the pinned
tab has to be there" still holds for every agent whose chats are all used,
which is the case it was about; that is pinned as its own mount case rather
than left to the comment. No backend change: nothing here asks the server for
anything it does not already do.

Three existing specs carried fixtures with neither message field, which under
the old rule meant "a thread exists" and under the ruling means "an empty chat
to reuse" — the assertion they were written for. Each is updated to a USED row
and the empty case is asserted separately, so none of them went green by
accident: `portalAgentsAtCentre` (the landing arms), `workspaceAgentLanding`
(the `?agent=` door, both id shapes) and `workspaceNewChat`, where `?new=1` is
load-bearing again — it is the one way past the whole precedence, so the two
answers now have to differ.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…and the caret (Abilityai/trinity-enterprise#784 review)

ent#784 made /workspace/a/:name the URL a landed new chat RESTS on. Three
things had only been true because nobody stayed there; each was reproduced
in a browser against the rebased branch.

- Clicking the row of the agent whose new chat is already on stage cleared
  `startingNewChat` in `openAgentPage` and then pushed the URL it was
  already on, so the landing never re-ran. The composer still read "New
  chat" and sent its first message without `new_thread`, which the server
  resolves to the agent's Main chat. The click is now a no-op that hands
  the caret back.
- The rail rules still read the agent URL as "a page, not a conversation",
  so a landed new chat had no rail, and the first send (which moves the URL
  to /workspace/c/:id) slid one in beside a reply mid-stream. The URL is
  rail-free only until the landing has put the named agent on stage, and
  its column is reserved mid-load like any 1:1 route.
- A landing that reuses a chat (the agent's empty one) goes through
  `openThread`, and a thread mount focuses nothing, so the caret stayed on
  the sidebar row. The shell hands it over under the landing's own
  fine-pointer rule.

Tests: portalAgentLandingStage.mount.spec.js (10 cases, mounted shell).
Four were red before the fix, for these reasons.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…ine (abilityai/trinity-enterprise#798)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…inity-enterprise#797)

Hero alt text, H2, lede, "The solution" line and footer move onto the
ADR-0011 line; the enterprise-modules sentence drops the bare "SSO" (T3).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ne form (abilityai/trinity-enterprise#797)

install.sh banner (width kept at 60), AWS/DO listing titles, AMI
ami_description, CFN Description first line, DO/Vultr MOTD taglines, and
the CLI docstring, pyproject description, Homebrew desc, src/cli/README
and docs/CLI.md now read "the operating system for the AI-native company".
Description strings only; no other value changed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ilityai/trinity-enterprise#797)

The /setup lede now opens with the ADR-0011 line; the CTA and the
Governed/Auditable/Your-infrastructure chips are unchanged. Text only.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…011 line (abilityai/trinity-enterprise#800)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…0011 line (abilityai/trinity-enterprise#800)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…lines (abilityai/trinity-enterprise#801)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…facts line (abilityai/trinity-enterprise#800)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…here is no room above (#3264)

On the Workspace agent band the activity chart is 23px tall, sits just
under the page header and inside ScanlineReveal's clip-path. The
tooltip was an `absolute bottom-full` child of the bar, so it always
opened upward and its top rows (the date, the top buckets) were covered.

StackedBarChart now teleports one tooltip to <body>, position: fixed,
placed from the hovered bar's rect: above when it fits, below when it
does not, clamped 8px inside the viewport on every side. It stays
hidden until measured and follows the bar through scroll and resize.
Applies to every caller (Workspace band, agent Overview, canvas charts).

Fixes #3264

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…se) answer (#3242)

The #2376 sink now handles SOMETHING_ELSE = "(something else)" before
membership: accepted on any approval with an instruction in response_text,
refused by name otherwise (instruction_required, reserved_value), and refused
on platform-minted / gate approvals (not_off_menu). Every other unoffered
string stays refused. validate_response_choice takes response_text
keyword-only and required.

- ask_service: gate refusal; ask_operator refuses the literal as an option
- both writers map ReservedAnswerError to a named 422
- OperatorResponse.response_text bounded at 4000
- Workspace projection: decided_by_options boolean (T8 ruling)
- recent answers render "Something else: <instruction>"
- resume frame: platform sentence above the data fence

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…#3215)

`_build_payment_required` passed no `scheme=` to the SDK helper, so every 402 —
and every verify and settle built from the same helper — advertised the SDK
default `nvm:erc4337`. The facilitator is POSTed that requirements document
verbatim, so `accepts[0].scheme` is the only channel the scheme travels
through: a fiat (card) plan's token was always checked against crypto
requirements and rejected. Card plans were unpayable.

- `PlanScheme(scheme, network)` + `default_plan_scheme(env)`; the erc4337
  environment map stays an explicit frozen literal (a payments-py bump must not
  move a live agent's network, #3216).
- `_build_payment_required(..., plan_scheme=None)` — `None` is the pre-#3215
  document, so crypto bytes are unchanged per environment (golden test).
- `scheme_from_token` reads `accepted.{scheme,network,planId}` off the presented
  token, allow-listed against the SDK's scheme/network vocabularies and checked
  against this agent's plan id. This is what verify and settle use: it removes
  every network call from the money path and makes a settle re-driven hours
  later byte-stable with its own verify.
- `resolve_plan_scheme` (402 and `/info` only) parses `plans.get_plan` with the
  SDK's own keys — parity-tested against `resolve_scheme`/`resolve_network` —
  behind a per-(env, plan_id) cache, single-flight futures, its own 2-slot gate
  (never a facilitator slot, so an anonymous 402 flood cannot starve paying
  traffic), stale-while-revalidate and a WARN once per negative window. The SDK
  swallows these failures at DEBUG, which is how a card plan silently stayed
  crypto with nothing in Trinity's logs.

61 new tests; the 196-test ent#679/#3185 baseline stays green unchanged.
Mutations proven red: dropping `scheme=` (5 red), swapping settle's
token-derived scheme for a plan lookup (2 red).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…3215)

AC5. The x402 402's `resource.url` is what a buyer's token is minted and
verified against, and the facilitator compares origin+path — so one wrong scheme
character mints a token for a URL the client never calls. Both payment doors
built it from `request.base_url`, which is http on a standard Trinity
deployment for two independent reasons:

- `docker-compose.prod.yml` / `docker-compose.hosted.yml` override the image
  `command:` and drop the Dockerfile CMD's `--proxy-headers
  --forwarded-allow-ips=*`, so uvicorn does not trust `X-Forwarded-*` at all and
  `request.url.scheme` stays http behind ANY proxy. Restoring those flags is a
  trust change (it re-prices `request.client.host` for every per-IP limiter, and
  agents on the agent network reach `backend:8000` directly), so it stays
  deferred — which is why `utils/public_url.py` reads the RAW header and never
  `request.url.scheme`.
- the frontend nginx forwarded `X-Forwarded-Proto $scheme`, and `$scheme` is
  http on that hop (TLS terminates upstream), clobbering the real https before
  the backend saw it. Now a `$fwd_proto` map preserves an upstream https and
  otherwise falls back to this hop's scheme, on all three proxied locations
  (`/api/`, `/a2a/`, `/mcp`) — a door that lied about the scheme on one of them
  would mint tokens for the wrong origin just the same.

`public_base_url(request, configured=, frontend_url=)` owns the precedence: the
operator's declared origin (Settings `public_chat_url` → `PUBLIC_CHAT_URL` →
`FRONTEND_URL`) when it is the host the caller actually used, else the request
host with an https upgrade — never a downgrade — from the raw header. The
same-host restriction keeps a caller that reached a private host from being sent
to a public one whose narrow tunnel may not route that path. One helper now
serves the agent card, the paid door and the A2A door, so a single request
yields a single origin; `_base_url_from_request` also starts honouring the
Settings row, as the Telegram/WhatsApp webhook URLs already do.

`GET /api/paid/{name}/info` — the card's `paymentInfoUrl`, from which a
payment-aware client pays on its first request — now loads the config WITH the
key so it can resolve the plan's real scheme, and emits an absolute
`resource.url`. The key is used server-side only; a test pins that it never
reaches the body.

34 new tests; 291 green across the #3215 files and the ent#679/#3185 baseline.
Mutation proven red: removing the header upgrade → 9 red.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…e) (#3242)

- platform prompt (both authored copies) + agent guide: the approval bullet
  and the queue-file write-back say what the reserved value means and that
  the instruction is in response_text; never list it as an option
- test_1402 SENTINELS += "(something else)"
- MCP: types.ts exports SOMETHING_ELSE; ask_operator, get_my_ask and
  respond_to_operator_queue descriptions name it; the respond tool keeps
  `error` and adds the backend's {status, code, message, offered_options?}
  (T5, additive)
- py <-> ts parity test for the literal

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- utils/operatorQueue.js: SOMETHING_ELSE (mirrored, parity-tested),
  offeredChips (drops the literal and the size-cap marker), decisionLabel,
  decidedByOptions; buildQueueResponse needs an instruction with the literal
- QueueCard + QueueItemDetail: chips from offeredChips, a 'Something else'
  chip outside the v-for (hidden on gate approvals); typing with no pick arms
  it, the label and Send copy flip; Enter sends only after an explicit pick;
  maxlength 4000; Send rule is the shared builder
- respondToItem returns true/false; the detail panel clears only on success
- ResolvedCard + detail resolved view render 'Something else', never the literal

Raw-colour counts unchanged (QueueCard 28/59, QueueItemDetail 26/55).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- /m: a 'Something else' chip in the options group (reuses ops-option-btn,
  hidden on gate approvals) opens the form; the consequence line says none of
  the options will run; Send reads 'Send instruction' and stays disabled until
  the instruction is typed; maxlength 4000
- PortalAsks: the chip after the offered chips, hidden when the projection
  says decided_by_options (T8 ruling); typing with no pick arms it, placeholder
  and Send aria-label flip; Enter never sends the auto-armed state; submit()
  sends the armed option through the shared builder. Diff kept to the approval
  template, helpers beside pick(), the body of submit() and one import line
  (open PR #3181 edits other hunks)
- portalAsksTestidPrefix: the id inventory gains the chip's id

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Document what #3242 built: the reserved (something else) decision an
approval accepts besides its own options, with the instruction in
response_text; the named 422s (instruction_required, reserved_value,
not_off_menu, invalid_options at raise); the Something else chip on the
desktop card and detail, /m and the Workspace, hidden on gate approvals
on every surface (decided_by_options on the Workspace projection); the
resume frame and contract text. Corrects every doc that still said an
approval's answer must be one of its options.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…dead optionsOf import (C1, I1) (#3242)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…-miss literals; public decided_by_options; learnings (I2 I4 I5 I6 I7 I12) (#3242)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…tor projections; gate- is only the fallback (I3) (#3242)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…lookalikes on the native raise (#3243)

One pure predicate in operator_queue_choices (shared with the file path in the
next checkpoint): at most 5 options, the reserved (something else) never
counted; at most 60 characters per option; an option reading as the chip
("Something else", any case, parenthesised or not) refused as invalid_options.
Agent raises also get a hard 120-character title (title_too_long). Checked
after the replay lookup, so a pre-cap ask's retry still replays, and before
the rate cap, so a refusal spends no token. Env-overridable, floored at load
(2 options / 16 chars / 40-char title).

Mutation: with the _refuse_over_caps call removed, 9 tests in
test_3243_atomic_asks.py go red (TestNativeRaise refusals).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…on the marker (#3243)

A NEW file entry over the option caps (or a chip lookalike) is held as
invalid_options, an over-long title as invalid_title — the same predicate as
the native raise, after the id check and before the depth/rate caps, so no
rate token is spent. The ids (<=10) and the limits ride on
platform.ingestion beside `reason`, so a self-clearing queue_full cannot hide
them; _marker_differs now compares the key set so a fixed entry's id drops
off. Cap holds never increment `held` (no "runaway or compromised agent"
flood alert). Rows already ingested are never re-judged.

Also replaces the CP1a env test's module reload (it handed other suites a
stale OperatorQueueSyncService) with a test of the loader itself.

Mutation: with the hold disabled, 4 TestFileHold tests go red.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Trinity Agent (trinity) and others added 18 commits October 6, 2026 10:48
#3243)

.env.example and both compose files gain OPERATOR_QUEUE_MAX_OPTIONS (5),
OPERATOR_QUEUE_OPTION_MAX_CHARS (60) and OPERATOR_QUEUE_ASK_TITLE_MAX_CHARS
(120). security.md §26, the operating-room flow (step + revision row) and the
api-endpoints catalog name the new codes, where they run, and the file hold.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…se so parity is green (#3243)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…irst, fold NFKC/zero-width chip lookalikes, register the test (#3243)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ion docs and add the §26.7 authoring-caps bullet (#3243)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… cap (#3243)

Claude Code shows a model only the first 2,048 characters of a tool
description (#3234); ask_operator was 2,676 at 3c14c23, so its last rules
(the refusal codes and their remedies) were invisible to every agent.

The description is now 1,710 characters, ordered by what a model acts on
first: what the tool does and the fire-and-park contract, then the five
atomic-ask rules, then the reserved (something else) rule. Field detail
moved into the parameter descriptions, which are not cut: the receipt
fields, differs and idempotency (request_id); title_too_long and its fix
(title); the bad-then-good option example, options_required,
too_many_options / option_too_long / invalid_options with their fixes
(options); field_too_large (question, context, proposal); role_unassigned
(to); how an ask ends (expires_at); reask_requires_link
(supersedes_expired).

Tests: the published descriptions of ask_operator, get_my_ask and
respond_to_operator_queue are <= 2,048 over a real listTools, and every
moved phrase is published on its field; the unit test keeps ask_operator
<= 1,800 for headroom.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…lus docs (#3215)

Operator scope-add (2026-10-04T11:48:59Z): a settled `message/send` returned a
Task with `metadata: null` while `a2a-protocol.md` promised "a normal A2A Task
whose metadata carries the payment status and a receipt". The metadata WAS built
and placed — on `status.message.metadata`, which is the spec location (the
provider SDK's own `X402A2AUtils` reads exactly there) — but top-level
`Task.metadata` was never set, so a typed `a2a-sdk` client rendered a charged
turn as `metadata: null`.

`_task_object` now mirrors it from the same assignment, the same dict object, so
the two locations cannot drift. The docs name `status.message.metadata` as the
primary location and the mirror as the compatibility half. The free path still
emits no `metadata` key at all, so an unpaid Task's bytes and every idempotency
snapshot built from them are unchanged.

Docs:
- `user-docs/integrations/a2a-protocol.md` — names both locations precisely.
- `user-docs/integrations/nevermined-payments.md` — crypto vs card plans and the
  scheme/network each advertises; why the 402's `resource.url` origin matters and
  which knob fixes it.
- `user-docs/guides/deploying/public-access.md` — the A2A door was published by
  #3213 and the narrow tunnel table had no row for it. Added with ANCHORED rules
  (`^/a2a/[^/]+$`, `^/a2a/[^/]+/\.well-known/agent-card\.json$`): an unanchored
  `/a2a/*` would also match `/api/agents/{name}/a2a/call`, Trinity's authenticated
  OUTBOUND caller, which has no business on a public hostname.
- `memory/requirements/public-access.md` §23.2–23.3, `memory/requirements/mcp.md`
  FR-4 + new FR-4a, `memory/feature-flows/nevermined-payments.md` (two new
  sections + the two new knobs), `memory/architecture/backend.md` one-liners for
  `nevermined_payment_service.py`, `paid.py` and `a2a.py`.

11 new tests; 382 green across the #3215 files and every payments/A2A neighbour.
Mutation proven red: removing the mirror → 6 red.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…st (#3215 review I3)

#3215 made ONE origin rule serve the agent card and both payment 402s. The
doors needed same-host — an x402 token is minted and facilitator-verified
against `resource.url`, so a 402 quoting another origin is unpayable. But the
card is the opposite kind of document, and same-host regressed it: the
`get_agent_a2a_card` MCP tool proxies the card route from `backend:8000`, so a
card fetched there advertised `http://backend:8000/a2a/{name}` to every external
buyer. That is dev's behaviour broken on an existing surface, and the plan's own
T5 text named configured-first as "keep the card as is".

Still one module and one upgrade rule, now one keyword apart:

* `public_base_url(..., configured_wins=False)` — default unchanged, so every
  door's bytes are identical. `True` returns the declared origin whatever host
  the caller used, and falls through to the SAME request-host + upgrade-only
  raw `X-Forwarded-Proto` logic when nothing is configured.
* `routers/a2a._card_base_url` wraps it for the two card routes only (the
  authenticated per-agent card and the public well-known card). The A2A
  JSON-RPC door, the paid door and `/info` keep T5 exactly as built.

Safe because the two only differ where it cannot matter: a buyer that follows
the card arrives on the configured host, so the door's own same-host rule mints
that 402 for that very origin — the 402 is byte-identical to the URL the card
sent it to (pinned by a test). With nothing configured the two are the same
function.

Tests: 14 new in `test_3215_public_base_url.py`, 10 red before the fix —
`TestConfiguredWins` at the helper, `TestCardAndDoorOverTheRealRoutes` over the
real routes (both card surfaces and the anonymous 402 on one internal host,
which is the only layer that can catch a call site passing the wrong
precedence), and `TestOneOriginPerRequest` adjusted to the new rule rather than
trimmed. The route-level fixture imports `dependencies` at collection time on
purpose: `tests/unit/conftest.py` pops it from `sys.modules` between tests, so a
fixture-local import makes `dependency_overrides` stop matching after the first
test. 405 green in the #3215/#3185/ent#679 batch (391 before).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…3215 review I4)

The single-flight leader's `finally` answered its followers either way, but with
the WRONG answer when it had nothing: a cancelled leader (its caller
disconnected, its request shut down) completed the shared future with
`default_plan_scheme(...)`. Every follower then returned crypto for a card plan
with no WARN, no negative window opened, and last-known-good never consulted —
the exact silent mis-advertisement #3215 exists to stop, reintroduced at the one
place the service is not looking.

A leader that resolved nothing now sets `PlanLookupLeaderLost` on the future
instead. Each follower already catches it, so it falls into its existing
`_record_plan_lookup_failure` branch: the 402 still gets an answer (fail-open is
the accepted design), but it is last-known-good when there is one, the negative
window opens, and the WARN names the reason. The leader's own `CancelledError`
still propagates unchanged — `finally` swallows nothing. The future's exception
is marked retrieved, so a cancelled 402 with no followers does not log "Future
exception was never retrieved" at GC.

Tests: 4 new in `test_3215_plan_scheme.py::TestCancelledLeaderFailsFollowersLoudly`
— the WARN and the negative window, last-known-good preferred over the default,
no asyncio noise with zero followers, and the successful-leader path proved
untouched. 2 go red with the fix reverted (verified against a scratch copy, then
restored byte-identical); the other 2 are the no-regression arms. 409 green in
the #3215/#3185/ent#679 batch.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…der (#3215)

`svc2._get_payments_client = svc._get_payments_client` handed the leader
svc's delay-less `_Plans`, so its lookup finished in microseconds and
whether `leader.cancel()` landed first was a thread-vs-event-loop race:
green on Linux CI, red on every macOS run. The cache and the in-flight
registry are module-global, so last-known-good is shared without sharing
the client — drop the override and let the 0.2 s delay do its job.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ui PR touches the frontend UI — triggers Playwright e2e tests

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants