diff --git a/docs/adr/0022-own-a-truthful-deep-harness-seam.md b/docs/adr/0022-own-a-truthful-deep-harness-seam.md index 97c14f60..accb7f8a 100644 --- a/docs/adr/0022-own-a-truthful-deep-harness-seam.md +++ b/docs/adr/0022-own-a-truthful-deep-harness-seam.md @@ -36,7 +36,11 @@ clarification; `interrupt` means Harness-confirmed termination of the active Tur and ending an Interactive agent step remains a Crucible control above the Seam. Expected races return an accepted or rejected receipt rather than throwing. Acceptance does not prove final effect: after an accepted interrupt the Adapter rejects new inputs, drains to native terminal evidence, and the Turn result confirms whether interruption occurred. Harness Requests are Turn-scoped, independently keyed, may coexist, and expire when the -Turn ends, is interrupted, or is lost; an ordinary assistant question at a Turn boundary is not a Harness Request. +Turn ends, is interrupted, or is lost; an ordinary assistant question at a Turn boundary is not a Harness Request. (Edited 2026-09-29: +[ADR 0035](./0035-interrupt-ends-only-the-turn-and-a-mid-turn-message-is-a-native-steer.md) widens `steer` to native mid-Turn delivery of the human's text at the Harness's +next boundary, keeps a Turn open until every accepted Steer is delivered, adds a delivered-or-dropped Steer lifecycle to the event stream, drops +undelivered Steers on interrupt, and qualifies Claude Code's `interrupt` control request as a confirmed active-Turn interruption with a process-stop +fallback.) Native conversation identifiers are opaque recovery coordinates, never Run truth. When one is observable before submission, the Adapter awaits a Crucible-owned durable recorder before sending content; recording failure proves `not-started`. When a Harness reveals it only after acceptance, the diff --git a/docs/adr/0035-interrupt-ends-only-the-turn-and-a-mid-turn-message-is-a-native-steer.md b/docs/adr/0035-interrupt-ends-only-the-turn-and-a-mid-turn-message-is-a-native-steer.md new file mode 100644 index 00000000..870df23d --- /dev/null +++ b/docs/adr/0035-interrupt-ends-only-the-turn-and-a-mid-turn-message-is-a-native-steer.md @@ -0,0 +1,85 @@ +# Interrupt Ends Only the Turn, and a Mid-Turn Message Is a Native Steer + +An **Interrupt** stops the live **Turn** and nothing more: the **Step Attempt** stays open, the **Run** waits on the human, and the human's next +message continues the same **Harness Session**. A message the human sends while a Turn is working is a **Steer** in both agent Step kinds, delivered +natively at the **Harness**'s next boundary, and the Turn does not end until every Steer has been delivered. The need came from the pre-public-release +reports ([#235](https://github.com/secantdev/secant/issues/235)): in the native clients Escape stops the Turn and the human types on in the same +Session, and a message typed mid-Turn is picked up without interrupting, while Secant halted the Run on Interrupt, offered Claude Code no Steer, and +refused a mid-Turn send. This supersedes the Run lifecycle glossary's Interrupt rule (the Attempt ends `cancelled` and the Run rests `halted`) and its +Steer definition ("not a new Turn, and unsupported Harnesses do not emulate it"), [Spec: M3](https://github.com/secantdev/secant/issues/107) stories +18 and 19 and its deferral of a Session-preserving Claude Code interrupt, and [Spec: M4](https://github.com/secantdev/secant/issues/137) story 19's +Codex-only Steer. It amends [ADR 0022](./0022-own-a-truthful-deep-harness-seam.md)'s `steer` and `interrupt` controls and Claude Code's interruption +mode. Unchanged: closing Secant or Ctrl+C halts the Run ([ADR 0019](./0019-failed-and-halted-runs-are-resumable-resting-states.md)), Cancel is the +only route to `cancelled`, an Agent call made in an interrupted Turn is dropped ([ADR +0032](./0032-let-opted-in-interactive-agent-steps-accept-agent-declared-completion.md)), and headless gains no Interrupt or Steer. + +**Interrupt, per Step kind.** In an **Interactive agent step** the interrupted Turn ends `interrupted` and the Run returns to `blocked` on the human, +whose next message goes into the same Session; no Attempt is published and the Iteration stays open. An interrupted **Entry Turn** is treated the same +way and is never re-sent. In an **Agent step** the interrupted Turn ends `interrupted`, the Attempt stays open, and the Run is `blocked` on the +human's message. That message starts a human-origin follow-up Turn in the same Session and the same Attempt, and the Attempt takes its outcome from +its last Turn: a clean end completes the Step and the Routing advances without the human, exactly as an uninterrupted Agent step does, and a failure +takes the Step's ordinary retry policy. A second Interrupt is the only way to hold the Step again. A `lost` Turn keeps its current meaning. When +Secant closes while a Run waits after an Interrupt, the Run halts as any live Run does, and resuming returns it to waiting on the human's message. + +**Steer.** Steer now means native mid-Turn delivery of the human's own text, landing at the Harness's next boundary; it is offered in Agent steps and +Interactive agent steps alike, wherever the profile's Steer evidence says the Harness supports it. Codex delivers it through `turn/steer` with +`expectedTurnId` and a Secant-minted `clientUserMessageId`, which comes back as the `clientId` of the `userMessage` item written when the text is +taken into history before the next model request. Claude Code delivers it as a stdin `user` frame on the existing `claude -p --input-format +stream-json` process, stamped with a Secant-minted `uuid`: a frame written while a tool round runs, including one waiting on a tool approval, is taken +with the round's last tool result in the same exchange and listed in `result.user_message_uuids`, and one written while text streams runs as the next +native exchange with its own `result`; `command_lifecycle` frames, advertised as `msg_lifecycle_v1`, say which frames are still queued. Delivery means +the Harness put the text in front of the model, never that the model followed it. + +**The Turn stretches until every Steer is delivered.** A Turn ends at the first Harness boundary after which no accepted Steer is pending, so one +Secant Turn may span several native exchanges. The human's message therefore always reaches the agent before an Agent-declared completion from that +Turn is applied. Codex can accept a Steer in the few milliseconds between its last check for pending input and the end of the Turn, or on a failing +Turn, write it into history, and complete without answering it; a Steer sent during its Stop hooks is answered, because pending input re-runs the +Turn. The Codex Adapter detects that leftover, a `userMessage` item with the Steer's `clientId` followed by `turn/completed` with no model output, and +re-delivers it inside the same Secant Turn with a native `turn/start` whose input is empty. Recorded on codex-cli 0.157.1, that start is accepted on +the idle thread and the model answers the leftover from history, which holds the text once. It is sent only on a detected leftover, since on a thread +with nothing pending it makes the model repeat itself or invent work. A Codex that refuses an empty input falls back to re-sending the same text, +which answers too but leaves the message in history twice. Either way the input is the human's own message, so a Steer is answered before its Turn +ends on both Harnesses and this is ordinary turn-taking, not emulation. A Steer that arrives after the Turn has ended is rejected and its draft kept: +in an Interactive agent step the human sends it as the next Turn, and in an Agent step that has advanced the human is told it was not delivered. + +**Interrupt drops what has not been delivered.** Every accepted Steer still pending when an Interrupt lands is dropped and its text returned to the +human's compose, so nothing runs after the human said stop. Codex discards pending steered input on `turn/interrupt`; Claude Code is interrupted with +`cancel_queued: true`, which hands the queued frames back instead of running them as the next Turn. + +**How Claude Code stops.** Claude Code is interrupted with a raw `control_request` `interrupt` written to the same stdin, which, recorded on Claude +Code 2.1.284, answers without the SDK's `initialize` in milliseconds, ends the Turn with a `result`, keeps the process and Session live for the next +frame, keeps the partial text in context, and stops the foreground tool. It holds while a tool waits on Secant's permission bridge too: Claude Code +cancels the pending approval call with an MCP `notifications/cancelled`, which expires the Harness Request, and the tool never runs. That makes Claude +Code's interruption a confirmed active-Turn interruption rather than a process-only stop. Qualification bounds the dependency: a Claude Code that does +not answer the request falls back to today's process stop, SIGTERM of the process group on POSIX and `taskkill /T /F` on Windows, and `--resume`, +which loses partial streamed text. This is the same wire and the same degrade rule [ADR +0034](./0034-choose-and-change-model-and-effort-as-one-run-wide-model-choice.md) adopted for `set_model`; adopting the Agent SDK itself stays rejected +for the reasons in [Establish Claude Code's viable structured transports](https://github.com/secantdev/secant/issues/3). An Interrupt stops the Turn's +foreground tool work; a shell the model moved to the background may outlive it until the Session closes. The same held on Windows +([recorded](../research/windows-live-interrupt-and-steer.md) on Claude Code 2.1.283 and codex-cli 0.155.0), with one limit: a script the Bash tool +runs through Git Bash leaves the process tree, so neither Claude Code's stop nor Secant's `taskkill /T` reaches it, and it can outlive the Interrupt +and the Session. The Turn stop itself is still confirmed, Claude Code's Windows interruption evidence states the limit, and containment is decided +separately on [Decide how Secant contains a Harness's descendant processes on Windows](https://github.com/secantdev/secant/issues/259). + +**Recording.** Each Steer is durable against its Turn: the human's text, when it was sent, and how it settled, as delivered within the Turn, delivered +after a native boundary, re-delivered, or dropped by an Interrupt or a lost Turn. An interrupted Turn records the stop used, the control request or +the process stop. A follow-up Turn after an Interrupt is a human-origin Turn inside the Agent step's Attempt. Steered text is shown as the human's +message inside its Turn, durable content kept distinct from live previews under [ADR +0024](./0024-use-one-deep-projection-port-for-tui-and-headless-clients.md). + +**Harness Interface.** `steer` keeps its `ControlReceipt`, whose acceptance now means the text was handed to the Harness. The Turn's closed event +stream gains one Steer lifecycle, delivered or dropped, beside the Harness Request lifecycle, and on terminal the Adapter publishes a dropped event +for any Steer never delivered before it closes the producer and settles the one result. Which Steers are pending, the native correlation ids, and +where the real Turn boundary falls stay private to each Adapter; execution still sees one result per Turn and learns nothing native. The profile's +Steer evidence states each Harness's delivery point, and Claude Code declares `active-turn` interruption when qualification confirms the control +request. + +Rejected: keeping Interrupt as a halt, which turns stopping one's own conversation into a Run-level stop outside the workflow's logic; resending an +interrupted Agent step's prompt on resume; an interrupted Agent step that waits for the human to end it, which silently turns it into an Interactive +agent step; a client-side queue sent as the next Turn, which is emulation and would race the one live Turn per Prepared Harness and the Agent-declared +completion's clean-end settlement; a separate queue key beside Steer; counting Claude Code's second native exchange as a new Turn the Harness started, +which would let an Agent-declared completion settle the Step while the human's message is still being answered; reporting a Codex leftover upward as +missed, which pushes a native race into two Step kinds; SIGINT, which exits the process and silently loses queued messages; and `queued_turn_count`, +which stayed 0 while a message waited. How the Run Workbench presents Interrupt, Steer, a dropped Steer's restored draft, and delivered and dropped +Steers is left to the Run Workbench prototype. The decision was made on [Decide how the human interrupts a Turn and sends a message while a Turn is +working](https://github.com/secantdev/secant/issues/255). diff --git a/docs/glossary/secant-run-lifecycle.md b/docs/glossary/secant-run-lifecycle.md index 67c259ac..60b73ab2 100644 --- a/docs/glossary/secant-run-lifecycle.md +++ b/docs/glossary/secant-run-lifecycle.md @@ -14,7 +14,8 @@ This cluster defines the target Secant terms for a **Run** and everything that h a Repeat group is a Step but not a Stage; the group is the Stage. **End Stage** and an **Agent-declared completion**'s stage done complete a human-controlled group's Stage. - **Agent step** — a **Step kind** running one autonomous **Turn** in a named **Harness Session**. It completes without the human, though the human - may **Steer** it while that Turn is live when the selected Harness supports native Steer. + may **Steer** it while a Turn is live when the selected Harness supports Steer. After an **Interrupt** the human's message starts a follow-up Turn + in the same Session and **Step Attempt**, and the Attempt takes its outcome from its last Turn. - **Command step** — the deterministic non-agent **Step kind**. Its attempt succeeds if the command ran to an exit; the exit status becomes a **Verdict** and the captured output a `text` **Run Artifact**. The attempt fails only when the command could not execute. - **Verdict** — a **Run Artifact** type holding `pass` or `fail`, produced from a deterministic **Step**'s exit status. The only thing a **Repeat @@ -91,8 +92,9 @@ This cluster defines the target Secant terms for a **Run** and everything that h - **Harness Session** — a named conversation with the selected **Harness**, owned by exactly one **Run** and never shared across Runs. The routing names the session each agent **Step** runs in; Secant opens it on first use and reuses it after. - **Turn** — one mechanical user-to-**Harness** exchange inside a **Harness Session**: submitted input, model and tool activity, streamed progress, - and the Harness's authoritative turn boundary. It is neither a Session nor a judgement that the **Step** reached its goal. An **Agent step** has - one Turn per attempt; an **Interactive agent step** may have many. + and the Harness's authoritative turn boundary, the first one after which no **Steer** is still pending, so one Turn may span several native + exchanges. It is neither a Session nor a judgement that the **Step** reached its goal. An **Agent step** has one autonomous Turn per attempt, plus + any follow-up Turns the human starts after an **Interrupt**; an **Interactive agent step** may have many. - **Session availability** — whether a **Harness Session** is `open` (a next Turn can be sent now), `detached` (not live, but holding a native recovery coordinate worth reattaching), or `unusable` (native evidence authoritatively says recovery cannot continue). - **Model choice** — the **Run**'s one current model and effort level, each a real value the selected **Harness** names, never "Harness default". It @@ -117,11 +119,14 @@ This cluster defines the target Secant terms for a **Run** and everything that h recognised phrase in a **Turn**. The legacy grill is this shape. A step may opt into an **Entry Turn**. Inside a **Repeat group** each iteration is its own **Step Attempt** with its own **Harness Session**; ending the step advances only that iteration. - **Entry Turn** — an **Interactive agent step**'s optional first **Turn**: its Bundle-authored prompt, rendered with **Launch inputs** and bundled - skill paths, sent once on entry so the human need not retype what the launch already carries. It is never re-sent: after a halt the human - continues the same **Harness Session**. -- **Steer** — sending native same-Turn guidance while a **Turn** is live. It is not a new Turn, and unsupported Harnesses do not emulate it. -- **Interrupt** — asking the **Harness** to stop the current live **Turn**, including its native tool work. Confirmation ends the **Step Attempt** - `cancelled` and leaves the **Run** `halted` and re-attemptable; it does not close the Harness or cancel the Run. + skill paths, sent once on entry so the human need not retype what the launch already carries. It is never re-sent: after an **Interrupt** or a halt the + human continues the same **Harness Session**. +- **Steer** — a message the human sends while a **Turn** is live, delivered natively at the **Harness**'s next boundary and inside that Turn, + which lasts until every Steer is delivered. Delivered means the Harness put it in front of the model, not that the model followed it. An + **Interrupt** drops a Steer not yet delivered and returns its text to the human. _Avoid_: queued message. +- **Interrupt** — asking the **Harness** to stop the current live **Turn** and its foreground tool work. It ends only the Turn: the **Step Attempt** + stays open and the **Run** waits `blocked` on the human, whose next message continues the same **Harness Session**. It does not close the + Harness, halt the Run, or cancel it. - **Cancel** — explicitly ending a **Run**. The only route to the terminal `cancelled` state. - **Preflight** — the precondition check performed before a **Run** exists: the **Composition check**, presence of required **Launch inputs**, the union of authored **Workspace prerequisites**, intrinsic **Step kind** preconditions and Harness capability needs, and resolution of each selected @@ -137,14 +142,14 @@ An Agent-bearing **Run** pins its semantic **Harness** selection with its **Work routing truth, while the Model choice may change between and during Turns; each Agent-step Attempt separately records the Harness executable, version, and Adapter evidence, and each Turn its requested and effective model and effort. -| State | Meaning | Terminal | -| ----------- | --------------------------------------------------------------- | -------- | -| `running` | a **Step Attempt** is executing | no | -| `blocked` | a **Human Gate** or **Harness Request** is waiting on the human | no | -| `halted` | stopped for a reason outside the workflow's logic | no | -| `failed` | the workflow concluded negatively | no | -| `succeeded` | the routing completed | yes | -| `cancelled` | the user explicitly ended the Run | yes | +| State | Meaning | Terminal | +| ----------- | ----------------------------------------------------------------------------- | -------- | +| `running` | a **Step Attempt** is executing | no | +| `blocked` | a **Human Gate**, **Harness Request**, or the human's next message is waiting | no | +| `halted` | stopped for a reason outside the workflow's logic | no | +| `failed` | the workflow concluded negatively | no | +| `succeeded` | the routing completed | yes | +| `cancelled` | the user explicitly ended the Run | yes | `blocked` is durable truth, not computed: since [#108](https://github.com/secantdev/secant/issues/108) the authored **Human Gate**'s `pending_gate` row and the `blocked` state are written in one transaction, and execution also stores `blocked` before a checkpoint pause, so a killed Run reconciles @@ -191,6 +196,8 @@ row and the `blocked` state are written in one transaction, and execution also s to ADR 0020's reasons, and the human-controlled group's **Review checkpoint**. - [ADR 0033](../adr/0033-carry-agent-calls-to-secant-over-a-per-session-loopback-mcp-server.md) owns the **Agent call**'s channel, its attribution to a **Harness Session**'s live **Turn**, and its reply. +- [ADR 0035](../adr/0035-interrupt-ends-only-the-turn-and-a-mid-turn-message-is-a-native-steer.md) owns what an **Interrupt** leaves behind, + **Steer** delivery, and why a Turn lasts until every Steer is delivered. - [ADR 0023](../adr/0023-own-durable-run-truth-in-isolated-run-stores.md) owns durable Run truth, Artifact publication, Workspace materialization, retention, and recovery storage. - [ADR 0031](../adr/0031-own-runs-per-run-not-per-workspace.md) owns Run ownership: many live Runs per Workspace, one owner per Run, and what a diff --git a/docs/headless-parity.md b/docs/headless-parity.md new file mode 100644 index 00000000..351fd065 --- /dev/null +++ b/docs/headless-parity.md @@ -0,0 +1,15 @@ +# Headless Parity + +What the Run Workbench (the TUI) offers that the headless `secant` commands do not. Headless is for scripts and CI, so it drives a Run to a resting +state without a human at the keyboard; anything that needs the human inside a live Turn is TUI-only. A gap listed here is deliberate unless its +decision says otherwise; closing one needs its own decision. + +| TUI capability | Headless | Decided in | +| --------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------ | ----------------------------------------------------------------------------------------------- | +| Interactive agent steps | `run launch` refuses a Bundle whose Routing reaches one ("Interactive agent Steps are TUI-only") | [#122](https://github.com/secantdev/secant/issues/122) | +| **Interrupt** a live Turn, leaving the Run waiting on the human | none; Ctrl+C is a signal that halts the Run (ADR 0019) | [ADR 0035](./adr/0035-interrupt-ends-only-the-turn-and-a-mid-turn-message-is-a-native-steer.md) | +| The follow-up message that continues an Agent step after an Interrupt | none, since headless cannot Interrupt | [ADR 0035](./adr/0035-interrupt-ends-only-the-turn-and-a-mid-turn-message-is-a-native-steer.md) | +| **Steer**: a message sent while a Turn is working | none | [ADR 0035](./adr/0035-interrupt-ends-only-the-turn-and-a-mid-turn-message-is-a-native-steer.md) | + +Wider headless parity for the agent screen, including Turn history that grows during a Turn and its frozen `--json` shapes, is still open on +[Chart the fixes Secant needs before its public release](https://github.com/secantdev/secant/issues/235). diff --git a/docs/research/claude-code-live-interrupt-and-mid-turn-send.md b/docs/research/claude-code-live-interrupt-and-mid-turn-send.md new file mode 100644 index 00000000..95589dbb --- /dev/null +++ b/docs/research/claude-code-live-interrupt-and-mid-turn-send.md @@ -0,0 +1,1088 @@ +# Claude Code Live Interrupt and Mid-Turn Send + +Research date: 2026-09-29 + +Harness version examined: Claude Code **2.1.284** (installed native executable, Linux +x64; `claude --version` returned `2.1.284 (Claude Code)`). + +Ticket: [#255](https://github.com/secantdev/secant/issues/255). This note settles +the Claude Code items left **Untested** or **Unknown** in +[Harness Interrupt and Queued Messages](harness-interrupt-and-queued-messages.md) +by running small live model sessions against the installed CLI. A second round, +[Round 2](#round-2-permission-waits-and-multi-tool-rounds), repeats the key stops and +sends through a stand-in for Secant's MCP permission bridge. A third round, +[Round 3](#round-3-sigterm-re-approval-and-denied-tool-calls), repeats the SIGTERM +during a permission wait and sends a user frame while the bridge denies a call. + +## Answer + +**SIGTERM loses streamed text but keeps a killed tool call.** A process-group SIGTERM +exits with code 143 and writes no `result`. Mid-text, the transcript keeps only the +user prompt: no partial text, no thinking, no interrupted marker. Mid-tool, the +transcript keeps the `tool_use` and a real `tool_result` of `Exit code 137` plus +the tool's partial stdout. On `--resume`, Claude Code first appends a synthetic +assistant message, `No response requested.` (`model: ""`). The resumed +model saw the killed tool call and its output. After a mid-text SIGTERM it saw +only its own prompt and the synthetic reply. Twice it said it had "declined" the +task. + +**SIGINT ends the Turn, then the process exits.** SIGINT gave one `result` +(`subtype: "error_during_execution"`, `is_error: true`, `terminal_reason` +`aborted_streaming` or `aborted_tools`). The process then exited with code 0 about +1 s later, with stdin still open. This held for the process group and for the pid +alone. The partial Turn is kept in the transcript: partial text is saved with +`isAbortedMidStream: true`, and `[Request interrupted by user]` is appended. A +`--resume` process sees both, after the synthetic `No response requested.`. A user +message still queued at the SIGINT was dropped and did not come back on resume. + +**A raw `control_request` `interrupt` works without `initialize` and keeps the +process alive.** The `control_response` came back in 2 to 4 ms and carried the +receipt `{"still_queued":[…]}`. It came before the Turn's `result`, which had the +same shape as the SIGINT one. The next stdin user frame then ran as a new Turn in +the same process and Session. + +- Mid-text, the partial text stays in context. The `assistant` frame carries + `aborted: true`, and the model later quoted the last line it had written. +- Mid-tool, the `tool_use` stays. Its `tool_result` is replaced by the standard + "The user doesn't want to proceed with this tool use… rejected" text, followed by + `[Request interrupted by user for tool use]`. Output the tool had already printed + is thrown away. The model then said the tool was "rejected before execution". +- The foreground Bash process tree was killed. +- A Bash command the model had moved to the background kept running after the + interrupt. It was killed only when the process exited. + +**A user frame written during a tool call is picked up in the same Turn. One written +while text streams becomes the next Turn.** + +- **During a tool call.** The frame was held until the tool finished. It then + reached the model as a `queued_command` system-reminder next to the tool result. + The Turn's one `result` listed its `uuid` in `user_message_uuids`, and + `num_turns` stayed 2. The model acted on it in that Turn. +- **Two frames** written 0.5 s apart were taken together at the same boundary, in + write order. +- **While text streamed.** The frame was not taken into the running Turn. It ran + as a second Turn with its own `result` and `result_index: 1`. +- `queued_turn_count` was `0` on every `result`, including the first `result` in + that last case, while the message was still waiting. + +**On an interrupt, a queued message runs, is cancelled, or is lost, depending on the +stop.** + +- **Plain `interrupt`.** The message is listed in `still_queued` and runs by itself + as the next Turn right after the interrupted `result`. +- **`interrupt` with `cancel_queued: true`.** The message is listed in `cancelled`, + and a `command_lifecycle` `cancelled` frame is written for it. No next Turn + starts, and the process stays alive. +- **SIGINT.** The message is dropped with no lifecycle frame, and the process + exits. + +**Version floor.** The changelog names none of `user_message_uuids`, +`queued_turn_count`, the `interrupt` control request, receipts, or `cancel_queued`. +The nearest entries are listed under [Version floor](#test-6-version-floor). + +**Which stop keeps the process alive and the partial Turn in context:** only the +`control_request` `interrupt`. SIGINT keeps the partial Turn but ends the process. +SIGTERM ends the process and loses streamed text. + +**Round 2: the same holds while a tool call waits on the permission bridge.** See +[Round 2](#round-2-permission-waits-and-multi-tool-rounds). + +- **Raw `interrupt` during a permission wait.** It was honoured in three of three + runs, with a `control_response` in 2 to 4 ms. Claude Code sent the bridge an MCP + `notifications/cancelled` for the pending `approve` call within 4 ms. The Turn + ended with the same `error_during_execution` / `aborted_tools` `result` as a + mid-tool interrupt, and `permission_denials` named the waiting call. The process + stayed alive and the next user frame ran in the same Session. The model saw the + call as "rejected before it could execute". The bridge's late `allow`, 30 s + later, had no effect: the command never ran. +- **SIGTERM during a permission wait.** The process exited 143 with no `result` + and no `tool_result`. No `notifications/cancelled` was sent. During shutdown, + Claude Code opened a new MCP session to the bridge and sent a second `approve` + call for the same `tool_use_id`. On `--resume`, the transcript gained a synthetic + `tool_result`: "[Tool call interrupted: the session ended before this call's + result was recorded, so its outcome is unknown…]". Round 3 repeated this; see + below. +- **A user frame is still taken at a tool boundary, but only after the whole tool + round.** This held when two MCP calls ran at once, when a Bash call and an MCP + call ran one after the other in one round, and for a single MCP call. It also + held for a frame written while a Bash call waited on the bridge: the frame was + held through the approval and the command. Each time, the frame reached the model + in the same Turn, and its `uuid` was in that `result`'s `user_message_uuids`. + +**Round 3: SIGTERM always re-asks, and a denial is a tool boundary.** See +[Round 3](#round-3-sigterm-re-approval-and-denied-tool-calls). + +- **The second `approve` after SIGTERM came every time.** It came in six of six + runs, with the SIGTERM 0.5 s, 3 s, or 10 s into the wait, 29 to 47 ms after the + signal, on a new MCP session, with the same `tool_use_id`. With R2-A2 that is + seven of seven. +- **Answering it did not run the command.** A quick `allow` (two runs) or `deny` + (one run) was delivered in full, but the command's side-effect file never + appeared, no frame was written, and the process still exited 143 about 0.9 s + after the signal. On `--resume` the call got the same synthetic "outcome is + unknown" result whatever the answer. +- **A user frame written during a wait that ends in a denial is taken with the + denial, in the same Turn, and listed.** Two of two runs: the frame reached the + model next to the denial's `tool_result`, and its `uuid` was in the Turn's one + `result`, which also named the call in `permission_denials`. +- **A frame written as the denial lands can miss the Turn.** Written 8 ms before + or 33 ms after the denial's `tool_result` frame, it ran as the next Turn, in + three of three runs. +- **A denial does not end a round early.** With one call denied and one allowed in + the same round, the frame was taken only after the allowed call's result, in the + same Turn (one run). + +Windows was not tested. + +## Evidence Vocabulary + +- **Recorded (2.1.284)**: seen in the live frame sequences or session transcripts + captured for this note on 2026-09-29 against Claude Code 2.1.284. Every fact + below without another label is Recorded (2.1.284). +- **Documented**: stated in the Claude Code changelog or official docs, cited + directly or through + [Harness Interrupt and Queued Messages](harness-interrupt-and-queued-messages.md). +- **Inferred**: a consequence drawn from Recorded facts; it needs a check before it + becomes a compatibility promise. +- **Unknown**: not settled by these runs. + +## Method + +A small Bun driver spawned `claude` with the stream flags that +`src/harness/claude-code.ts` uses. It spawned `claude` `detached`, so the CLI led its +own process group, as `src/process/process.ts` does: + +```text +claude -p --input-format stream-json --output-format stream-json --verbose \ + --include-partial-messages --model haiku \ + [--session-id | --resume ] \ + --allowedTools "Bash(sleep:*)" "Bash(echo:*)" "Bash(./slowjob.sh)" +``` + +The driver worked like this: + +- It wrote user frames in the Adapter's `encodeTurn` shape plus a `uuid`: + `{"type":"user","message":{"role":"user","content":""},"parent_tool_use_id":null,"uuid":""}`. +- It logged every stdout line with a millisecond offset, and recorded exit code + and signal. +- It signalled with `process.kill(-pid, sig)` for the process group, as Secant's + `killGroup` does, or with `process.kill(pid, sig)` for the pid alone. +- It checked the Bash child with `ps -eo pid,ppid,pgid,sid,args` before and + after each stop. +- It read the session transcript at + `~/.claude/projects//.jsonl` after each run. +- Each process ran with a scratch directory as cwd, not the repository. It had + no `CLAUDE*` environment variables, used subscription auth + (`apiKeySource: "none"`), and ran in permission mode `default`. + +The two Turn shapes: + +- **Mid-tool.** The prompt asked for `./slowjob.sh` in the foreground, then the + code word `PINEAPPLE`. `slowjob.sh` is `echo started-27; sleep 27; echo done-27`. + The stop or send came 3 s after the `tool_use` frame. The brief's `sleep 20` could + not be used: 2.1.284's Bash tool refuses long sleeps on its own. It answered + `sleep 27` with `Blocked: standalone sleep 27…` and `sleep 27 && echo done-27` + with `Blocked: sleep 27 followed by: echo done-27…`. The first `sleep 27` run + (test 3, run 1) was moved to the background by the model. +- **Mid-text.** The prompt was "count from one to three hundred in English words, + one number per line". The stop or send came 2.5 s after the first text delta, + about 1,400 characters in. + +After each stop the driver asked the same process, or a `--resume` process, +"Without running any tools: what was the last thing you wrote or did in this +conversation before this message? Quote the last line you wrote…". + +Differences from Secant's launch: + +- Secant adds the loopback MCP permission bridge (`--mcp-config`, + `--permission-prompt-tool`). The round 1 driver did not. Round 2 added a stand-in + bridge; see [Round 2](#round-2-permission-waits-and-multi-tool-rounds). +- The host's user settings load no-op `SessionStart`, `PreToolUse`, `PostToolUse`, + and `Stop` hooks, plus a user `CLAUDE.md`. + +Frame excerpts below are trimmed. `t` is milliseconds since spawn. Hook, status, +`thinking_tokens`, and `stream_event` frames are omitted, and so are most +`command_lifecycle` frames. Raw logs were kept outside the repository. + +## Test 1: SIGTERM mid-Turn, then `--resume` + +### 1a: during a tool call (three runs) + +```text +t=2916 assistant tool_use {"command":"./slowjob.sh","run_in_background":false} + ps: 68103 claude (pgid 68103) + 68275 /bin/bash -c … eval ./slowjob.sh (pgid 68275, sid 68275) + 68277 /bin/bash ./slowjob.sh (pgid 68277) + 68278 sleep 27 (pgid 68277) +t=5986 SIGNAL SIGTERM group pid=68103 +t=6000 system/task_notification status=stopped +t=6021 user tool_result "Exit code 137\nstarted-27" is_error +t=6865 EXIT code=143 signal=null +t=8427 ps: no slowjob or sleep 27 left +``` + +No `result` frame was written. The transcript kept the whole tool round, and resume +added one synthetic entry: + +```text +user "Use the Bash tool … ./slowjob.sh …" +assistant thinking, tool_use toolu_01… stop_reason=tool_use +user tool_result "Exit code 137\nstarted-27" is_error + toolUseResult="Error: Exit code 137\nstarted-27" +--- written by the --resume process --- +assistant "No response requested." model= stop_reason=stop_sequence +user +``` + +The resumed model answered: "I ran the Bash command `./slowjob.sh` in the +foreground, but it was killed before completion—the exit code was 137 (SIGKILL) and +it only printed "started-27" before terminating." + +- **The Bash tree was killed, although it was outside the signalled group.** The + Bash tool runs its shell as a new session and process group (`pgid 68275`), so + `kill(-pid)` does not reach it. Claude Code killed it itself (exit 137). This + matches the documented 2.1.212 fix for orphaned trees on SIGTERM. +- The killed tool call is kept as an ordinary error result. There is no + `[Request interrupted…]` marker and no synthetic denial. This matches 2.1.236. + +### 1b: while text streams (two runs) + +```text +t=5139 partial text so far (1372 chars) tail="…One Hundred Eighteen\nOne" +t=5140 SIGNAL SIGTERM group +t=6401 EXIT code=143 signal=null +``` + +No `result` frame and no `assistant` frame were written. The transcript held only +the user prompt, then the resume's synthetic `No response requested.`. The resumed +model answered, in both runs: "The last line I wrote was: "No response requested." +No commands ran—you asked me to count without using tools, which I declined to do." + +- **The partial text and its thinking are gone.** The resumed model has no trace + that it started to answer. +- The synthetic `No response requested.` reads, to the model, as a refusal of the + prompt. **Inferred.** + +## Test 2: SIGINT to a `-p` stream-json process + +### 2a: during a tool call (process group, then pid only) + +```text +t=3282 assistant tool_use {"command":"./slowjob.sh",…} +t=6325 SIGNAL SIGINT group +t=6382 user tool_result "The user doesn't want to proceed with this tool use. The tool + use was rejected (eg. if it was a file edit, the new_string was NOT written to + the file). STOP what you are doing and wait for the user t…" is_error +t=6385 user text "[Request interrupted by user for tool use]" +t=6401 result subtype=error_during_execution is_error=true terminal_reason=aborted_tools + stop_reason=tool_use num_turns=3 user_message_uuids=[] + errors=["[ede_diagnostic] result_type=user last_content_type=n/a stop_reason=tool_use"] +t=7255 EXIT code=0 signal=null +``` + +SIGINT to the pid alone (2c) gave the same frames and `EXIT code=0` after 1 s. + +### 2b: while text streams (two runs) + +```text +t=5239 SIGNAL SIGINT group +t=5277 assistant text "one\ntwo\n…" (1822 chars) aborted=true +t=5278 user text "[Request interrupted by user]" +t=5283 result subtype=error_during_execution terminal_reason=aborted_streaming + stop_reason=null num_turns=2 total_cost_usd=0 duration_api_ms=0 +t=6564 EXIT code=0 signal=null +``` + +Transcript, then the `--resume` process: + +```text +assistant thinking stop_reason=null +assistant text "one\ntwo\n…" (1822 chars) isAbortedMidStream=true +user "[Request interrupted by user]" +--- written by the --resume process --- +assistant "No response requested." model= +user +``` + +The resumed model answered: "The last line I wrote was "one hundred" as part of +counting from one to three hundred in words. You interrupted the counting task +mid-way through the one-hundreds range." + +- **SIGINT ends the Turn and then the process.** The process did not wait for more + stdin. It exited 0 about 1 s after the `result` in all four runs, so no next + user frame can run in the same process. +- **The partial Turn is kept.** On resume, the model sees the partial text, the + interrupt marker, and the synthetic reply. For a tool call, it sees the rejected + result. The second mid-text run's resumed model quoted the synthetic line as its + last line but still said "you interrupted the counting task". +- The Bash tree was gone after the SIGINT. +- The interrupted mid-text `result` reported `total_cost_usd: 0` and zero usage, + even though about 1,400 characters had been generated. The next `result` in the + process (test 3b) carried the running total again. Whether an aborted stream's + tokens are ever counted is Unknown. + +## Test 3: raw `control_request` `interrupt` without `initialize` + +The first frame of each process was the user prompt. No `initialize` was ever sent. + +### 3a: during a tool call + +```text +t=2732 assistant tool_use {"command":"./slowjob.sh",…} +t=5774 in {"type":"control_request","request_id":"int1","request":{"subtype":"interrupt"}} +t=5778 control_response {"subtype":"success","request_id":"int1","response":{"still_queued":[]}} +t=5792 user tool_result "The user doesn't want to proceed with this tool use. The tool use was + rejected …" is_error +t=5793 user text "[Request interrupted by user for tool use]" +t=5800 result subtype=error_during_execution is_error=true terminal_reason=aborted_tools + num_turns=3 user_message_uuids=[] queued_turn_count=0 result_index=0 +t=5800 command_lifecycle state=cancelled +t=7343 ps: no slowjob or sleep 27 left; process alive +t=7343 in user +t=7403 system/init (same session_id) +t=11845 assistant text "I attempted to call the Bash tool to run `./slowjob.sh`, but the + tool use was rejected before execution. No command finished and I received no + output—the user interrupted the tool call." +t=11891 result subtype=success num_turns=1 result_index=1 +``` + +In the transcript, the `tool_result` is stored with `toolUseResult: "User rejected +tool use"` and `toolDenialKind: "user-rejected"`. In tests 5a and 5b, `slowjob.sh` +had already printed `started-27` when the interrupt came, and the stored +`tool_result` was the same rejection text. The partial stdout was discarded. No +synthetic `No response requested.` is written when the next Turn runs in the same +process. + +Run 1 used a bare `sleep 27` and hit the Bash tool's sleep guard. The model then +started `sleep 27` with `run_in_background: true`. The interrupt at t=5612 ended the +Turn, but `sleep 27` was still running at t=7154. The background task was reported +`killed` only at t=15993, after the driver closed stdin and the process began to +exit. The second Turn's model said the background command "is still running". + +### 3b: while text streams + +```text +t=5444 in control_request interrupt int1 +t=5446 control_response {"subtype":"success","request_id":"int1","response":{"still_queued":[]}} +t=5481 assistant text "One\nTwo\n…" (1412 chars) aborted=true +t=5484 user text "[Request interrupted by user]" +t=5488 result subtype=error_during_execution terminal_reason=aborted_streaming stop_reason=null + total_cost_usd=0 num_turns=2 result_index=0 +t=7027 in user +t=10714 assistant text "The last line I wrote was "One" (after "One Hundred Twenty"), + completing the number list you requested. … I was interrupted partway through + the three-hundred count." +t=10775 result subtype=success num_turns=1 result_index=1 +t=11673 EXIT code=0 (after the driver closed stdin) +``` + +- **Honoured without `initialize`.** The `control_response` is a `success` with the + documented receipt, written before the interrupted `result`, as the SDK typings + describe. +- **The process stays alive and the Session continues.** The next user frame ran + in the same process with the same `session_id`. +- **Partial text is in context. A partial tool result is not.** The model quoted + the exact last partial line. For the tool call, it saw only the rejection text, + and it believed the command never ran. +- **The foreground Bash tree is killed. A background Bash task is not.** +- The `init` frame advertised `capabilities: ["interrupt_receipt_v1", +"interrupt_cancel_queued_v1", "msg_lifecycle_v1", "mcp_read_resource_v1", +"mcp_tool_ui_meta_v1"]`. `command_lifecycle` frames (`queued`, `started`, + `completed`, `cancelled`) were written by default on this raw stream, with no + opt-in. + +## Test 4: user frames written mid-Turn + +### 4a: one frame during a tool call + +```text +t=2978 assistant tool_use {"command":"./slowjob.sh",…} +t=6018 in user uuid=aaaaaaaa-…-000000000001 "Additional instruction: after the command + finishes, also say the word MANGO." +t=6022 command_lifecycle aaaaaaaa-… state=queued +t=30101 user tool_result "started-27\ndone-27" +t=30139 command_lifecycle aaaaaaaa-… state=started +t=34711 assistant text "PINEAPPLE\nMANGO" +t=34741 command_lifecycle aaaaaaaa-… state=completed +t=34746 result subtype=success num_turns=2 queued_turn_count=0 result_index=0 + user_message_uuid= + user_message_uuids=[, "aaaaaaaa-0000-4000-8000-000000000001"] +t=34748 command_lifecycle state=completed +``` + +The transcript shows how the frame reached the model: + +```text +queue-operation enqueue "Additional instruction: …" +user tool_result "started-27\ndone-27" +attachment queued_command source_uuid=aaaaaaaa-… commandMode=prompt + rendered: "\nThe user sent a new message while you were + working:\nAdditional instruction: after the command finishes, also say + the word MANGO.\n\nThis is how Claude Code surfaces messages the user + sends mid-turn — within the running turn, often alongside the next tool + result, rather than as a separate conversation turn. Address the message + above as you continue this turn.\n" +queue-operation remove reason=absorbed_mid_turn +assistant "PINEAPPLE\nMANGO" +``` + +- **Picked up in the same Turn.** The frame is held until the running tool + finishes, not injected into it, and it did not cut the 27 s command short. It + reached the model with the tool result. +- The Turn's one `result` lists the picked-up `uuid` in `user_message_uuids`. + `user_message_uuid` stays the prompt's. `num_turns` is 2, the same as a Turn with + no pickup. +- The picked-up frame is never echoed on stdout as a `user` frame; the driver did + not pass `--replay-user-messages`. `command_lifecycle` `completed` for it comes + before the Turn's `result`. + +### 4b: one frame while text streams, with no tool round left + +```text +t=5584 in user uuid=bbbbbbbb-…-000000000001 "Now reply with the single word MANGO." +t=5585 command_lifecycle bbbbbbbb-… state=queued +t=10247 assistant text "One\nTwo\n…" (5488 chars) +t=10290 result subtype=success num_turns=1 queued_turn_count=0 result_index=0 + user_message_uuids=[] +t=10293 command_lifecycle state=completed +t=10294 command_lifecycle bbbbbbbb-… state=started +t=10333 system/init +t=11500 assistant text "MANGO" +t=11542 result subtype=success num_turns=1 queued_turn_count=0 result_index=1 + user_message_uuid="bbbbbbbb-…" user_message_uuids=["bbbbbbbb-…"] +``` + +- **Not picked up. It runs as the next Turn.** The first Turn finished its text + with no sign of the message. A second `system/init` and a second `result` followed + at once, with no new stdin write. +- **`queued_turn_count` was `0`** on the first `result`, while `bbbbbbbb-…` was + still queued and about to run. It was `0` on every `result` captured for this + note. It cannot be used to tell that another Turn will follow. The + `command_lifecycle` `queued` frame, without a matching `started` before the + `result`, did show this. + +### 4c: two frames during a tool call + +```text +t=5767 in user uuid=cccccccc-…-01 "Additional instruction one: … say the word MANGO." +t=6269 in user uuid=cccccccc-…-02 "Additional instruction two: after that, say the word KIWI." +t=29830 user tool_result "started-27\ndone-27" +t=29890 command_lifecycle cccccccc-…-01 state=started +t=29890 command_lifecycle cccccccc-…-02 state=started +t=33241 assistant text "PINEAPPLE" +t=33277 result subtype=success num_turns=2 queued_turn_count=0 + user_message_uuids=[, "cccccccc-…-01", "cccccccc-…-02"] +``` + +- **Both frames are taken at the same boundary, in write order.** The transcript + has two `queued_command` attachments (one, then two) after the tool result. The + Turn made one more model call, not one per message. +- **Listed does not mean acted on.** Both `uuid`s are in `user_message_uuids`, but + the model answered only `PINEAPPLE`. The original prompt said "and nothing else", + and Haiku kept to it. `user_message_uuids` shows delivery to the model, not + compliance. + +## Test 5: a queued message, then a stop + +Each run wrote a user frame during the tool call and stopped the Turn 1.5 s later, +before the tool finished. + +### 5a: plain `interrupt` + +```text +t=5955 in user uuid=dddddddd-…-01 "Queued message: reply with the single word MANGO." +t=7456 in control_request interrupt int1 +t=7460 control_response {"subtype":"success","request_id":"int1", + "response":{"still_queued":["dddddddd-0000-4000-8000-000000000001"]}} +t=7472 user tool_result "The user doesn't want to proceed with this tool use. …" is_error +t=7480 result subtype=error_during_execution terminal_reason=aborted_tools result_index=0 + user_message_uuids=[] +t=7484 command_lifecycle dddddddd-… state=started +t=9103 assistant text "MANGO" +t=9136 result subtype=success num_turns=1 result_index=1 user_message_uuids=["dddddddd-…"] +``` + +- **The queued message runs by itself as the next Turn**, 4 ms after the + interrupted `result`, with no new stdin write. The receipt names it in + `still_queued`. + +### 5b: `interrupt` with `cancel_queued: true` + +```text +t=5728 in user uuid=eeeeeeee-…-01 "Queued message: reply with the single word MANGO." +t=7229 in {"type":"control_request","request_id":"int1", + "request":{"subtype":"interrupt","cancel_queued":true}} +t=7234 command_lifecycle eeeeeeee-… state=cancelled +t=7235 control_response {"subtype":"success","request_id":"int1", + "response":{"still_queued":[],"cancelled":["eeeeeeee-0000-4000-8000-000000000001"]}} +t=7262 result subtype=error_during_execution terminal_reason=aborted_tools result_index=0 +t=18827 process alive, no further frames; in user +t=23444 result subtype=success result_index=1 +``` + +- **`cancel_queued` works on the raw stream.** The message is dropped, and the + receipt names it in `cancelled`. It never reaches the model: the transcript has + `queue-operation` `enqueue` then `remove`, and no `queued_command` attachment. No + Turn starts until the next stdin frame. + +### 5c: SIGINT + +```text +t=6501 in user uuid=ffffffff-…-01 "Queued message: reply with the single word MANGO." +t=6502 command_lifecycle ffffffff-… state=queued +t=8001 SIGNAL SIGINT group +t=8050 result subtype=error_during_execution terminal_reason=aborted_tools +t=8052 command_lifecycle state=cancelled +t=8961 EXIT code=0 signal=null +``` + +- **The queued message is lost.** No lifecycle frame is written for + `ffffffff-…`, and the process exits. The transcript holds a `queue-operation` + `enqueue` for it with no `dequeue` or `remove`. A `--resume` of the Session did + not replay it. Asked to list every user message, the resumed model listed the + prompt, `[Request interrupted by user for tool use]`, the synthetic `No response +requested.`, and the question. MANGO was not among them. + +## Test 6: version floor + +Read from the [Claude Code CHANGELOG](https://github.com/anthropics/claude-code/blob/main/CHANGELOG.md) +at its 2.1.284 head (**Documented**). No entry names `user_message_uuids`, +`user_message_uuid`, `queued_turn_count`, the `interrupt` control request, +`interrupt_receipt`, `still_queued`, `cancel_queued`, `command_lifecycle`, or a +`priority` field on user messages. The nearest entries: + +| Version | Entry (quoted or trimmed) | +| ------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| 2.1.19 | "[SDK] Added replay of `queued_command` attachment messages as `SDKUserMessageReplay` events when `replayUserMessages` is enabled" | +| 2.1.94 | "Fixed SDK/print mode not preserving the partial assistant response in conversation history when interrupted mid-stream" | +| 2.1.162 | "Fixed an interrupt (Esc) sent at the very start of a turn being silently dropped in stream-json/SDK sessions" | +| 2.1.212 | "Fixed SIGTERM during a running Bash tool orphaning the command's process tree in print/SDK mode; the CLI now aborts the turn, kills the tree, and exits 143" | +| 2.1.236 | "SIGTERM in print/SDK mode no longer records an interrupted turn or synthetic tool denials before exiting; running commands are still terminated and the process still exits with code 143" | +| 2.1.246 | "Fixed MCP tool calls interrupted by an incoming message in headless/remote sessions being reported to the model as "completed with no output" instead of an explicit interrupted error" | +| 2.1.261 | "Fixed SDK and cloud sessions ignoring a Stop or interrupt sent just after the first prompt, before the turn had started" | +| 2.1.274 | "Fixed a local `claude -p --resume` started with `CLAUDE_CODE_RESUME_INTERRUPTED_TURN` not reporting background tasks the previous process left unfinished" | + +The SDK reference's own version notes still apply: `interrupt_receipt_v1` from +v2.1.205 and `interrupt_cancel_queued_v1` from v2.1.219 (**Documented**, cited in +[Harness Interrupt and Queued Messages](harness-interrupt-and-queued-messages.md)). +The installed 2.1.284 advertises both, and also `msg_lifecycle_v1`, in +`system/init.capabilities`. + +## Test 7: Windows + +Not tested. No Windows host was available. Signal delivery, `taskkill` tree +behaviour, and the SIGINT and SIGTERM results above are all Linux-only. + +## Round 2: permission waits and multi-tool rounds + +Run on 2026-09-29 against the same Claude Code **2.1.284**. Every fact in this +section is Recorded (2.1.284) unless it carries another label. + +### Round 2 method + +The round 2 driver used the same launch flags as `src/harness/claude-code.ts`: +`-p --input-format stream-json --output-format stream-json --verbose +--include-partial-messages --model haiku --session-id ` (or `--resume `), +then the bridge fragment. It passed no `--allowedTools` and no permission-mode flag, +so `system/init` reported `permissionMode: "default"`. + +The stand-in bridge copied `src/harness/permission-bridge.ts`: + +- a Streamable HTTP MCP server on `127.0.0.1` with a random port and a bearer + token, one transport per MCP session, built on the repository's + `@modelcontextprotocol/sdk` 1.29.0; +- server `secant-permissions` with one `approve` tool (`tool_name`, `input`, + `tool_use_id`), returning `{"behavior":"allow","updatedInput":…}` or + `{"behavior":"deny",…}` as JSON text; +- launch fragment `--mcp-config --permission-prompt-tool +mcp__secant-permissions__approve`. + +Unlike Secant, the stand-in let the experiment set a delay for each `approve` +answer, and it did not stop when the call was cancelled. It logged every HTTP +request body, every closed response stream, and the handler's abort signal. For +round 2 it also served a second MCP server, `slowtools`, from the same process. Its +one tool, `slow_echo(text, delay_s)`, waits `delay_s` seconds and returns `echo: +`. It is marked `readOnlyHint: true`. + +Other differences from round 1: + +- `ORCA_*` variables were removed from the child environment, as well as + `CLAUDE*`. The host's user hooks, including a `PermissionRequest` hook, then + printed `{}` and did not decide anything. Every `Bash ./slowjob.sh` and + `./slow15.sh` call reached the stand-in `approve`. A Bash `echo` the model ran + on its own did not: Claude Code allowed it without a prompt. +- The user's other MCP servers (context7 and the claude.ai connectors) also loaded, + as they would under Secant. MCP tools were deferred, so the model sometimes + called `ToolSearch` before `slow_echo`. +- `slow15.sh` is `echo started-15; sleep 15; echo done-15`. + +`HTTP` and `MCP` lines below are the stand-in's log, on the same clock as the +stream frames. + +### R2-A1: raw `interrupt` while Bash waits on the bridge (two runs, plus one with `cancel_queued`) + +The stand-in held the Bash `approve` for 30 s, then answered `allow`. The +interrupt came 3 s into the wait. Run 1: + +```text +t=5832 assistant tool_use toolu_01Y7… Bash {"command":"./slowjob.sh",…} +t=5854 HTTP POST /mcp tools/call#2 name=approve +t=5857 MCP approve#1 CALLED tool=Bash tool_use_id=toolu_01Y7… -> will allow after 30000ms +t=8937 ps: no slowjob or sleep 27 +t=8938 in {"type":"control_request","request_id":"int1","request":{"subtype":"interrupt"}} +t=8940 control_response {"subtype":"success","request_id":"int1","response":{"still_queued":[]}} +t=8942 HTTP POST /mcp notifications/cancelled params={"requestId":2,"reason":"AbortError: remote-cancel"} +t=8944 MCP approve#1 extra.signal ABORTED +t=8946 user tool_result "The user doesn't want to proceed with this tool use. The tool use was + rejected …" is_error +t=8948 user text "[Request interrupted by user for tool use]" +t=8958 result subtype=error_during_execution is_error=true terminal_reason=aborted_tools + stop_reason=tool_use num_turns=3 result_index=0 queued_turn_count=0 + user_message_uuids=[] + permission_denials=[{"tool_name":"Bash","tool_use_id":"toolu_01Y7…",…}] +t=8960 command_lifecycle state=cancelled +t=10495 in user +t=17396 assistant text "I attempted to call the Bash tool to run `./slowjob.sh`, but the tool use + was rejected before it could execute. No command finished and no output was + produced—the tool call was denied by the user." +t=17424 result subtype=success num_turns=1 result_index=1 +t=35858 MCP approve#1 RETURNING {"behavior":"allow",…} (signal.aborted=true) +t=39891 ps: no slowjob or sleep 27; no further stream frames +t=42932 HTTP approve response stream closed (only when the driver closed stdin) +t=43767 EXIT code=0 +``` + +Run 2 (test C, the repeat) gave the same frames: `control_response` in 2 ms, +`notifications/cancelled` 4 ms after the interrupt, the same `result`, and the same +`permission_denials`. Asked afterwards, the model said "My only action was an +attempted Bash tool call to run `./slowjob.sh`, which was rejected before +execution." + +- **Honoured, and the process stays alive.** The next user frame ran as + `result_index: 1` in the same process and Session. +- **Claude Code cancels the bridge call.** The bridge receives a standard MCP + `notifications/cancelled` with the `approve` call's JSON-RPC id and `reason: +"AbortError: remote-cancel"`. It is sent on a new POST. The HTTP response stream + for the cancelled call is not closed. It stayed open until the process exited. +- **A late answer does nothing.** The stand-in's `allow` came 27 s after the + interrupt. The SDK server does not send a response for a cancelled request, so + nothing reached Claude Code. **Inferred** from the SDK, and consistent with what + was seen: `slowjob.sh` never started, and no frame followed. +- **The transcript matches a mid-tool interrupt.** The `tool_result` is stored with + `toolUseResult: "User rejected tool use"` and `toolDenialKind: "user-rejected"`, + followed by `[Request interrupted by user for tool use]`. Nothing marks that the + call was waiting on approval and never ran. Only `permission_denials` in the + `result` names it. + +With `cancel_queued: true` and a user frame written 1.5 s before the interrupt: + +```text +t=5571 MCP approve#1 CALLED tool=Bash … -> will allow after 30000ms +t=7110 in user uuid=a1a1a1a1-…-01 "Queued message: reply with the single word MANGO." +t=7111 command_lifecycle a1a1a1a1-… state=queued +t=8610 in control_request interrupt int1 cancel_queued=true +t=8613 command_lifecycle a1a1a1a1-… state=cancelled +t=8614 control_response {"still_queued":[],"cancelled":["a1a1a1a1-0000-4000-8000-000000000001"]} +t=8615 HTTP POST /mcp notifications/cancelled params={"requestId":2,…} +t=8631 result subtype=error_during_execution terminal_reason=aborted_tools result_index=0 +t=13667 in user (no Turn started before this) +t=18831 result subtype=success result_index=1 +``` + +The queued message was dropped as in test 5b, and the bridge call was cancelled the +same way. + +### R2-A2: SIGTERM while Bash waits on the bridge + +```text +t=4926 assistant tool_use toolu_01Ub… Bash {"command":"./slowjob.sh",…} +t=4975 MCP approve#1 CALLED tool=Bash tool_use_id=toolu_01Ub… reqId=2 (held 30 s) +t=8033 SIGNAL SIGTERM group +t=8037 HTTP response streams closed: both GET streams and the approve#1 stream +t=8055 HTTP POST /mcp server/discover, then initialize → new MCP session cea0a024 +t=8072 HTTP POST /mcp sid=cea0a024 tools/call#1 name=approve +t=8073 MCP approve#2 CALLED tool=Bash tool_use_id=toolu_01Ub… (same call, again) +t=8896 approve#2 response stream closed +t=8897 EXIT code=143 signal=null +``` + +- **The process exits 143 with no `result`**, 864 ms after the signal. No + `notifications/cancelled` was sent. The bridge saw its streams drop. +- **A second `approve` call arrives during shutdown.** Claude Code opened a new MCP + session and asked again for the same `tool_use_id`, 39 ms after the SIGTERM. It + exited without waiting for the answer. The bash command never ran. Seen in the + one run. +- **The original process writes nothing for the waiting call.** The transcript ends + at the `tool_use`, with no `tool_result`. +- **`--resume` fills the gap with a synthetic result.** The resume process + (launched with the bridge flags) appended, before its first Turn: + +```text +user tool_result "[Tool call interrupted: the session ended before this call's result + was recorded, so its outcome is unknown. Check whether it took effect before + relying on it or running it again.]" is_error toolDenialKind="interrupted" +assistant "No response requested." model= +user +``` + +The resumed model answered: "The last line I wrote was "No response requested." The +command did not finish — the session ended before the tool result was recorded, so +I received no output from `./slowjob.sh`." The resume process made no `approve` +call. + +This differs from test 1a. There, a SIGTERM during a running Bash call wrote a real +`Exit code 137` result before the exit. + +### R2-B1: a user frame during a round with two slow tool calls + +**Two `slow_echo` calls in one message (8 s and 20 s).** Both ran at once. Both +`approve` calls came first, and both `slow_echo` calls started within 350 ms. + +```text +t=5514 assistant tool_use toolu_013T… slow_echo {"text":"alpha","delay_s":8} +t=5543 MCP slow_echo#1 START alpha +t=5855 assistant tool_use toolu_012b… slow_echo {"text":"beta","delay_s":20} +t=5890 MCP slow_echo#2 START beta +t=8585 in user uuid=b1a00000-…-01 "Additional instruction: when you reply, also say the word MANGO." +t=8587 command_lifecycle b1a00000-… state=queued +t=13562 user tool_result toolu_013T… "echo: alpha" +t=25904 user tool_result toolu_012b… "echo: beta" +t=25921 command_lifecycle b1a00000-… state=started +t=30394 assistant text "echo: alpha\necho: beta" +t=30430 command_lifecycle b1a00000-… state=completed +t=30434 result subtype=success num_turns=3 result_index=0 + user_message_uuids=[, "b1a00000-0000-4000-8000-000000000001"] +``` + +**Bash `./slow15.sh` and `slow_echo` (5 s) in one message.** These ran one after +the other. The `approve` call for `slow_echo` came only after Bash finished. + +```text +t=10142 assistant tool_use toolu_01E8… Bash {"command":"./slow15.sh",…} +t=10166 MCP approve#1 CALLED tool=Bash (allowed at once) +t=10347 assistant tool_use toolu_01UN… slow_echo {"text":"gamma","delay_s":5} +t=14186 in user uuid=b1b00000-…-01 "Additional instruction: … MANGO." +t=14188 command_lifecycle b1b00000-… state=queued +t=25233 user tool_result toolu_01E8… "started-15\ndone-15" +t=25246 MCP approve#2 CALLED tool=mcp__slowtools__slow_echo +t=25253 MCP slow_echo#1 START gamma +t=30266 user tool_result toolu_01UN… "echo: gamma" +t=30299 command_lifecycle b1b00000-… state=started +t=33981 assistant text "PINEAPPLE" +t=34012 result subtype=success num_turns=4 + user_message_uuids=[, "b1b00000-0000-4000-8000-000000000001"] +``` + +An earlier run wrote the frame during the second call, the MCP one. It was taken +after that call's result, in the same way. + +- **A frame is taken only after the whole round.** It is not taken after the first + call's result, whether the calls run together (first result at 13.6 s, pickup at + 25.9 s) or one after the other (the Bash result, then the whole MCP call, then + pickup). In the transcript, the `queued_command` attachment follows the last + `tool_result` of the round, and the `queue-operation` `remove` has `reason: +"absorbed_mid_turn"`. +- **Same Turn, listed.** Each time there was one `result`, and the picked-up `uuid` + was in its `user_message_uuids`. +- Haiku did not say MANGO in either run. Delivery does not mean compliance, as in + test 4c. + +### R2-B2: a user frame during one MCP call, and during a permission wait + +**During one `slow_echo` (15 s).** + +```text +t=9359 assistant tool_use toolu_01Qq… slow_echo {"text":"delta","delay_s":15} +t=9387 MCP slow_echo#1 START delta +t=12410 in user uuid=b2a00000-…-01 "Additional instruction: … MANGO." +t=12412 command_lifecycle b2a00000-… state=queued +t=24400 user tool_result toolu_01Qq… "echo: delta" +t=24415 command_lifecycle b2a00000-… state=started +t=28050 assistant text "delta" +t=28085 result subtype=success num_turns=4 + user_message_uuids=[, "b2a00000-0000-4000-8000-000000000001"] +``` + +The MCP call was not cut short, and no `notifications/cancelled` was sent. The +frame was taken after the MCP result, as for Bash in test 4a. In a first run, the +frame arrived while the model was streaming the call that became `slow_echo`. It +was held through the whole MCP call and taken after its result, in the same Turn. + +**While Bash waits on the bridge (held 10 s, then `allow`).** + +```text +t=29043 assistant tool_use toolu_01Nn… Bash {"command":"./slow15.sh",…} +t=29068 MCP approve#1 CALLED tool=Bash -> will allow after 10000ms +t=32092 in user uuid=b2b00000-…-01 "Additional instruction: … MANGO." +t=32095 command_lifecycle b2b00000-… state=queued +t=39070 MCP approve#1 RETURNING {"behavior":"allow",…} +t=42120 system/task_started local_bash +t=54146 user tool_result toolu_01Nn… "started-15\ndone-15" +t=54168 command_lifecycle b2b00000-… state=started +t=56342 assistant text "PINEAPPLE MANGO" +t=56369 result subtype=success num_turns=2 + user_message_uuids=[, "b2b00000-0000-4000-8000-000000000001"] +``` + +- **The frame does not end or answer the permission wait.** It was held through + the rest of the wait, the approval, and the 15 s command. It was taken after the + tool result, in the same Turn, and listed. Haiku acted on it this time. + +### `command_lifecycle` in round 2 + +In every run: + +- A written frame got `queued` within 1 to 3 ms. +- A picked-up frame got `started` 15 to 33 ms after the round's last `tool_result` + frame, and `completed` just before the Turn's `result`. The prompt's `completed` + came just after the `result`. +- After an interrupt, the prompt got `cancelled` just after the `result`. +- With `cancel_queued`, the queued frame got `cancelled` just before the + `control_response`. + +## Round 3: SIGTERM re-approval and denied tool calls + +Run on 2026-09-29 against the same Claude Code **2.1.284**. Every fact in this +section is Recorded (2.1.284) unless it carries another label. + +### Round 3 method + +The round 3 driver and stand-in bridge were the round 2 ones, with the same launch +flags, the same `default` permission mode, no `--allowedTools`, and `CLAUDE*` and +`ORCA_*` removed from the child environment. Two changes: + +- The stand-in's answer policy could depend on the call's position, so the first + `approve` for a Bash call could be held and a second one for the same call + answered differently. It also called back when an answer was returned, so the + driver could write a frame at that moment. +- The command was `./mark.sh`, which appends a timestamp to `mark.out` and prints + `marked`. The driver deleted `mark.out` before each run and checked for it + before the signal, 2 s after the exit, and after the resume. `mark.out` was + absent at every check in every round 3 run. + +The B prompts said "run exactly this command: ./mark.sh ." and Haiku sometimes sent +the command as `./mark.sh .`. It was denied every time, so this made no difference. + +### R3-A: SIGTERM while Bash waits on the bridge (six runs) + +Each run held the first Bash `approve` for 120 s, then sent a process-group SIGTERM +after a delay. A second `approve` for the same call was held, answered `allow` after +50 ms, or answered `deny` after 50 ms. The `allow` run with a 3 s delay: + +```text +t=5168 MCP approve#1 CALLED tool=Bash tool_use_id=toolu_01Ct… reqId=2 (held 120 s) +t=8211 mark.out before SIGTERM: ABSENT +t=8212 SIGNAL SIGTERM group +t=8216 HTTP approve#1 response stream closed +t=8242 HTTP POST /mcp server/discover, then initialize → new MCP session bb09799d +t=8257 MCP approve#2 CALLED tool=Bash tool_use_id=toolu_01Ct… reqId=1 (same call, again) +t=8307 MCP approve#2 RETURNING {"behavior":"allow","updatedInput":{"command":"./mark.sh",…}} +t=8309 HTTP approve#2 response written in full (writableFinished=true) +t=9119 EXIT code=143 signal=null +t=11123 mark.out 2s after exit: ABSENT +``` + +| Run | SIGTERM after the first `approve` | Second `approve` after SIGTERM | Answer to it | Exit after SIGTERM | `mark.out` | +| ---------- | --------------------------------- | ------------------------------ | -------------- | ------------------ | ---------- | +| a_hold_05 | 0.5 s | 47 ms | held | 1,306 ms, 143 | absent | +| a_hold_3 | 3 s | 37 ms | held | 933 ms, 143 | absent | +| a_hold_10 | 10 s | 33 ms | held | 920 ms, 143 | absent | +| a_allow_05 | 0.5 s | 29 ms | `allow`, 50 ms | 978 ms, 143 | absent | +| a_allow_3 | 3 s | 45 ms | `allow`, 50 ms | 907 ms, 143 | absent | +| a_deny_3 | 3 s | 38 ms | `deny`, 50 ms | 951 ms, 143 | absent | + +- **The second `approve` came in all six runs.** Each time Claude Code opened a new + MCP session (`server/discover`, then `initialize`) and called `approve` again with + the same `tool_use_id` and input, 29 to 47 ms after the signal. It happened at + every delay tried, 0.5 s to 10 s. With R2-A2 that makes seven of seven runs. +- **Answering it did not run the command.** An `allow` delivered 50 ms after the + call, 80 to 95 ms after the signal, did not start `./mark.sh`: `mark.out` never + appeared, and no `task_started`, `tool_result`, or other stream frame followed + the signal. The process exited 812 to 899 ms after the answer, about as long as + when the call was held. A `deny` changed nothing either. +- **No frame after the signal.** No stream frame of any kind was written after the + SIGTERM, and no `result`. +- **The answer is not recorded.** In all six transcripts nothing follows the + `tool_use` until the resume. The `--resume` process wrote the same synthetic + `tool_result` as in R2-A2, with `toolDenialKind: "interrupted"`, whether the second + `approve` was held, allowed, or denied. It made no `approve` call. Each resumed + model said the command did not finish and its outcome was unknown. +- Claude Code opens the second `approve` but does not act on its answer before it + exits. **Inferred** from the six runs. Whether a slower shutdown could act on it + is Unknown. + +### R3-B1: a user frame while a Bash call waits on a denial (two runs) + +The stand-in denied the Bash `approve` after 8 s. The driver wrote a user frame 3 s +into the wait. Run 1: + +```text +t=4952 assistant tool_use toolu_01XY… Bash {"command":"./mark.sh",…} +t=4977 MCP approve#1 CALLED tool=Bash -> will deny after 8000ms +t=7985 in user uuid=b3100000-…-01 "Additional instruction: when you reply, also say the word MANGO." +t=7987 command_lifecycle b3100000-… state=queued +t=12978 MCP approve#1 RETURNING {"behavior":"deny","message":"Denied by the stand-in bridge."} +t=12987 user tool_result toolu_01XY… "Denied by the stand-in bridge." is_error +t=13003 command_lifecycle b3100000-… state=started +t=14954 assistant text "PINEAPPLE MANGO" +t=14970 command_lifecycle b3100000-… state=completed +t=14974 result subtype=success terminal_reason=completed num_turns=2 result_index=0 + queued_turn_count=0 + permission_denials=[{"tool_name":"Bash","tool_use_id":"toolu_01XY…",…}] + user_message_uuids=[, "b3100000-0000-4000-8000-000000000001"] +t=14975 command_lifecycle state=completed +``` + +Run 2 gave the same sequence: `queued` 2 ms after the write, `started` 15 ms after +the denied `tool_result`, the reply "PINEAPPLE MANGO", and one `result` listing the +frame. + +- **Taken at the denial's boundary, in the same Turn, and listed.** The frame was + held through the rest of the wait. It reached the model next to the denial's + `tool_result`, and its `uuid` was in the Turn's one `result`. Haiku acted on it in + both runs. +- In the transcript, the `tool_result` is stored with `toolUseResult: "Error: Denied +by the stand-in bridge."` and `toolDenialKind: "permission-rule"`. The + `queued_command` attachment follows it, and the `queue-operation` `remove` has + `reason: "absorbed_mid_turn"`, as in rounds 1 and 2. +- The denial's `message` reaches the model as the `tool_result` text. + +### R3-B2: a user frame written as the denial is returned (three runs) + +The stand-in denied after 5 s. The driver wrote the frame from the stand-in's +return callback: at once (two runs) or 40 ms later (one run). + +```text +t=9663 MCP approve#1 RETURNING {"behavior":"deny",…} +t=9663 in user uuid=b3200000-…-01 "Additional instruction: … MANGO." +t=9671 user tool_result toolu_013H… "Denied by the stand-in bridge." is_error +t=9698 command_lifecycle b3200000-… state=queued +t=11285 assistant text "PINEAPPLE" +t=11307 result subtype=success num_turns=2 result_index=0 user_message_uuids=[] +t=11309 command_lifecycle state=completed +t=11310 command_lifecycle b3200000-… state=started +t=15623 assistant text "PINEAPPLE" +t=15673 result subtype=success num_turns=1 result_index=1 + user_message_uuids=["b3200000-0000-4000-8000-000000000001"] +``` + +| Run | Frame written, relative to the denied `tool_result` frame | `queued` after the write | Taken in the Turn? | +| ----- | --------------------------------------------------------- | ------------------------ | ------------------ | +| b2 | 8 ms before | 35 ms | No, next Turn | +| b2-r2 | 8 ms before | 31 ms | No, next Turn | +| b2d | 33 ms after | 2 ms | No, next Turn | + +- **Not taken in the running Turn.** In all three runs the frame was not in the + first `result`'s `user_message_uuids`. It ran by itself as the next Turn + (`result_index: 1`) with no new stdin write, as in test 4b. In each transcript + the `enqueue` comes after the denial's `tool_result`, with a `dequeue` for the + next Turn and no `queued_command` attachment. +- **The window closes very close to the `tool_result`.** A frame written 8 ms + before the `tool_result` frame was acknowledged `queued` 31 to 35 ms after the + write, which was after the pickup point. In R3-B1 the pickup came 15 to 16 ms + after the `tool_result`. So a frame that is not yet `queued` when the denial + lands can miss the Turn. **Inferred** from three runs; the exact cut-off was not + measured. +- In b2d the second Turn's model refused the instruction, saying it had already + answered. Delivery does not mean compliance. + +### R3-B3: one call denied, one allowed, in one round (one run) + +The prompt asked for Bash `./mark.sh` and `slow_echo("gamma", 5)` in one message. +The model first called `ToolSearch` for `slow_echo` in a round of its own. The +stand-in denied the Bash `approve` after 6 s and allowed `slow_echo` at once. The +driver wrote the frame 3 s into the Bash wait. + +```text +t=7098 assistant tool_use toolu_01V2… Bash {"command":"./mark.sh",…} +t=7118 MCP approve#1 CALLED tool=Bash -> will deny after 6000ms +t=7296 assistant tool_use toolu_01Qx… mcp__slowtools__slow_echo {"text":"gamma","delay_s":5} +t=10135 in user uuid=b3300000-…-01 "Additional instruction: … MANGO." +t=10137 command_lifecycle b3300000-… state=queued +t=13119 MCP approve#1 RETURNING {"behavior":"deny",…} +t=13129 user tool_result toolu_01V2… "Denied by the stand-in bridge." is_error +t=13142 MCP approve#2 CALLED tool=mcp__slowtools__slow_echo (allowed at once) +t=13152 MCP slow_echo#1 START gamma +t=18166 user tool_result toolu_01Qx… "echo: gamma" +t=18182 command_lifecycle b3300000-… state=started +t=22955 assistant text "PINEAPPLE" +t=22976 result subtype=success num_turns=4 result_index=0 + permission_denials=[{"tool_name":"Bash",…}] + user_message_uuids=[, "b3300000-0000-4000-8000-000000000001"] +``` + +- **A denied call does not end the round early.** The frame was not taken after the + denial. The `slow_echo` call's `approve` came 13 ms after the denial, and the + frame was taken 16 ms after the `slow_echo` result, the round's last. This is + the same whole-round rule as R2-B1. +- **Same Turn, listed.** One `result`, with the frame's `uuid` in + `user_message_uuids` and the denial in `permission_denials`. Haiku did not say + MANGO. + +### `command_lifecycle` in round 3 + +- A frame written during a wait got `queued` within 2 ms. A frame written within a + few milliseconds of the denial landing got it 31 to 35 ms later. +- A frame taken at a boundary got `started` 15 to 16 ms after the round's last + `tool_result`, `completed` 3 to 4 ms before the `result`, and the prompt's + `completed` 1 to 2 ms after it. +- A frame that missed the Turn got `started` 1 ms after the prompt's + `completed`, then ran as the next Turn. + +## The Four Original Unknowns + +| Unknown in [Harness Interrupt and Queued Messages](harness-interrupt-and-queued-messages.md) | Result on 2.1.284 | Evidence | +| -------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------- | +| Does SIGTERM in `-p` mode keep the partial assistant text in the transcript that `--resume` loads? | **No for streamed text:** nothing of the unfinished model call is saved, not even thinking. **Yes for a finished tool round:** the `tool_use` and a real `Exit code 137` result with partial stdout are saved. There is no interrupted marker. Resume adds a synthetic `No response requested.`. **Round 2:** for a call waiting on the permission bridge nothing is saved after the `tool_use`; resume adds a synthetic "Tool call interrupted… outcome is unknown" `tool_result`. **Round 3:** the same in six more runs, whether the bridge held, allowed, or denied the second `approve` that shutdown sends. | Recorded (2.1.284) | +| Does SIGINT to a `-p --input-format stream-json` process end only the Turn and keep reading stdin? | **No.** It writes one `error_during_execution` `result`, then exits with code 0 about 1 s later, with stdin still open. The partial Turn is kept for `--resume`. A still-queued message is lost. | Recorded (2.1.284) | +| Is a raw `control_request` `interrupt` honoured without the SDK's `initialize`? | **Yes.** `control_response` success with the receipt in 2 to 4 ms, then an `error_during_execution` `result`. The process stays alive, and the next user frame runs in the same Session with the partial Turn in context. `cancel_queued` works raw. **Round 2:** the same during a permission-bridge wait, and the bridge call gets MCP `notifications/cancelled`. | Recorded (2.1.284) | +| Minimum version for headless mid-Turn pickup and for `priority` | **Not named** in the changelog. Pickup between tool rounds is Recorded on 2.1.284. `priority` was not sent. | Documented (none found); Recorded (2.1.284) | + +## Capability Table + +| Stop or send on the raw `-p` stream | Turn ends with | Process | Partial text in context | Tool call in context | Foreground Bash tree | Queued message | +| -------------------------------------------------------------- | ----------------------------------------------------------------------------------------- | ------------------ | -------------------------- | ------------------------------------------------------------------------------------- | ------------------------------------------------------------------ | ---------------------------------------------------------------------------------------- | +| SIGTERM (group) | No `result` | Exits 143 | No, not even on resume | Yes: `tool_use` plus `Exit code 137` and partial stdout | Killed by the CLI | Not tested | +| SIGINT (group or pid) | `result` `error_during_execution`, `aborted_streaming` / `aborted_tools` | Exits 0 after ~1 s | Yes, on resume | Yes, as a "rejected" result; partial stdout discarded | Killed | Lost; not replayed on resume | +| `control_request` `interrupt` | `control_response` receipt, then the same `result` | Stays alive | Yes, in the next Turn | Yes, as a "rejected" result; partial stdout discarded | Killed | Runs next by itself; listed in `still_queued` | +| `control_request` `interrupt` with `cancel_queued: true` | Same, with `cancelled: [uuid]` | Stays alive | Not tested (mid-tool only) | Same as above | Killed | Dropped; `command_lifecycle` `cancelled` | +| User frame during a tool call | Same Turn continues | Stays alive | n/a | n/a | Runs to completion | Delivered with the tool result; in `user_message_uuids` | +| `interrupt` while a tool call waits on the permission bridge | Same `result`; `permission_denials` names the call; bridge gets `notifications/cancelled` | Stays alive | n/a | Yes, as a "rejected" result; the tool never ran | Never started; a late `allow` has no effect | With `cancel_queued`: dropped, as above | +| SIGTERM while a tool call waits on the permission bridge | No `result`; a second `approve` for the same call during shutdown, in seven of seven runs | Exits 143 | n/a | Only on resume: a synthetic "Tool call interrupted… outcome is unknown" result | Never started, even when the second `approve` was answered `allow` | Not tested | +| User frame during an MCP call, or a round of two or more calls | Same Turn continues | Stays alive | n/a | n/a | Runs to completion | Delivered after the round's last tool result, not between calls; in `user_message_uuids` | +| User frame while a tool call waits on the permission bridge | Same Turn continues | Stays alive | n/a | n/a | Runs after approval | Held through approval and the tool; delivered with its result; in `user_message_uuids` | +| User frame while a tool call waits on a bridge denial | Same Turn continues | Stays alive | n/a | Yes, as an error `tool_result` with the denial message; `permission_denials` names it | Never started | Held through the wait; delivered with the denial's result; in `user_message_uuids` | +| User frame written as the bridge denial lands | First Turn ends normally | Stays alive | n/a | Same as above | Never started | Missed the boundary in three of three runs; runs next as its own Turn | +| User frame while text streams, no tool round left | First Turn ends normally | Stays alive | n/a | n/a | n/a | Runs next as its own Turn with its own `result` | + +## Still Unknown + +- Whether a stdin user frame with `priority: "now"` preempts the running Turn on + the raw stream, and whether `"later"` holds a frame past a tool boundary. + `priority` was never sent. +- Why `queued_turn_count` stayed `0` while a message was waiting, and what makes it + non-zero. +- Whether an interrupted mid-text Turn's tokens are counted anywhere. Its `result` + reported `total_cost_usd: 0` and zero usage. +- What SIGINT does while a tool call waits on the permission bridge. Round 2 + tested only the raw `interrupt` and SIGTERM there. +- Why SIGTERM sends a second `approve` for the waiting call, and whether any + answer to it can take effect. It came in seven of seven runs. Round 3 answered it + within 50 ms in three runs and the command never ran, but a slower shutdown, or + an answer that lands at another moment, was not tested. +- Whether a user frame is taken between calls when a round has a call that fails + on its own, rather than being denied. Round 3 found that a denied call does not + end the round early (one run). +- The exact cut-off for a frame to be taken at a denial's boundary. Frames written + 5 s before the denial were taken; frames written 8 ms before or 33 ms after its + `tool_result` frame were not. Nothing in between was tried. +- Whether MCP tools without `readOnlyHint` run at the same time in one round. The + stand-in's `slow_echo` set it, and a Bash call plus an MCP call ran one after the + other. +- How a SIGTERM that arrives while a user frame is queued treats that frame. It was + not tested, but the SIGINT result suggests it is lost. **Inferred.** +- Whether `CLAUDE_CODE_RESUME_INTERRUPTED_TURN=1` changes what a resume after + SIGTERM or SIGINT sees. It was not set. +- Whether a stronger model follows a mid-Turn message that conflicts with the + original prompt. Haiku ignored two in test 4c. +- Windows: every behaviour in this note. diff --git a/docs/research/codex-live-leftover-steer-and-empty-turn-start.md b/docs/research/codex-live-leftover-steer-and-empty-turn-start.md new file mode 100644 index 00000000..ebf61d92 --- /dev/null +++ b/docs/research/codex-live-leftover-steer-and-empty-turn-start.md @@ -0,0 +1,252 @@ +# Codex Live Leftover Steer and Empty Turn Start + +Research date: 2026-09-29 + +Harness version examined: Codex **codex-cli 0.157.1** (installed standalone build, Linux x64, `codex --version` returned `codex-cli 0.157.1`), +source read at upstream tag `rust-v0.157.1`, commit +[`36650394c5b38c2990ccf2a3457165ca3e9d9726`](https://github.com/openai/codex/commit/36650394c5b38c2990ccf2a3457165ca3e9d9726). + +Ticket: [#255](https://github.com/secantdev/secant/issues/255). This note runs live `codex app-server` sessions to settle three questions that +[Harness Interrupt and Queued Messages](harness-interrupt-and-queued-messages.md) left open and that ADR 0035 (on `docs/interrupt-steer`) depends +on. ADR 0035 says that when Codex takes a steered text into history after its last model request and completes the Turn without answering it, +the Codex Adapter re-delivers that leftover with a native `turn/start`. It would use an empty input if Codex accepts one, and the same text +otherwise. + +## Answer + +**1. Codex accepts an empty-input `turn/start` on an idle thread, and the model answers the leftover.** `turn/start` with `input: []` returns a +normal `{ turn }`. Codex then emits `turn/started`, no `userMessage` item, one model response, and `turn/completed` `status: "completed"`. The +model responds to whatever the history already holds. With a leftover steer at the end of the history, it answered the steered question in all +4 of 4 runs, and `thread/read` then shows the steered text exactly once. An input of one empty text item is accepted too, but it records an +empty `userMessage` and the model returned an empty reply. While a Turn is active, an empty `turn/start` is refused with +`-32603 "failed to submit turn input: EmptyInput"`, and an empty `turn/steer` with `-32600 "input must not be empty"`.[^empty-src] + +**2. The leftover window was reproduced, but not during Stop hooks.** A `turn/steer` sent while a 12-second Stop hook runs is accepted, and Codex +answers it **in the same Turn** (2 of 2 runs). After the hooks, it drains the steer as a `userMessage` with the `clientId`, makes another model +request, runs the Stop hooks again, and only then completes. The source shows why: the regular task re-runs the turn loop whenever pending input +remains after it returns.[^regular-loop] The real window is the gap between that last pending-input check and task finish. A `turn/steer` +written the moment the Stop hook's `hook/completed` arrives hit it in **6 of 22** trials (the pilot's 1 of 2 plus 5 of 20 in the main run). In +those trials Codex accepted the steer, emitted its `userMessage` item with the matching `clientId`, and then sent `turn/completed` +`status: "completed"` with no model output after the item. The other 16 were refused with `-32600 "no active turn to steer"`. With a 1 to 3 ms +delay, all 4 trials were refused. + +**3. Re-delivering the same text writes it into history twice.** Given a leftover, `turn/start` with the same text (and the same +`clientUserMessageId`) is answered normally. `thread/read` then lists two `userMessage` items with the same text and the same `clientId`: one in +the leftover Turn and one in the re-delivery Turn. Asked how many times it had been asked the question, the model answered `2` (2 of 2). After an +empty re-delivery it answered `1` (3 of 3; the pilot run asked a different question). + +**4. Nothing experimental was used.** Every method, field, and notification in these runs is in the stable schema that the installed binary +generates (`codex app-server generate-json-schema`, without `--experimental`), and every session initialized with +`capabilities.experimentalApi: false`. `TurnStartParams.input` is a plain array with no `minItems`, both there and in Secant's recorded fixture +`tests/harness/fixtures/codex/codex-qualification/stable-schema.generated.json`. Only the test setup was out of the ordinary: the Stop hook was +added through the per-thread `config` override `bypass_hook_trust: true`, which Codex flags as dangerous (see Method). + +## Evidence Vocabulary + +- **Observed**: seen in these live runs, quoted from the raw frames. +- **Not observed**: looked for in these runs and absent. +- **Source-observed**: read in the Codex source at the commit above. +- **Untested**: implied by the source but not run. + +## Method + +A Python driver spawned `codex app-server` (stdio JSONL) in a throwaway working directory under the session scratchpad, not in the repository. +It sent Secant's handshake from `src/harness/codex/qualification.ts`: + +```json +{"method":"initialize","params":{"clientInfo":{"name":"secant","title":"Secant","version":"0.0.0-dev"},"capabilities":{"experimentalApi":false}}} +{"method":"initialized"} +``` + +It then sent `thread/start {cwd}` and `turn/start {threadId, input:[{type:"text",text}], model}` in the shape `src/harness/codex.ts` uses, and +`turn/steer {threadId, expectedTurnId, input, clientUserMessageId}`, the ADR 0035 shape. It logged every frame with a millisecond offset +(`t`, seconds since spawn below). Any server request would have been declined automatically. + +Differences from Secant's launch: + +- Every `turn/start` also set `effort: "low"` to keep cost down. Secant sets no effort. The model was `gpt-6-luna` ("Fast and affordable model + for easier tasks" in `model/list`), not the user's configured `gpt-6-sol`. +- The user's real `~/.codex` was the Codex home, unmodified, as in Secant's user-compatible launch. Its `hooks.json` registers + Orca command hooks, including `SessionStart`, `UserPromptSubmit`, and `Stop`. They ran on every Turn and appear in the frames. Its MCP servers + (`context7`, `codex_apps`, and two that failed to start) loaded too. +- To keep a Turn open during a Stop hook, experiment 2 started its threads with a per-thread config override: + + ```json + { + "bypass_hook_trust": true, + "hooks.Stop": [ + { "hooks": [{ "type": "command", "command": "sleep 12", "timeout": 60 }] } + ] + } + ``` + + The race variant used the same override with `"command": "true"`. Hooks from a config layer run only when they are trusted, and a + `bypass_hook_trust` request override lifts that gate for the session.[^hook-trust] Codex answered with + `configWarning` "`--dangerously-bypass-hook-trust` is enabled. Enabled hooks may run without review for this invocation." The hook ran + from the `sessionFlags` source (`"sourcePath": "//config.toml"`). No file in `~/.codex` was written by hand, and + `config.toml` and `hooks.json` kept their earlier modification times. Codex still wrote its own session rollouts, as any run does. + +- No temporary `CODEX_HOME` was needed, so none was created and `auth.json` was never touched. +- Every `codex app-server` the driver started exited when the driver closed stdin. The only `codex` processes left afterwards predate the runs: + the user's 0.158.0 app-server daemon, the VS Code extension, and an interactive `codex resume`. + +## Experiment 1: empty-input `turn/start` + +Each run used one thread and three Turns, one after another, each started after the previous `turn/completed`. + +**Observed.** T1 `"Remember the code word ZEBRA-42. Reply with just: OK"` completed with `OK`. T2 then sent `input: []`: + +```text +[4.748] out turn/start {"threadId":"…d1ad6","input":[],"model":"gpt-6-luna","effort":"low"} +[4.778] in response {"turn":{"id":"…d1d46","items":[],"status":"inProgress",…}} +[4.781] in turn/started {"turn":{"id":"…d1d46",…}} +[6.187] in item/completed agentMessage text="OK" +[6.224] in hook/started stop (user hooks.json) +[6.247] in turn/completed {"turn":{"id":"…d1d46","status":"completed","error":null,…}} +``` + +- **Not observed** in T2: a `userMessage` item, or a `UserPromptSubmit` hook run. Both appear in T1 and in T3. +- The model's reply repeated T1's answer: with no new input, it answered the last user message in history again. +- `thread/read` (`includeTurns: true`) lists T2 as `completed` with only `agentMessage "OK"`. + +T3 sent `input: [{"type":"text","text":""}]`. Codex recorded +`userMessage content=[{"type":"text","text":"","text_elements":[]}]`, ran the `UserPromptSubmit` hook, and the model returned +`agentMessage text=""`. `turn/completed` had `status: "completed"` and `items: []`. + +**Addendum, fresh thread (1 run).** An empty `turn/start` as the first Turn of a new thread was also accepted. With no user message in history, +the model made up a task from the injected context. It wrote a commentary message ("I'll check the current Codex docs and your local config +guidance first…"), called two `context7` MCP tools, tried a shell command (which failed), and gave a final answer about running Codex subagents in +parallel. + +**Observed, empty input during an active Turn** (experiment 2 mode `c`, sent while the Stop hook ran): + +```text +[7.440] turn/steer {"expectedTurnId":"…dd47","input":[]} -> {"error":{"code":-32600,"message":"input must not be empty"}} +[7.442] turn/start {"input":[]} -> {"error":{"code":-32603,"message":"failed to submit turn input: EmptyInput"}} +``` + +The Turn then completed normally with `ALPHA`. + +**Source-observed.** `start_or_steer` treats an empty `UserInput` as `has_explicit_input = false`. With no active Turn, it spawns the regular +task without pushing any input, so the model request is built from the existing history alone. With an active Turn, `steer_input` returns +`EmptyInput`, which `turn/start` reports as `failed to submit turn input: EmptyInput`.[^empty-src] + +## Experiment 2: reproducing a leftover steer + +### 2a. Steer during a Stop hook (answered in the same Turn) + +Prompt `"Reply with just: ALPHA"`. The driver waited for the session-flags Stop hook's `hook/started`, then 1 s, then sent +`turn/steer` with `clientUserMessageId: "secant-steer-n2"`. Run `n2` (run `n1` matched it): + +```text +[ 6.108] item/completed agentMessage text="ALPHA" +[ 6.144] hook/started stop sourcePath=//config.toml (sleep 12) +[ 7.146] turn/steer -> {"result":{"turnId":"…b256"}} +[18.146] hook/completed stop +[18.202] item/completed userMessage clientId="secant-steer-n2" "New question: what is 17+25? Reply with just the number." +[19.944] item/completed agentMessage text="42" +[19.977] hook/started stop (second run of the Stop hooks) +[31.985] hook/completed stop +[31.993] turn/completed {"turn":{"id":"…b256","status":"completed",…}} +``` + +**Observed.** A steer sent during a Stop hook is accepted, drained after the hooks, and answered within the same Turn, and the Stop hooks run +again. `thread/read` shows one Turn holding `userMessage ALPHA`, `agentMessage ALPHA`, `userMessage (clientId secant-steer-n2) 17+25`, +`agentMessage 42`. So a Stop hook does not open a leftover window at this version. The pending-input check that runs before the Stop hooks sees +nothing, but the regular task checks again after `run_turn` returns and re-enters it with the pending input.[^regular-loop] + +### 2b. Steer at the Stop hook's `hook/completed` (the leftover) + +The window left is the gap between that final check (`regular.rs` line 120) and `on_task_finished`. There Codex takes the active task and then +records any pending input to history, running `UserPromptSubmit` hooks on it and emitting its `userMessage` item.[^task-finish] To aim at it, the +driver used a fresh thread per trial with a no-op session-flags Stop hook (`true`). The reader thread wrote `turn/steer` as soon as it parsed +the first Stop `hook/completed` of the Turn (0 ms delay), or after a fixed delay. + +| Delay after Stop `hook/completed` | Trials | Leftover (accepted, item, no answer) | Refused `no active turn to steer` | Answered in the Turn | +| --------------------------------- | -----: | -----------------------------------: | --------------------------------: | -------------------: | +| 0 ms (pilot and main run) | 22 | 6 | 16 | 0 | +| 1 ms | 2 | 0 | 2 | 0 | +| 3 ms | 2 | 0 | 2 | 0 | + +Hit rate at 0 ms: 6 of 22 (27%). The steer response always came back within 1 to 2 ms. By the server's own `emittedAtMs` stamps, +`turn/completed` followed the Stop `hook/completed` by 3 to 6 ms in refused trials and by 43 to 58 ms in leftover trials. The extra time is +spent recording the leftover, including the user's `UserPromptSubmit` hook. + +**Observed, decisive frames** (trial `r0-0`, emitted in this order within about 4 ms): + +```json +{"method":"item/started","params":{"item":{"type":"userMessage","id":"01a0ebee-5b21-…","clientId":"secant-steer-r0-0","content":[{"type":"text","text":"New question: what is 17+25? Reply with just the number.","text_elements":[]}]},"turnId":"01a0ebee-3d03-…"}} +{"method":"item/completed","params":{"item":{"type":"userMessage","id":"01a0ebee-5b21-…","clientId":"secant-steer-r0-0",…},"turnId":"01a0ebee-3d03-…"}} +{"method":"turn/completed","params":{"turn":{"id":"01a0ebee-3d03-…","items":[{"type":"agentMessage","text":"OK","phase":"final_answer",…}],"itemsView":"summary","status":"completed","error":null,…}}} +``` + +- The `turn/steer` response was `{"result":{"turnId":"01a0ebee-3d03-…"}}`, the same success shape as a delivered steer. +- **Not observed** after the item: any `agentMessage`, any further model request, or any `error` notification. +- The `turn/completed` summary lists only the final `agentMessage`. It never lists `userMessage` items in any Turn. +- `thread/read` places the leftover `userMessage` inside the completed Turn, after `agentMessage "OK"`. +- In the pilot trial, the user's `UserPromptSubmit` hook ran between the steer and the leftover item. + +**Source-observed, not reproduced.** The regular task also returns early, without re-checking pending input, when the Turn has a terminal +error. Pending input then reaches `on_task_finished` the same way, which is the "failing Turn" case.[^regular-loop] Post-turn compaction runs +inside `run_turn` before it returns, so the same re-check covers it. That case is Untested. + +## Experiment 3: re-delivery of a leftover + +Each leftover in 2b was followed on the same thread by a re-delivery `turn/start`, then an account Turn. The account Turn asked: "Ignore any +AGENTS.md or environment context. Counting only messages I typed in this chat before this one: how many times did I ask you what 17+25 is? Reply +with just the number." After that came a `thread/read`. The pilot asked for the messages listed verbatim instead, and the model listed its +injected AGENTS.md context, so only its `thread/read` counts. + +### 3a. Empty input (4 leftovers: pilot, `r0-0`, `r0-14`, `r0-19`) + +```text +out turn/start {"threadId":"…20fa68","input":[],"model":"gpt-6-luna","effort":"low"} +in response {"turn":{"id":"01a0ebee-5b2d-…","status":"inProgress",…}} +in item/completed agentMessage text="42" phase=final_answer +in turn/completed {"turn":{"id":"01a0ebee-5b2d-…","status":"completed",…}} +account Turn -> "1" +thread/read: + TURN …3d03 completed: userMessage "Reply with just: OK" | agentMessage "OK" | userMessage clientId=secant-steer-r0-0 "New question: what is 17+25? …" + TURN …5b2d completed: agentMessage "42" +``` + +**Observed.** In 4 of 4 runs, the empty `turn/start` produced an answer to the leftover (`42`) with no new `userMessage` item. The steered text +appears once in history, and the model counted it once in 3 of 3 account Turns. + +### 3b. Same text and same `clientUserMessageId` (2 leftovers: `r0-2`, `r0-17`) + +```text +out turn/start {"input":[{"type":"text","text":"New question: what is 17+25? Reply with just the number."}],"clientUserMessageId":"secant-steer-r0-2",…} +in item/completed userMessage clientId="secant-steer-r0-2" "New question: what is 17+25? …" +in item/completed agentMessage text="42" +in turn/completed status=completed +account Turn -> "2" +thread/read: + TURN …8b66 completed: userMessage "Reply with just: OK" | agentMessage "OK" | userMessage clientId=secant-steer-r0-2 "New question: what is 17+25? …" + TURN …9b50 completed: userMessage clientId=secant-steer-r0-2 "New question: what is 17+25? …" | agentMessage "42" +``` + +**Observed.** In 2 of 2 runs, the same-text re-delivery is answered. History then holds the text twice, as two `userMessage` items with the same +`clientId` in two Turns, and Codex did not de-duplicate on `clientUserMessageId`. The model counted the question twice in both account Turns. + +## Unknowns and caveats + +- The race hit rate depends on this host and on the user's hooks. The Orca `UserPromptSubmit` hook lengthens the finishing path, and a client + that is not reacting to `hook/completed` will land in the window only by chance. The rate is not a property of Codex. +- The failing-Turn path (terminal error, then pending input recorded at task finish) was read in source and not reproduced. +- Only one model (`gpt-6-luna`, effort `low`) was used, with trivial prompts. Whether larger models answer the leftover as reliably after an + empty `turn/start` is Untested. In a longer history, an empty start makes the model respond to "whatever history holds", which was a + made-up task on a thread with no user message. +- Timestamps are the client's read times, rounded to 1 ms, except where `emittedAtMs` is named. +- The Stop hook in experiment 2 depended on `bypass_hook_trust` and the session-flags hook layer. Secant uses neither. They were only a way to + open the window on purpose. + +## Primary Sources + +[^empty-src]: OpenAI Codex source at `36650394`: [`core/src/session/turn_input.rs` lines 289-300](https://github.com/openai/codex/blob/36650394c5b38c2990ccf2a3457165ca3e9d9726/codex-rs/core/src/session/turn_input.rs#L289-L300) (`has_explicit_input`), [lines 356-362](https://github.com/openai/codex/blob/36650394c5b38c2990ccf2a3457165ca3e9d9726/codex-rs/core/src/session/turn_input.rs#L356-L362) (an empty input spawns the task with no pushed input), [lines 632-638 and 665-667](https://github.com/openai/codex/blob/36650394c5b38c2990ccf2a3457165ca3e9d9726/codex-rs/core/src/session/turn_input.rs#L632-L667) (`NoActiveTurn`, `EmptyInput` in `steer_input`); [`app-server/src/request_processors/turn_processor.rs` lines 596-620 and 672-682](https://github.com/openai/codex/blob/36650394c5b38c2990ccf2a3457165ca3e9d9726/codex-rs/app-server/src/request_processors/turn_processor.rs#L596-L682) (`turn/start` input mapping and `failed to submit turn input: {reason:?}`), [lines 1084-1085 and 1133-1134](https://github.com/openai/codex/blob/36650394c5b38c2990ccf2a3457165ca3e9d9726/codex-rs/app-server/src/request_processors/turn_processor.rs#L1084-L1134) (`turn/steer` errors `no active turn to steer`, `input must not be empty`). + +[^regular-loop]: OpenAI Codex source, [`core/src/tasks/regular.rs` lines 104-122](https://github.com/openai/codex/blob/36650394c5b38c2990ccf2a3457165ca3e9d9726/codex-rs/core/src/tasks/regular.rs#L104-L122) (loop over `run_turn`, early return on `terminal_error`, re-run while `has_pending_input`); [`core/src/session/turn.rs` lines 551-563 and 640-733](https://github.com/openai/codex/blob/36650394c5b38c2990ccf2a3457165ca3e9d9726/codex-rs/core/src/session/turn.rs#L551-L733) (post-sampling pending-input check, then Stop hooks and post-turn compaction before `break`). + +[^task-finish]: OpenAI Codex source, [`core/src/tasks/mod.rs` lines 621-684](https://github.com/openai/codex/blob/36650394c5b38c2990ccf2a3457165ca3e9d9726/codex-rs/core/src/tasks/mod.rs#L621-L684) (`on_task_finished` takes the task, takes pending input, and records it via `run_hooks_and_record_inputs`); [`core/src/session/turn.rs` lines 839-887](https://github.com/openai/codex/blob/36650394c5b38c2990ccf2a3457165ca3e9d9726/codex-rs/core/src/session/turn.rs#L839-L887) (`UserPromptSubmit` inspection, then `record_pending_input`). + +[^hook-trust]: OpenAI Codex source, [`app-server/src/config_manager.rs` lines 438-445](https://github.com/openai/codex/blob/36650394c5b38c2990ccf2a3457165ca3e9d9726/codex-rs/app-server/src/config_manager.rs#L438-L445) (`bypass_hook_trust` request override), [`hooks/src/engine/discovery.rs` lines 714-719 and 831](https://github.com/openai/codex/blob/36650394c5b38c2990ccf2a3457165ca3e9d9726/codex-rs/hooks/src/engine/discovery.rs#L714-L831) (only trusted, managed, or bypassed handlers run; `SessionFlags` hook source), [`config/src/hook_config.rs` lines 19-25 and 57-58](https://github.com/openai/codex/blob/36650394c5b38c2990ccf2a3457165ca3e9d9726/codex-rs/config/src/hook_config.rs#L19-L58) (`hooks.` and `hooks.state` in config), [`features/src/lib.rs` lines 1211-1216](https://github.com/openai/codex/blob/36650394c5b38c2990ccf2a3457165ca3e9d9726/codex-rs/features/src/lib.rs#L1211-L1216) (`hooks` feature stable, on by default). diff --git a/docs/research/harness-interrupt-and-queued-messages.md b/docs/research/harness-interrupt-and-queued-messages.md index c17e5f20..de982c49 100644 --- a/docs/research/harness-interrupt-and-queued-messages.md +++ b/docs/research/harness-interrupt-and-queued-messages.md @@ -279,7 +279,9 @@ installed version's tag. T3 Code was read from a local clone at the commit above retried. A `lost` Turn maps to `indeterminate`, and the Run also rests `halted`. `resume-run` re-attempts the Step. Because the Session record is `detached`, the retry sends the Step's prompt again into the same native Session.[^sec-exec-cancelled][^sec-agent-recovery] - **Interactive agent step.** - - Each human Turn runs with `detachAfterTurn`, and the Harness is closed after each Turn.[^sec-agent-interactive] + - Each human Turn runs with `detachAfterTurn`, which records the Session `detached` after each Turn.[^sec-agent-interactive] (Corrected + 2026-09-29: the Harness is not closed after each Turn. The wiring keeps one prepared Harness, and one Claude Code process, across the + Step's Turns; only Codex reissues `thread/resume` before each human Turn because of the `detached` record.) - An interrupted or lost human Turn rests the Run `halted`, not `blocked`, and publishes no Attempt. An interrupted Entry Turn does the same.[^sec-app-interactive][^sec-exec-entry] - The glossary says that "after a halt the human continues the same Harness Session". That takes a resume back to `blocked` before the diff --git a/docs/research/windows-live-interrupt-and-steer.md b/docs/research/windows-live-interrupt-and-steer.md new file mode 100644 index 00000000..9139e29c --- /dev/null +++ b/docs/research/windows-live-interrupt-and-steer.md @@ -0,0 +1,634 @@ +# Windows Live Interrupt and Steer + +Research date: 2026-09-29 + +Harness versions examined, on Windows 11 Home 10.0.26200 (x64): + +- Claude Code **2.1.283**, the WinGet native executable (`claude --version` returned + `2.1.283 (Claude Code)`), resolved on PATH as + `%LOCALAPPDATA%\Microsoft\WinGet\Links\claude.exe`, a symbolic link into the WinGet package. +- Codex **codex-cli 0.155.0**, the standalone build at + `%LOCALAPPDATA%\Programs\OpenAI\Codex\bin\codex.exe`. + +Ticket: [#255](https://github.com/secantdev/secant/issues/255). The Linux rounds in +[Claude Code Live Interrupt and Mid-Turn Send](claude-code-live-interrupt-and-mid-turn-send.md) +(on `research/255-claude-live-interrupt`, Claude Code 2.1.284) and +[Codex Live Leftover Steer and Empty Turn Start](codex-live-leftover-steer-and-empty-turn-start.md) +(codex-cli 0.157.1) left Windows untested. This note repeats the stops and sends that +[ADR 0035](../adr/0035-interrupt-ends-only-the-turn-and-a-mid-turn-message-is-a-native-steer.md) +relies on, against the installed Windows CLIs. The Windows builds are one patch release (Claude +Code) and two minor releases (Codex) older than the Linux ones. + +## Answer + +**W1. The raw `control_request` `interrupt` works on Windows without `initialize`. The Turn +stops and the process and Session stay alive. But it does not kill a script that the Bash tool +runs.** The `control_response` came back 6 to 30 ms after the write, in 12 of 12 runs, with the +receipt `{"still_queued":[…]}`. The Turn then ended with the same `result` as on Linux: +`subtype: "error_during_execution"`, `terminal_reason` `aborted_tools` or `aborted_streaming`. The +next stdin user frame ran as `result_index: 1` in the same process and `session_id`. Mid-text, +the partial text stayed in context (`aborted: true`, `isAbortedMidStream: true`), and the model +quoted its last partial line. Mid-tool, the `tool_result` was the standard "The user doesn't want +to proceed with this tool use… rejected" text, and the model said the command "never executed". +To stop the tool, Claude Code runs `taskkill /PID /T /F` (seen in the process +table). That reached a native command run directly by the Bash tool (`ping`) and a PowerShell-tool +script, and both were killed. It did not reach `./slowjob.sh` run by the Bash tool, which runs +through Git Bash. The script's `bash.exe` had a Windows parent pid that no longer existed, so it +was outside the tree. It ran to completion in all 6 runs where the raw interrupt stopped it, and in +5 of them it was still running after Claude Code had exited. A backgrounded Bash task was not stopped by the interrupt. Claude Code +reported it `killed` when the process exited, but its script also ran on. + +**W2. `cancel_queued: true` behaves as on Linux.** The queued frame was listed in `cancelled`, +with a `command_lifecycle` `cancelled` frame just before the `control_response`. No Turn ran in +the next 10 s, and the process stayed alive (2 of 2 runs). The transcript holds `enqueue`, then +`remove`, and no `queued_command` attachment. A plain `interrupt` listed the frame in +`still_queued` and ran it by itself as the next Turn, 10 ms after the interrupted `result` +(1 run). + +**W3. Mid-Turn pickup behaves as on Linux.** A frame written during a Bash call was held until +the call finished. It then reached the model as a `queued_command` attachment, next to the tool +result, in the same Turn. Its `uuid` was in the one `result`'s `user_message_uuids`, `num_turns` +was 2, and the transcript's `remove` had `reason: "absorbed_mid_turn"` (1 run). A frame written +while text streamed ran as the next native exchange, with its own `system/init` and `result` +(`result_index: 1`) (1 run). `system/init.capabilities` advertised `msg_lifecycle_v1`, and +`command_lifecycle` frames (`queued`, `started`, `completed`, `cancelled`) came on the raw stream +in every run. `queued_turn_count` was `0` on every `result`. + +**W4. An interrupt during a permission-bridge wait behaves as on Linux.** The stand-in bridge +held the `approve` answer for 30 s. After the interrupt, the stand-in got MCP +`notifications/cancelled` for the pending call (`"reason":"AbortError: remote-cancel"`) 3 ms after +the `control_response`, and its handler's abort signal fired. The `result` named the +call in `permission_denials`. The process stayed alive and the next frame ran in the same +Session. The late `allow` had no effect: the script never started (2 of 2 runs, one with +`cancel_queued`). + +**W5. Secant's Windows stop today (`taskkill /pid /T /F`) kills Claude Code outright. +The Bash tool's script still survives.** Claude Code exited with code 1, 29 to 41 ms after +`taskkill` returned, and wrote no `result` and no frame after the kill. Mid-text, nothing of the +partial answer was saved. On `--resume`, Claude Code added a synthetic `No response requested.`, +and the resumed model said it had "declined" the task, as after a Linux SIGTERM. Mid-tool, the +transcript ended at the `tool_use`. `--resume` added a synthetic `tool_result`, "[Tool call +interrupted: the session ended before this call's result was recorded, so its outcome is +unknown…]" (`toolDenialKind: "interrupted"`), plus the synthetic reply. That is the Linux result +for a SIGTERM during a bridge wait, not for a SIGTERM during a running Bash call. `taskkill /T` +listed the Claude Code process, its two tool shells, and their consoles as killed. The script and +its `ping` were not in that list, and the script finished 24 s after the kill (1 run each). + +**W6. Codex on Windows behaves as on Linux.** `turn/interrupt` answered `{}` in 92 ms, and +`turn/completed` `status: "interrupted"` followed 3 ms later. The same thread took the next +`turn/start` (2 of 2). A `turn/steer` with `clientUserMessageId` was accepted at once while text +streamed. It was taken into history only after the whole answer, as a `userMessage` item whose +`clientId` echoed the id, and it was answered in the same Turn (2 of 2). An empty-input +`turn/start` on an idle thread after a completed Turn was accepted. It wrote no `userMessage` +item, and the model answered from history by repeating its last reply (2 of 2). The leftover race +was not attempted. + +## Evidence Vocabulary + +- **Observed**: seen in the live frames, process tables, marker files, or session transcripts + captured for this note on 2026-09-29 against the versions above. Every fact below without + another label is Observed. +- **Source-observed**: read in the Secant source at this branch's base, `docs/interrupt-steer` + (`1174b55`). +- **Inferred**: a consequence drawn from Observed facts that needs a check before it becomes a + compatibility promise. +- **Not observed**: looked for in these runs and absent. + +## Method + +### Driver + +A small Bun 1.4.2 driver spawned each CLI the way `src/process/process.ts` spawns an owned +process on win32: `node:child_process` `spawn` with `overlapped` stdio pipes, +`windowsHide: true`, and `detached: false`, since Windows has no process groups +(**Source-observed**, `spawnOwnedProcessWithNode`, lines 399-414). The driver: + +- wrote stdin frames and logged every stdout line with a millisecond offset from spawn (`t` + below), and logged the exit code and signal; +- read the whole Windows process table (`Get-CimInstance Win32_Process`: pid, parent pid, name, + command line) before and after each stop, and walked Claude Code's descendants by parent pid; +- copied the session transcript from `%USERPROFILE%\.claude\projects\\.jsonl` + after each run. + +Each run used its own throwaway directory under the session scratchpad as cwd. Scripts and raw +logs stayed there, outside the repository. + +### Claude Code launch + +The flags were those of `src/harness/claude-code.ts` (**Source-observed**, lines 764-780): + +```text +claude.exe -p --input-format stream-json --output-format stream-json --verbose \ + --include-partial-messages --model haiku --session-id (or --resume ) +``` + +- `system/init` reported `permissionMode: "default"` and + `capabilities: ["interrupt_receipt_v1","interrupt_cancel_queued_v1","msg_lifecycle_v1","mcp_read_resource_v1","mcp_tool_ui_meta_v1"]` + in every run. The tool list included both `Bash` and `PowerShell`. +- Every `CLAUDE*` and `ORCA_*` variable was removed from the child environment. The host's user + settings still load Orca hook commands (`claude-hook.cmd || echo {}`), a user `CLAUDE.md`, and + the user's MCP servers, as they would under Secant. +- W1 to W3 and W5 pre-allowed only the slow script, with `--allowedTools "Bash(./slowjob.sh)" +"Bash(./slowjob.sh:*)"`. The PowerShell and direct-`ping` variants pre-allowed only their own + command. W4 passed no `--allowedTools`. +- User frames were in the Adapter's shape plus a `uuid`: + `{"type":"user","message":{"role":"user","content":""},"parent_tool_use_id":null,"uuid":""}`. + The interrupt was `{"type":"control_request","request_id":"int1","request":{"subtype":"interrupt"}}`, + with `"cancel_queued":true` added in W2. No `initialize` was ever sent. + +### The slow tool + +`slowjob.sh` waits about 27 s on a native Windows child and leaves marker files, so that its +survival shows as a side effect: + +```bash +#!/bin/bash +echo started-27 +date +%s%3N > started.marker +ping -n 28 127.0.0.1 > /dev/null +echo done-27 +date +%s%3N > done.marker +``` + +The prompt asked for `./slowjob.sh` in the foreground with the Bash tool, then the word +`PINEAPPLE`. Claude Code's Bash tool ran it through Git for Windows: `Git\bin\bash.exe -c "source +…shell-snapshots\snapshot-bash-….sh …"`, then `Git\usr\bin\bash.exe`, then the script's own +`bash.exe`. Two variants checked other tool paths: + +- `slowjob.ps1`, the same steps in PowerShell, run by the PowerShell tool + (`cmd.exe /d /s /c "chcp 65001 & pwsh.exe -NoProfile -NonInteractive … -Command …"`); +- `ping -n 28 127.0.0.1` typed directly as the Bash tool's command. + +In one early W1a run, Haiku sent the command as `./slowjob.sh .`, which the pre-allow rule did not +match. Claude Code denied it ("This command requires approval"), and the interrupt landed while +the model streamed its next step (`terminal_reason: aborted_streaming`). That run counts toward +the receipt timings only. The prompt was then reworded to name the command in backticks with "no +arguments". + +The mid-text prompt was "count from one to three hundred in English words, one number per line". +The stop or send came 2.5 s after the first `text_delta`, or 3 to 5 s after the `tool_use` frame. +After each stop, the same process (or a `--resume` process) was asked, "Without running any +tools: what was the last thing you wrote or did…? Quote the last line you wrote…". + +### Permission-bridge stand-in (W4) + +The stand-in copied `src/harness/permission-bridge.ts` (**Source-observed**): a Streamable HTTP +MCP server on `127.0.0.1` with a random port and a 256-bit bearer token, one transport per MCP +session, built on the repository's `@modelcontextprotocol/sdk` 1.29.0. It served +`secant-permissions` with one `approve` tool (`tool_name`, `input`, `tool_use_id`) returning +`{"behavior":"allow","updatedInput":…}` as JSON text. The launch fragment was +`--mcp-config --permission-prompt-tool mcp__secant-permissions__approve`. Unlike +Secant, it held each answer for 30 s, did not stop when cancelled, and logged every HTTP request +body and the handler's abort signal. + +### Secant's Windows stop (W5) + +Secant's process Module has no graceful stage on Windows. `interrupt` calls `killGroup`, which +spawns `taskkill /pid /T /F` at once and reports `escalated: true` for a live child +(**Source-observed**, `safeInterrupt` lines 561-595, `killGroup` lines 756-776, and +[`src/process/AGENTS.md`](../../src/process/AGENTS.md)). The Claude Code Adapter then settles the +Turn `lost` with `interruption-unknown` (`src/harness/claude-code.ts` lines 545-580), and the +profile's Windows interruption evidence says so (lines 1542-1548). The driver ran the same +`taskkill` command against the Claude Code pid. + +### Codex launch (W6) + +The driver spawned `codex.exe app-server` over JSONL with the handshake of +`src/harness/codex/qualification.ts`: + +```json +{"id":1,"method":"initialize","params":{"clientInfo":{"name":"secant","title":"Secant","version":"0.0.0-dev"},"capabilities":{"experimentalApi":false}}} +{"method":"initialized"} +``` + +- `ORCA_*`, `CLAUDE*`, and the inherited `CODEX_HOME` (which pointed at an Orca runtime home) + were removed. The `initialize` response then reported + `"codexHome":"C:\\Users\\rg\\.codex"`, the user's real home. Nothing in it was changed. +- `model/list` offered `gpt-6-luna` ("Fast and affordable model for easier tasks"), which every + `turn/start` used, with `effort: "low"`. Secant sets no effort. +- `thread/start {cwd}` and `turn/start {threadId, input, model}` followed the shape in + `src/harness/codex.ts`. `turn/steer` carried `threadId`, `expectedTurnId`, `input`, and + `clientUserMessageId`, the ADR 0035 shape. The 0.155.0 stable schema + (`codex app-server generate-json-schema`, without `--experimental`) lists + `clientUserMessageId` on both `TurnSteerParams` and `TurnStartParams`. +- No hook ran in these sessions. No server request arrived. + +### Process hygiene + +Every process started for this note exited. A final process-table check against a snapshot taken +before the first run found no `claude`, `codex`, `bash`, or `ping` left from the runs. The +scripts that outlived Claude Code (W1, W5) ended on their own when their `ping` finished. + +## W1: raw `interrupt` without `initialize` + +### W1a: during a Bash-tool script (`w1-tool2`, run 2) + +```text +t=5772 assistant tool_use Bash {"command":"./slowjob.sh"} +t=7355 started.marker written +t=10483 ps: 7880 claude.exe + └ 18100 Git\bin\bash.exe -c "source …snapshot-bash-….sh …" + └ 5132 Git\usr\bin\bash.exe -c "source …" + 8180 Git\usr\bin\bash.exe ./slowjob.sh (parent 17956: no such process) + └ 2920 Git\usr\bin\bash.exe ./slowjob.sh + └ 9792 PING.EXE -n 28 127.0.0.1 +t=10484 in control_request interrupt int1 +t=10502 control_response {"subtype":"success","request_id":"int1","response":{"still_queued":[]}} +t=10506 system/task_notification status=stopped +t=10526 user tool_result "The user doesn't want to proceed with this tool use. The tool use was + rejected …" is_error +t=10569 result subtype=error_during_execution is_error=true terminal_reason=aborted_tools + stop_reason=tool_use num_turns=3 result_index=0 queued_turn_count=0 +t=12232 ps: 7880 claude.exe + └ 11088 taskkill.exe /PID 18100 /T /F + 8180, 2920, 9792 PING.EXE still running +t=12233 in user +t=16819 assistant text "…I immediately attempted to run the Bash tool with the command + `./slowjob.sh`, but that tool use was rejected…, so the command never executed and I + received no output from it." +t=17017 result subtype=success result_index=1 (same session_id) +t=18803 EXIT code=0 (after the driver closed stdin) +t=29681 ps: 8180, 2920, 9792 PING.EXE still running +t=34803 done.marker written: the script ran to completion +``` + +- **The interrupt is honoured, and the process and Session stay alive.** The receipt came + 18 ms after the write, before the `result`. The next frame ran in the same process and Session. +- **Claude Code stops the tool with `taskkill /T /F` on the outer shell.** The `taskkill.exe` + child of `claude.exe` names pid 18100, the Bash tool's `Git\bin\bash.exe`. +- **That kill does not reach the script.** The script's first `bash.exe` had a parent pid that no + process held, even before the interrupt. The same held for every Bash-tool script whose tree was + captured (four scripts in three runs: the parent pids 17956, 2176, 20276, and 15548 were never + live). `taskkill /T` walks + live parent links only, so it could not find the script's subtree. **Inferred**: Git Bash's + Cygwin-style fork and exec leave the child's Windows parent pointing at a process that has + already exited. +- **The script ran to completion in all 6 runs where the raw interrupt stopped it:** W1a runs 1 + and 2, W1c, W2 runs 1 and 2, and the plain interrupt in W2. It also did in W5a, Secant's own + kill. `done.marker` was written about 27.5 s after `started.marker`, the script's full length. + In every raw-interrupt run except W1a run 1, Claude Code had exited before then. (In W1c the two + scripts share one marker file, and both were seen running after the exit.) +- **What the model is told does not match what happened.** The stored `tool_result` is the + rejection text (`toolUseResult: "User rejected tool use"`, `toolDenialKind: "user-rejected"`), + followed by `[Request interrupted by user for tool use]`. The model said the command never ran. + Its side effects all happened. +- The Linux run's "foreground Bash tree was killed" does not carry over to this Windows tool path. + +### W1a': a PowerShell-tool script and a direct native command (one run each) + +```text +PowerShell tool, ./slowjob.ps1: +t=9441 ps: claude.exe └ cmd.exe /d /s /c "chcp 65001 & pwsh.exe …" └ pwsh.exe -Command … └ PING.EXE +t=9442 in control_request interrupt +t=9461 control_response {"still_queued":[]} +t=9526 result error_during_execution terminal_reason=aborted_tools +t=11185 ps: claude.exe └ taskkill.exe /PID 17180 /T /F (cmd, pwsh, and PING gone) + done.marker absent 32 s after exit + +Bash tool, `ping -n 28 127.0.0.1` typed directly: +t=11969 ps: claude.exe └ Git\bin\bash.exe └ Git\usr\bin\bash.exe └ Git\usr\bin\bash.exe └ PING.EXE +t=11991 control_response {"still_queued":[]} +t=13750 ps: claude.exe └ taskkill.exe /PID 17844 /T /F (bash chain and PING gone) +``` + +- **A tool tree with unbroken parent links is killed.** That covers the PowerShell tool's + `cmd.exe`, `pwsh.exe`, and `ping`, and a native command that the Bash tool starts itself. What + escapes is a process that Git Bash starts as a script interpreter. + +### W1b: while text streams (one run) + +````text +t=10817 stream first text_delta +t=13320 in control_request interrupt int1 +t=13329 control_response {"subtype":"success","request_id":"int1","response":{"still_queued":[]}} +t=13369 assistant text "one\ntwo\n…one hundred forty\none hundred forty-" (1896 chars) aborted=true +t=13373 user text "[Request interrupted by user]" +t=13382 result subtype=error_during_execution terminal_reason=aborted_streaming stop_reason=null + total_cost_usd=0 num_turns=2 +t=14883 alive=true; in user +t=19805 assistant text "The last line I wrote was:\n\n```\none hundred forty-\n```\n\nMy response + was cut off mid-word…" +t=20022 result subtype=success result_index=1 +```` + +The transcript stores the partial text with `isAbortedMidStream: true`, then +`[Request interrupted by user]`. This is the same as Linux test 3b, down to `total_cost_usd: 0` +on the interrupted `result`. + +### W1c: a background task and a foreground call in one Turn (one run) + +The model started `./slowjob.sh` with `run_in_background: true`, then again in the foreground. +The interrupt came 4 s into the foreground call. + +```text +t=5837 system/task_started task_id=but651f8j is_backgrounded=true +t=9558 control_response {"still_queued":[]} +t=9618 result error_during_execution terminal_reason=aborted_tools +t=11864 ps: background shell (Git\bin\bash.exe 20356 └ 132) still a child of claude.exe; + both scripts and both PING.EXE running outside the tree +t=20633 system/task_updated task_id=but651f8j status=killed (after the driver closed stdin) +t=22958 EXIT code=1 +t=25187 ps: both scripts and both PING.EXE still running +``` + +- **A background task survives the interrupt**, as on Linux. Claude Code killed its shell only as + the process exited. +- **On Windows, its script outlived the exit too.** So did the killed foreground call's script. + Both ran to completion. + +### Exit code at stdin close + +After the driver closed stdin, Claude Code exited with code 0 when the last `result` was a +success. It exited with code 1 in the two runs where the last `result` was the interrupted one +(W1a', direct `ping`; W1c). This was not compared on Linux. + +## W2: `cancel_queued: true` with a queued frame + +Two runs. Each wrote a user frame 3 s into the Bash call and interrupted 1.5 s later. Run 1: + +```text +t=6726 in user uuid=eeeeeeee-…-01 "Queued message: reply with the single word MANGO." +t=6730 command_lifecycle eeeeeeee-… state=queued +t=8229 in {"type":"control_request","request_id":"int1","request":{"subtype":"interrupt","cancel_queued":true}} +t=8257 command_lifecycle eeeeeeee-… state=cancelled +t=8259 control_response {"subtype":"success","request_id":"int1", + "response":{"still_queued":[],"cancelled":["eeeeeeee-0000-4000-8000-000000000001"]}} +t=8349 result subtype=error_during_execution terminal_reason=aborted_tools result_index=0 +t=8353 command_lifecycle state=cancelled +t=18351 after 10 s: one result, process alive; in user +t=24246 result subtype=success result_index=1 +``` + +- Run 2 gave the same sequence, with the receipt 28 ms after the write. +- The transcript has `queue-operation` `enqueue`, then `remove`, for the queued frame, and no + `queued_command` attachment. The model never saw it. + +**Plain `interrupt` with a queued frame (one run):** + +```text +t=10119 in control_request interrupt int1 +t=10140 control_response {"still_queued":["dddddddd-0000-4000-8000-000000000001"]} +t=10208 result subtype=error_during_execution terminal_reason=aborted_tools result_index=0 +t=10218 command_lifecycle dddddddd-… state=started +t=11968 assistant text "MANGO" +t=12193 result subtype=success result_index=1 user_message_uuids=["dddddddd-…"] +``` + +The queued frame ran by itself as the next Turn, with no new stdin write, as in Linux test 5a. + +## W3: mid-Turn pickup + +### W3a: one frame during a Bash call (one run) + +```text +t=4061 assistant tool_use Bash {"command":"./slowjob.sh"} +t=7063 in user uuid=aaaaaaaa-…-01 "Additional instruction: after the command finishes, also say the word MANGO." +t=7067 command_lifecycle aaaaaaaa-… state=queued +t=33192 user tool_result "started-27\ndone-27" +t=33386 command_lifecycle aaaaaaaa-… state=started +t=35944 assistant text "PINEAPPLE MANGO" +t=36161 command_lifecycle aaaaaaaa-… state=completed +t=36171 result subtype=success num_turns=2 result_index=0 queued_turn_count=0 + user_message_uuids=[, "aaaaaaaa-0000-4000-8000-000000000001"] +``` + +- The transcript has the `queued_command` attachment (`source_uuid: aaaaaaaa-…`, + `commandMode: "prompt"`) after the tool result, then `queue-operation` `remove` with + `reason: "absorbed_mid_turn"`. +- `started` came 194 ms after the `tool_result` frame. The Linux runs measured 15 to 33 ms. This + is one run. + +### W3b: one frame while text streams (one run) + +```text +t=12969 stream first text_delta +t=15471 in user uuid=bbbbbbbb-…-01 "Now reply with the single word MANGO." +t=15478 command_lifecycle bbbbbbbb-… state=queued +t=21606 assistant text "One\nTwo\n…Three Hundred" (5488 chars) +t=21804 result subtype=success num_turns=1 result_index=0 user_message_uuids=[] +t=21809 command_lifecycle bbbbbbbb-… state=started +t=22004 system/init +t=23314 assistant text "MANGO" +t=23524 result subtype=success result_index=1 user_message_uuids=["bbbbbbbb-…"] +``` + +The frame was not taken into the running Turn. It ran as the next native exchange, with no new +stdin write, as in Linux test 4b. `queued_turn_count` was `0` on the first `result` while the +frame waited. + +## W4: raw `interrupt` while Bash waits on the permission bridge + +Run 1. The stand-in held the Bash `approve` for 30 s. The interrupt came 3.7 s into the wait. + +```text +t=8169 assistant tool_use toolu_01Wr… Bash {"command":"./slowjob.sh",…} +t=8411 HTTP POST /mcp tools/call#2 name=approve +t=8417 MCP approve#1 CALLED tool=Bash tool_use_id=toolu_01Wr… reqId=2 -> will allow after 30000ms +t=12130 in control_request interrupt int1 +t=12136 control_response {"subtype":"success","request_id":"int1","response":{"still_queued":[]}} +t=12139 HTTP POST /mcp notifications/cancelled params={"requestId":2,"reason":"AbortError: remote-cancel"} +t=12149 MCP approve#1 extra.signal ABORTED +t=12152 user tool_result "The user doesn't want to proceed with this tool use. …" is_error +t=12154 user text "[Request interrupted by user for tool use]" +t=12196 result subtype=error_during_execution terminal_reason=aborted_tools result_index=0 + permission_denials=[{"tool_name":"Bash","tool_use_id":"toolu_01Wr…",…}] +t=20092 result subtype=success result_index=1 (the question, same Session) +t=38420 MCP approve#1 RETURNING {"behavior":"allow",…} (signal.aborted=true) +t=53799 started.marker absent, done.marker absent; process alive +t=54877 EXIT code=0 (after the driver closed stdin) +``` + +Run 2 wrote a user frame 1.5 s into the wait and interrupted with `cancel_queued: true` 2.3 s +later. The queued frame got `cancelled`, the receipt listed it in `cancelled`, and +`notifications/cancelled` for the `approve` call followed 3 ms after the receipt. The same +`result` and `permission_denials` followed. No Turn ran until the next stdin frame, and the +script never started. + +- **Same as Linux R2-A1.** The pending approval is cancelled over MCP, and the late `allow` does + nothing. The tool never ran, so the Git Bash survival in W1 cannot arise here. + +## W5: Secant's Windows stop (`taskkill /T /F`), then `--resume` + +### W5a: during a Bash-tool script (one run) + +```text +t=3735 assistant tool_use Bash {"command":"./slowjob.sh"} +t=5279 started.marker written +t=8472 ps: 13320 claude.exe └ 13336 Git\bin\bash.exe └ 7064 Git\usr\bin\bash.exe + 7880 bash.exe ./slowjob.sh (parent 2176: no such process) └ 18776 └ PING.EXE 3352 +t=8473 SIGNAL taskkill /pid 13320 /T /F +t=8690 taskkill exit=0 "SUCCESS: … PID 6052 (child process of PID 7064) … PID 7064 (child + process of PID 13336) … PID 18888 (child process of PID 13320) … PID 13336 (child + process of PID 13320) … PID 13320 (child process of PID 15680) has been terminated." +t=8731 EXIT code=1 signal=null (no result frame, no frame after the kill) +t=10947 ps: 7880, 18776, PING.EXE still running +t=32704 done.marker written: the script ran to completion +``` + +The transcript ended at the `tool_use`. The `--resume` process appended, before its first Turn: + +```text +user tool_result "[Tool call interrupted: the session ended before this call's result + was recorded, so its outcome is unknown. Check whether it took effect before relying + on it or running it again.]" is_error toolDenialKind="interrupted" +assistant "No response requested." model= stop_reason=stop_sequence +user +``` + +The resumed model answered: "The last line I wrote was: "No response requested." The command I +ran (`./slowjob.sh`) did not finish. It was interrupted…". The script had in fact finished. + +- **A forced kill writes nothing.** On Linux, a SIGTERM during a running Bash call let Claude Code + kill the tree and write a real `Exit code 137` result (Linux test 1a). `taskkill /F` gives it no + chance, so the resume sees the same "outcome is unknown" result as after a Linux SIGTERM during + a bridge wait (R2-A2). +- **Secant's `taskkill /T` misses the script for the same reason Claude Code's own kill does.** + The script was not among the processes `taskkill` listed. + +### W5b: while text streams (one run) + +```text +t=4222 stream first text_delta +t=7373 SIGNAL taskkill /pid 2768 /T /F +t=7575 EXIT code=1 signal=null (no result frame) +``` + +The transcript held only the prompt. The `--resume` process added the synthetic +`No response requested.`, and the resumed model answered: "The last line I wrote was "No response +requested." No commands were run — you explicitly asked me to do that task "without using any +tools," so I declined rather than execute it." This is the Linux SIGTERM result (test 1b): the +partial text and its thinking are lost. + +## W6: Codex app-server + +Two runs of the same three-part session, each on fresh threads. + +### W6a: `turn/interrupt` mid-Turn + +```text +[ 7.672] item/agentMessage/delta "one" (86 deltas before the interrupt) +[ 9.174] out turn/interrupt {"threadId":"…0ff8","turnId":"…3532"} +[ 9.266] in response {} +[ 9.270] in turn/completed {"turn":{"id":"…3532","status":"interrupted","error":null,"items":[]}} +[ 9.271] out turn/start {"threadId":"…0ff8","input":[{"type":"text","text":"Without using tools: quote the last line you wrote…"}],…} +[12.593] in turn/completed status=completed + agentMessage "I didn't write a count before your message, so I can't quote a last line. + I didn't finish the count." +``` + +- **Confirmed interrupt, and the thread is reused.** The response came in 92 ms in both runs, + and `turn/completed` `status: "interrupted"` 3 to 4 ms later. The same thread took the next + `turn/start`, which completed. +- **The partial answer was not in context.** No `item/completed` `agentMessage` was emitted for + the interrupted Turn, and in both runs the next Turn's model said it had written no count. + Whether Codex keeps any of an interrupted message in history was not checked with `thread/read`. + +### W6b: `turn/steer` with `clientUserMessageId` while text streams + +```text +[17.068] item/started agentMessage +[18.072] out turn/steer {"threadId":"…40c9","expectedTurnId":"…94be","input":[{"type":"text", + "text":"New instruction: stop counting now and reply with just the word MANGO."}], + "clientUserMessageId":"secant-steer-r1"} +[18.081] in response {"turnId":"…94be"} +[40.667] item/completed agentMessage (5488 chars, "one … three hundred") +[40.754] item/started userMessage clientId="secant-steer-r1" "New instruction: stop counting now…" +[40.779] item/completed userMessage clientId="secant-steer-r1" +[42.204] item/completed agentMessage "MANGO" +[42.299] turn/completed {"turn":{"id":"…94be","status":"completed",…}} +``` + +`thread/read` (`includeTurns: true`) shows one Turn holding the prompt, the full count, the +steered `userMessage` with `clientId: "secant-steer-r1"`, and `MANGO`. Run 2 matched. + +- **The steer is accepted at once but taken only at the next model-request boundary.** Here that + was after the whole streamed answer. It is answered in the same Turn, and the `clientId` + echoes the Secant-minted id. +- 0.155.0 also printed a `deprecationNotice` for `thread/read` with `includeTurns: true` on a + paginated thread, pointing at `thread/turns/list` and `thread/items/list`. + +### W6c: empty-input `turn/start` on an idle thread + +```text +T1 turn/start "Remember the code word ZEBRA-42. Reply with just: OK" -> agentMessage "OK", completed +T2 turn/start {"threadId":"…","input":[],"model":"gpt-6-luna","effort":"low"} + -> response {"turn":{"id":"…","status":"inProgress",…}} + -> turn/started, agentMessage "OK", turn/completed status=completed + (no userMessage item in T2) +T3 turn/start "What code word did I ask you to remember?" -> agentMessage "ZEBRA-42" +``` + +The same held in both runs. As on Linux, with no new input the model answered the last user +message in history again. The leftover race (a steer landing after Codex's last pending-input +check) was not attempted. + +## Windows Compared With Linux + +| Case | Linux (Claude Code 2.1.284, codex-cli 0.157.1) | Windows (Claude Code 2.1.283, codex-cli 0.155.0) | +| --------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| Raw `interrupt` without `initialize`: receipt | `control_response` success in 2 to 4 ms, before the `result` | Same, in 6 to 30 ms (12 of 12) | +| Raw `interrupt`: Turn end | `error_during_execution`, `aborted_streaming` / `aborted_tools` | Same | +| Raw `interrupt`: process and Session | Alive; next frame runs in the same Session | Same | +| Raw `interrupt` mid-text: partial text | Kept (`aborted: true`); model quotes the last partial line | Same | +| Raw `interrupt` mid-tool: what the model sees | "rejected" `tool_result`; partial stdout dropped | Same | +| Raw `interrupt` mid-tool: foreground tool processes | Bash tree killed | Claude Code runs `taskkill /PID /T /F`. PowerShell tool and a direct native command: killed. A script run by the Bash tool (Git Bash): **not killed, ran to completion** (6 of 6) | +| Background Bash task | Survives the interrupt; killed at process exit | Survives the interrupt; its shell is killed at exit, but **its script outlives the exit** | +| Plain `interrupt` with a queued frame | Listed in `still_queued`; runs next by itself | Same | +| `interrupt` with `cancel_queued: true` | Listed in `cancelled`; lifecycle `cancelled`; no next Turn; alive | Same (2 of 2) | +| Frame during a tool call | Taken after the round's last result, same Turn, in `user_message_uuids` | Same (1 run) | +| Frame while text streams | Runs as the next native exchange, own `result` | Same (1 run) | +| `command_lifecycle` / `msg_lifecycle_v1` | Present by default on the raw stream | Same | +| `queued_turn_count` | Always `0` | Always `0` | +| `interrupt` during a bridge wait | `notifications/cancelled` within 4 ms; `permission_denials` names the call; late `allow` ignored; tool never runs | Same (2 of 2) | +| Today's stop, mid-text | SIGTERM: exit 143, no `result`, partial text lost; resume adds `No response requested.`; model says it "declined" | `taskkill /T /F`: exit 1, no `result`, partial text lost; same resume and same "declined" answer | +| Today's stop, mid-tool | SIGTERM: Claude Code kills the tree and records `Exit code 137` with partial stdout; resume shows it | `taskkill /T /F`: nothing recorded after the `tool_use`; resume adds "Tool call interrupted… outcome is unknown"; the Git Bash script survives and completes | +| Codex `turn/interrupt` | Not re-run in the leftover note; ADR 0035 relies on `turn/completed` `status: "interrupted"` | Response `{}` in 92 ms, then `turn/completed` `status: "interrupted"`; thread reused (2 of 2) | +| Codex `turn/steer` with `clientUserMessageId` | Accepted; `userMessage` item carries the `clientId` | Same; taken after the streamed answer, answered in the same Turn (2 of 2) | +| Codex empty-input `turn/start` on an idle thread | Accepted; no `userMessage` item; model answers from history | Same (2 of 2) | +| Codex leftover steer race | 6 of 22 at 0 ms after the Stop hook's `hook/completed` | Not attempted | + +## Against ADR 0035 + +These Observed results differ from what ADR 0035 states. The note makes no recommendation. + +- **"kills the foreground tool tree."** On Windows, the raw `interrupt` did not kill a script that + Claude Code's Bash tool ran through Git Bash. The script ran to completion in 6 of 6 runs, while + the model was told the call was rejected and said it never ran. It did kill the PowerShell + tool's tree and a native command that the Bash tool started directly. +- **"a shell the model moved to the background may outlive it until the Session closes."** On + Windows, a background Bash task's script outlived the Session's process as well. +- **The SIGTERM fallback.** On Windows, today's fallback is Secant's `taskkill /T /F`, not SIGTERM. + It also misses the Git Bash script. After it, the resumed model sees "outcome is unknown" rather + than a killed tool's `Exit code 137`. The ADR's statement that the fallback loses partial + streamed text holds on Windows. + +Everything else the ADR relies on held on Windows. The raw `interrupt` answered without +`initialize`, ended the Turn with a `result`, kept the process, Session, and partial text, and +cancelled a pending bridge approval with `notifications/cancelled`. `cancel_queued` handed the +queued frame back. Pickup at the tool round's end and next-exchange delivery during text both +held, as did `command_lifecycle` frames. For Codex, `turn/interrupt`, `turn/steer` with +`clientUserMessageId`, and an empty `turn/start` all held. + +## Still Unknown + +- Whether Claude Code 2.1.284, the Linux version, behaves differently on Windows. Only 2.1.283 was + installed here. +- Why the Git Bash script's Windows parent pid is dead. The Cygwin fork and exec explanation is + **Inferred**. It was not traced, and neither was whether a `CLAUDE_CODE_GIT_BASH_PATH` or + other shell setting changes it. +- Which other Bash-tool commands escape the tree: a pipeline, `bash -c`, `npm`, `bun`, or a + command that runs a `.sh` through `sh`. Only a script run as `./slowjob.sh` escaped, and only + one direct native command was tried. +- Whether a Windows Job Object or a later Claude Code release would make `taskkill /T` reach + these processes. +- SIGINT, or `GenerateConsoleCtrlEvent`, to a hidden Windows Claude Code child. Neither was + tried, since ADR 0035 uses neither. +- Whether a user frame is taken between calls in a multi-call round on Windows, and whether the + denial-boundary timing of Linux round 3 holds. Neither was repeated. +- Why W3a's `started` came 194 ms after the `tool_result`, against 15 to 33 ms on Linux, and + whether that widens the window in which a frame written near a boundary misses the Turn. +- The Codex leftover race on Windows, and whether its hit rate differs from Linux. +- Whether Codex keeps any of an interrupted Turn's partial agent message in history. The next + Turn's model said it had written nothing, in 2 of 2 runs. +- Whether Codex's `turn/interrupt` stops a running shell command's process tree on Windows. W6 + interrupted only streamed text. diff --git a/src/headless/AGENTS.md b/src/headless/AGENTS.md index f2641488..7e568445 100644 --- a/src/headless/AGENTS.md +++ b/src/headless/AGENTS.md @@ -69,3 +69,7 @@ Inherits the engineering baseline; records only non-obvious local facts. Ownersh `Bun.main` separator normalisation, the #62 Windows entry quirk (`:48-53`); and the top-level error handler that prints a stack and sets exit 1 (`:55-62`). The file records at `:40-47` why it cannot be unit-tested as written; if a fourth branch appears, revisit a small `tests/cli` rather than widening the smoke. + +## Read next + +- [Headless parity](../../docs/headless-parity.md) lists the TUI capabilities headless deliberately lacks; update it when a decision adds or closes one.