From d699cdaaad94b12edd76890d95bb47886e98f0bd Mon Sep 17 00:00:00 2001 From: rohan Date: Tue, 29 Sep 2026 10:30:55 +0530 Subject: [PATCH 01/11] docs(research): record live Claude Code interrupt and mid-Turn send on 2.1.284 (#255) --- ...e-code-live-interrupt-and-mid-turn-send.md | 564 ++++++++++++++++++ 1 file changed, 564 insertions(+) create mode 100644 docs/research/claude-code-live-interrupt-and-mid-turn-send.md diff --git a/docs/research/claude-code-live-interrupt-and-mid-turn-send.md b/docs/research/claude-code-live-interrupt-and-mid-turn-send.md new file mode 100644 index 0000000..c85f058 --- /dev/null +++ b/docs/research/claude-code-live-interrupt-and-mid-turn-send.md @@ -0,0 +1,564 @@ +# Claude Code Live Interrupt and Mid-Turn Send + +Research date: 2026-09-29 + +Harness version examined: Claude Code **2.1.284** (installed native executable, Linux +x64; `claude --version` returned `2.1.284 (Claude Code)`). + +Ticket: [#255](https://github.com/secantdev/secant/issues/255). This note settles +the Claude Code items left **Untested** or **Unknown** in +[Harness Interrupt and Queued Messages](harness-interrupt-and-queued-messages.md) +by running small live model sessions against the installed CLI. + +## Answer + +**SIGTERM loses streamed text but keeps a killed tool call.** A process-group SIGTERM +exits with code 143 and writes no `result`. Mid-text, the transcript keeps only the +user prompt: no partial text, no thinking, no interrupted marker. Mid-tool, the +transcript keeps the `tool_use` and a real `tool_result` of `Exit code 137` plus +the tool's partial stdout. On `--resume`, Claude Code first appends a synthetic +assistant message, `No response requested.` (`model: ""`). The resumed +model saw the killed tool call and its output. After a mid-text SIGTERM it saw +only its own prompt and the synthetic reply. Twice it said it had "declined" the +task. + +**SIGINT ends the Turn, then the process exits.** SIGINT gave one `result` +(`subtype: "error_during_execution"`, `is_error: true`, `terminal_reason` +`aborted_streaming` or `aborted_tools`). The process then exited with code 0 about +1 s later, with stdin still open. This held for the process group and for the pid +alone. The partial Turn is kept in the transcript: partial text is saved with +`isAbortedMidStream: true`, and `[Request interrupted by user]` is appended. A +`--resume` process sees both, after the synthetic `No response requested.`. A user +message still queued at the SIGINT was dropped and did not come back on resume. + +**A raw `control_request` `interrupt` works without `initialize` and keeps the +process alive.** The `control_response` came back in 2 to 4 ms and carried the +receipt `{"still_queued":[…]}`. It came before the Turn's `result`, which had the +same shape as the SIGINT one. The next stdin user frame then ran as a new Turn in +the same process and Session. + +- Mid-text, the partial text stays in context. The `assistant` frame carries + `aborted: true`, and the model later quoted the last line it had written. +- Mid-tool, the `tool_use` stays. Its `tool_result` is replaced by the standard + "The user doesn't want to proceed with this tool use… rejected" text, followed by + `[Request interrupted by user for tool use]`. Output the tool had already printed + is thrown away. The model then said the tool was "rejected before execution". +- The foreground Bash process tree was killed. +- A Bash command the model had moved to the background kept running after the + interrupt. It was killed only when the process exited. + +**A user frame written during a tool call is picked up in the same Turn. One written +while text streams becomes the next Turn.** + +- **During a tool call.** The frame was held until the tool finished. It then + reached the model as a `queued_command` system-reminder next to the tool result. + The Turn's one `result` listed its `uuid` in `user_message_uuids`, and + `num_turns` stayed 2. The model acted on it in that Turn. +- **Two frames** written 0.5 s apart were taken together at the same boundary, in + write order. +- **While text streamed.** The frame was not taken into the running Turn. It ran + as a second Turn with its own `result` and `result_index: 1`. +- `queued_turn_count` was `0` on every `result`, including the first `result` in + that last case, while the message was still waiting. + +**On an interrupt, a queued message runs, is cancelled, or is lost, depending on the +stop.** + +- **Plain `interrupt`.** The message is listed in `still_queued` and runs by itself + as the next Turn right after the interrupted `result`. +- **`interrupt` with `cancel_queued: true`.** The message is listed in `cancelled`, + and a `command_lifecycle` `cancelled` frame is written for it. No next Turn + starts, and the process stays alive. +- **SIGINT.** The message is dropped with no lifecycle frame, and the process + exits. + +**Version floor.** The changelog names none of `user_message_uuids`, +`queued_turn_count`, the `interrupt` control request, receipts, or `cancel_queued`. +The nearest entries are listed under [Version floor](#test-6-version-floor). + +**Which stop keeps the process alive and the partial Turn in context:** only the +`control_request` `interrupt`. SIGINT keeps the partial Turn but ends the process. +SIGTERM ends the process and loses streamed text. + +Windows was not tested. + +## Evidence Vocabulary + +- **Recorded (2.1.284)**: seen in the live frame sequences or session transcripts + captured for this note on 2026-09-29 against Claude Code 2.1.284. Every fact + below without another label is Recorded (2.1.284). +- **Documented**: stated in the Claude Code changelog or official docs, cited + directly or through + [Harness Interrupt and Queued Messages](harness-interrupt-and-queued-messages.md). +- **Inferred**: a consequence drawn from Recorded facts; it needs a check before it + becomes a compatibility promise. +- **Unknown**: not settled by these runs. + +## Method + +A small Bun driver spawned `claude` with the stream flags that +`src/harness/claude-code.ts` uses. It spawned `claude` `detached`, so the CLI led its +own process group, as `src/process/process.ts` does: + +```text +claude -p --input-format stream-json --output-format stream-json --verbose \ + --include-partial-messages --model haiku \ + [--session-id | --resume ] \ + --allowedTools "Bash(sleep:*)" "Bash(echo:*)" "Bash(./slowjob.sh)" +``` + +The driver worked like this: + +- It wrote user frames in the Adapter's `encodeTurn` shape plus a `uuid`: + `{"type":"user","message":{"role":"user","content":""},"parent_tool_use_id":null,"uuid":""}`. +- It logged every stdout line with a millisecond offset, and recorded exit code + and signal. +- It signalled with `process.kill(-pid, sig)` for the process group, as Secant's + `killGroup` does, or with `process.kill(pid, sig)` for the pid alone. +- It checked the Bash child with `ps -eo pid,ppid,pgid,sid,args` before and + after each stop. +- It read the session transcript at + `~/.claude/projects//.jsonl` after each run. +- Each process ran with a scratch directory as cwd, not the repository. It had + no `CLAUDE*` environment variables, used subscription auth + (`apiKeySource: "none"`), and ran in permission mode `default`. + +The two Turn shapes: + +- **Mid-tool.** The prompt asked for `./slowjob.sh` in the foreground, then the + code word `PINEAPPLE`. `slowjob.sh` is `echo started-27; sleep 27; echo done-27`. + The stop or send came 3 s after the `tool_use` frame. The brief's `sleep 20` could + not be used: 2.1.284's Bash tool refuses long sleeps on its own. It answered + `sleep 27` with `Blocked: standalone sleep 27…` and `sleep 27 && echo done-27` + with `Blocked: sleep 27 followed by: echo done-27…`. The first `sleep 27` run + (test 3, run 1) was moved to the background by the model. +- **Mid-text.** The prompt was "count from one to three hundred in English words, + one number per line". The stop or send came 2.5 s after the first text delta, + about 1,400 characters in. + +After each stop the driver asked the same process, or a `--resume` process, +"Without running any tools: what was the last thing you wrote or did in this +conversation before this message? Quote the last line you wrote…". + +Differences from Secant's launch: + +- Secant adds the loopback MCP permission bridge (`--mcp-config`, + `--permission-prompt-tool`). The driver did not. A tool call parked on a + permission prompt at the moment of a stop was not tested. +- The host's user settings load no-op `SessionStart`, `PreToolUse`, `PostToolUse`, + and `Stop` hooks, plus a user `CLAUDE.md`. + +Frame excerpts below are trimmed. `t` is milliseconds since spawn. Hook, status, +`thinking_tokens`, and `stream_event` frames are omitted, and so are most +`command_lifecycle` frames. Raw logs were kept outside the repository. + +## Test 1: SIGTERM mid-Turn, then `--resume` + +### 1a: during a tool call (three runs) + +```text +t=2916 assistant tool_use {"command":"./slowjob.sh","run_in_background":false} + ps: 68103 claude (pgid 68103) + 68275 /bin/bash -c … eval ./slowjob.sh (pgid 68275, sid 68275) + 68277 /bin/bash ./slowjob.sh (pgid 68277) + 68278 sleep 27 (pgid 68277) +t=5986 SIGNAL SIGTERM group pid=68103 +t=6000 system/task_notification status=stopped +t=6021 user tool_result "Exit code 137\nstarted-27" is_error +t=6865 EXIT code=143 signal=null +t=8427 ps: no slowjob or sleep 27 left +``` + +No `result` frame was written. The transcript kept the whole tool round, and resume +added one synthetic entry: + +```text +user "Use the Bash tool … ./slowjob.sh …" +assistant thinking, tool_use toolu_01… stop_reason=tool_use +user tool_result "Exit code 137\nstarted-27" is_error + toolUseResult="Error: Exit code 137\nstarted-27" +--- written by the --resume process --- +assistant "No response requested." model= stop_reason=stop_sequence +user +``` + +The resumed model answered: "I ran the Bash command `./slowjob.sh` in the +foreground, but it was killed before completion—the exit code was 137 (SIGKILL) and +it only printed "started-27" before terminating." + +- **The Bash tree was killed, although it was outside the signalled group.** The + Bash tool runs its shell as a new session and process group (`pgid 68275`), so + `kill(-pid)` does not reach it. Claude Code killed it itself (exit 137). This + matches the documented 2.1.212 fix for orphaned trees on SIGTERM. +- The killed tool call is kept as an ordinary error result. There is no + `[Request interrupted…]` marker and no synthetic denial. This matches 2.1.236. + +### 1b: while text streams (two runs) + +```text +t=5139 partial text so far (1372 chars) tail="…One Hundred Eighteen\nOne" +t=5140 SIGNAL SIGTERM group +t=6401 EXIT code=143 signal=null +``` + +No `result` frame and no `assistant` frame were written. The transcript held only +the user prompt, then the resume's synthetic `No response requested.`. The resumed +model answered, in both runs: "The last line I wrote was: "No response requested." +No commands ran—you asked me to count without using tools, which I declined to do." + +- **The partial text and its thinking are gone.** The resumed model has no trace + that it started to answer. +- The synthetic `No response requested.` reads, to the model, as a refusal of the + prompt. **Inferred.** + +## Test 2: SIGINT to a `-p` stream-json process + +### 2a: during a tool call (process group, then pid only) + +```text +t=3282 assistant tool_use {"command":"./slowjob.sh",…} +t=6325 SIGNAL SIGINT group +t=6382 user tool_result "The user doesn't want to proceed with this tool use. The tool + use was rejected (eg. if it was a file edit, the new_string was NOT written to + the file). STOP what you are doing and wait for the user t…" is_error +t=6385 user text "[Request interrupted by user for tool use]" +t=6401 result subtype=error_during_execution is_error=true terminal_reason=aborted_tools + stop_reason=tool_use num_turns=3 user_message_uuids=[] + errors=["[ede_diagnostic] result_type=user last_content_type=n/a stop_reason=tool_use"] +t=7255 EXIT code=0 signal=null +``` + +SIGINT to the pid alone (2c) gave the same frames and `EXIT code=0` after 1 s. + +### 2b: while text streams (two runs) + +```text +t=5239 SIGNAL SIGINT group +t=5277 assistant text "one\ntwo\n…" (1822 chars) aborted=true +t=5278 user text "[Request interrupted by user]" +t=5283 result subtype=error_during_execution terminal_reason=aborted_streaming + stop_reason=null num_turns=2 total_cost_usd=0 duration_api_ms=0 +t=6564 EXIT code=0 signal=null +``` + +Transcript, then the `--resume` process: + +```text +assistant thinking stop_reason=null +assistant text "one\ntwo\n…" (1822 chars) isAbortedMidStream=true +user "[Request interrupted by user]" +--- written by the --resume process --- +assistant "No response requested." model= +user +``` + +The resumed model answered: "The last line I wrote was "one hundred" as part of +counting from one to three hundred in words. You interrupted the counting task +mid-way through the one-hundreds range." + +- **SIGINT ends the Turn and then the process.** The process did not wait for more + stdin. It exited 0 about 1 s after the `result` in all four runs, so no next + user frame can run in the same process. +- **The partial Turn is kept.** On resume, the model sees the partial text, the + interrupt marker, and the synthetic reply. For a tool call, it sees the rejected + result. The second mid-text run's resumed model quoted the synthetic line as its + last line but still said "you interrupted the counting task". +- The Bash tree was gone after the SIGINT. +- The interrupted mid-text `result` reported `total_cost_usd: 0` and zero usage, + even though about 1,400 characters had been generated. The next `result` in the + process (test 3b) carried the running total again. Whether an aborted stream's + tokens are ever counted is Unknown. + +## Test 3: raw `control_request` `interrupt` without `initialize` + +The first frame of each process was the user prompt. No `initialize` was ever sent. + +### 3a: during a tool call + +```text +t=2732 assistant tool_use {"command":"./slowjob.sh",…} +t=5774 in {"type":"control_request","request_id":"int1","request":{"subtype":"interrupt"}} +t=5778 control_response {"subtype":"success","request_id":"int1","response":{"still_queued":[]}} +t=5792 user tool_result "The user doesn't want to proceed with this tool use. The tool use was + rejected …" is_error +t=5793 user text "[Request interrupted by user for tool use]" +t=5800 result subtype=error_during_execution is_error=true terminal_reason=aborted_tools + num_turns=3 user_message_uuids=[] queued_turn_count=0 result_index=0 +t=5800 command_lifecycle state=cancelled +t=7343 ps: no slowjob or sleep 27 left; process alive +t=7343 in user +t=7403 system/init (same session_id) +t=11845 assistant text "I attempted to call the Bash tool to run `./slowjob.sh`, but the + tool use was rejected before execution. No command finished and I received no + output—the user interrupted the tool call." +t=11891 result subtype=success num_turns=1 result_index=1 +``` + +In the transcript, the `tool_result` is stored with `toolUseResult: "User rejected +tool use"` and `toolDenialKind: "user-rejected"`. In tests 5a and 5b, `slowjob.sh` +had already printed `started-27` when the interrupt came, and the stored +`tool_result` was the same rejection text. The partial stdout was discarded. No +synthetic `No response requested.` is written when the next Turn runs in the same +process. + +Run 1 used a bare `sleep 27` and hit the Bash tool's sleep guard. The model then +started `sleep 27` with `run_in_background: true`. The interrupt at t=5612 ended the +Turn, but `sleep 27` was still running at t=7154. The background task was reported +`killed` only at t=15993, after the driver closed stdin and the process began to +exit. The second Turn's model said the background command "is still running". + +### 3b: while text streams + +```text +t=5444 in control_request interrupt int1 +t=5446 control_response {"subtype":"success","request_id":"int1","response":{"still_queued":[]}} +t=5481 assistant text "One\nTwo\n…" (1412 chars) aborted=true +t=5484 user text "[Request interrupted by user]" +t=5488 result subtype=error_during_execution terminal_reason=aborted_streaming stop_reason=null + total_cost_usd=0 num_turns=2 result_index=0 +t=7027 in user +t=10714 assistant text "The last line I wrote was "One" (after "One Hundred Twenty"), + completing the number list you requested. … I was interrupted partway through + the three-hundred count." +t=10775 result subtype=success num_turns=1 result_index=1 +t=11673 EXIT code=0 (after the driver closed stdin) +``` + +- **Honoured without `initialize`.** The `control_response` is a `success` with the + documented receipt, written before the interrupted `result`, as the SDK typings + describe. +- **The process stays alive and the Session continues.** The next user frame ran + in the same process with the same `session_id`. +- **Partial text is in context. A partial tool result is not.** The model quoted + the exact last partial line. For the tool call, it saw only the rejection text, + and it believed the command never ran. +- **The foreground Bash tree is killed. A background Bash task is not.** +- The `init` frame advertised `capabilities: ["interrupt_receipt_v1", +"interrupt_cancel_queued_v1", "msg_lifecycle_v1", "mcp_read_resource_v1", +"mcp_tool_ui_meta_v1"]`. `command_lifecycle` frames (`queued`, `started`, + `completed`, `cancelled`) were written by default on this raw stream, with no + opt-in. + +## Test 4: user frames written mid-Turn + +### 4a: one frame during a tool call + +```text +t=2978 assistant tool_use {"command":"./slowjob.sh",…} +t=6018 in user uuid=aaaaaaaa-…-000000000001 "Additional instruction: after the command + finishes, also say the word MANGO." +t=6022 command_lifecycle aaaaaaaa-… state=queued +t=30101 user tool_result "started-27\ndone-27" +t=30139 command_lifecycle aaaaaaaa-… state=started +t=34711 assistant text "PINEAPPLE\nMANGO" +t=34741 command_lifecycle aaaaaaaa-… state=completed +t=34746 result subtype=success num_turns=2 queued_turn_count=0 result_index=0 + user_message_uuid= + user_message_uuids=[, "aaaaaaaa-0000-4000-8000-000000000001"] +t=34748 command_lifecycle state=completed +``` + +The transcript shows how the frame reached the model: + +```text +queue-operation enqueue "Additional instruction: …" +user tool_result "started-27\ndone-27" +attachment queued_command source_uuid=aaaaaaaa-… commandMode=prompt + rendered: "\nThe user sent a new message while you were + working:\nAdditional instruction: after the command finishes, also say + the word MANGO.\n\nThis is how Claude Code surfaces messages the user + sends mid-turn — within the running turn, often alongside the next tool + result, rather than as a separate conversation turn. Address the message + above as you continue this turn.\n" +queue-operation remove reason=absorbed_mid_turn +assistant "PINEAPPLE\nMANGO" +``` + +- **Picked up in the same Turn.** The frame is held until the running tool + finishes, not injected into it, and it did not cut the 27 s command short. It + reached the model with the tool result. +- The Turn's one `result` lists the picked-up `uuid` in `user_message_uuids`. + `user_message_uuid` stays the prompt's. `num_turns` is 2, the same as a Turn with + no pickup. +- The picked-up frame is never echoed on stdout as a `user` frame; the driver did + not pass `--replay-user-messages`. `command_lifecycle` `completed` for it comes + before the Turn's `result`. + +### 4b: one frame while text streams, with no tool round left + +```text +t=5584 in user uuid=bbbbbbbb-…-000000000001 "Now reply with the single word MANGO." +t=5585 command_lifecycle bbbbbbbb-… state=queued +t=10247 assistant text "One\nTwo\n…" (5488 chars) +t=10290 result subtype=success num_turns=1 queued_turn_count=0 result_index=0 + user_message_uuids=[] +t=10293 command_lifecycle state=completed +t=10294 command_lifecycle bbbbbbbb-… state=started +t=10333 system/init +t=11500 assistant text "MANGO" +t=11542 result subtype=success num_turns=1 queued_turn_count=0 result_index=1 + user_message_uuid="bbbbbbbb-…" user_message_uuids=["bbbbbbbb-…"] +``` + +- **Not picked up. It runs as the next Turn.** The first Turn finished its text + with no sign of the message. A second `system/init` and a second `result` followed + at once, with no new stdin write. +- **`queued_turn_count` was `0`** on the first `result`, while `bbbbbbbb-…` was + still queued and about to run. It was `0` on every `result` captured for this + note. It cannot be used to tell that another Turn will follow. The + `command_lifecycle` `queued` frame, without a matching `started` before the + `result`, did show this. + +### 4c: two frames during a tool call + +```text +t=5767 in user uuid=cccccccc-…-01 "Additional instruction one: … say the word MANGO." +t=6269 in user uuid=cccccccc-…-02 "Additional instruction two: after that, say the word KIWI." +t=29830 user tool_result "started-27\ndone-27" +t=29890 command_lifecycle cccccccc-…-01 state=started +t=29890 command_lifecycle cccccccc-…-02 state=started +t=33241 assistant text "PINEAPPLE" +t=33277 result subtype=success num_turns=2 queued_turn_count=0 + user_message_uuids=[, "cccccccc-…-01", "cccccccc-…-02"] +``` + +- **Both frames are taken at the same boundary, in write order.** The transcript + has two `queued_command` attachments (one, then two) after the tool result. The + Turn made one more model call, not one per message. +- **Listed does not mean acted on.** Both `uuid`s are in `user_message_uuids`, but + the model answered only `PINEAPPLE`. The original prompt said "and nothing else", + and Haiku kept to it. `user_message_uuids` shows delivery to the model, not + compliance. + +## Test 5: a queued message, then a stop + +Each run wrote a user frame during the tool call and stopped the Turn 1.5 s later, +before the tool finished. + +### 5a: plain `interrupt` + +```text +t=5955 in user uuid=dddddddd-…-01 "Queued message: reply with the single word MANGO." +t=7456 in control_request interrupt int1 +t=7460 control_response {"subtype":"success","request_id":"int1", + "response":{"still_queued":["dddddddd-0000-4000-8000-000000000001"]}} +t=7472 user tool_result "The user doesn't want to proceed with this tool use. …" is_error +t=7480 result subtype=error_during_execution terminal_reason=aborted_tools result_index=0 + user_message_uuids=[] +t=7484 command_lifecycle dddddddd-… state=started +t=9103 assistant text "MANGO" +t=9136 result subtype=success num_turns=1 result_index=1 user_message_uuids=["dddddddd-…"] +``` + +- **The queued message runs by itself as the next Turn**, 4 ms after the + interrupted `result`, with no new stdin write. The receipt names it in + `still_queued`. + +### 5b: `interrupt` with `cancel_queued: true` + +```text +t=5728 in user uuid=eeeeeeee-…-01 "Queued message: reply with the single word MANGO." +t=7229 in {"type":"control_request","request_id":"int1", + "request":{"subtype":"interrupt","cancel_queued":true}} +t=7234 command_lifecycle eeeeeeee-… state=cancelled +t=7235 control_response {"subtype":"success","request_id":"int1", + "response":{"still_queued":[],"cancelled":["eeeeeeee-0000-4000-8000-000000000001"]}} +t=7262 result subtype=error_during_execution terminal_reason=aborted_tools result_index=0 +t=18827 process alive, no further frames; in user +t=23444 result subtype=success result_index=1 +``` + +- **`cancel_queued` works on the raw stream.** The message is dropped, and the + receipt names it in `cancelled`. It never reaches the model: the transcript has + `queue-operation` `enqueue` then `remove`, and no `queued_command` attachment. No + Turn starts until the next stdin frame. + +### 5c: SIGINT + +```text +t=6501 in user uuid=ffffffff-…-01 "Queued message: reply with the single word MANGO." +t=6502 command_lifecycle ffffffff-… state=queued +t=8001 SIGNAL SIGINT group +t=8050 result subtype=error_during_execution terminal_reason=aborted_tools +t=8052 command_lifecycle state=cancelled +t=8961 EXIT code=0 signal=null +``` + +- **The queued message is lost.** No lifecycle frame is written for + `ffffffff-…`, and the process exits. The transcript holds a `queue-operation` + `enqueue` for it with no `dequeue` or `remove`. A `--resume` of the Session did + not replay it. Asked to list every user message, the resumed model listed the + prompt, `[Request interrupted by user for tool use]`, the synthetic `No response +requested.`, and the question. MANGO was not among them. + +## Test 6: version floor + +Read from the [Claude Code CHANGELOG](https://github.com/anthropics/claude-code/blob/main/CHANGELOG.md) +at its 2.1.284 head (**Documented**). No entry names `user_message_uuids`, +`user_message_uuid`, `queued_turn_count`, the `interrupt` control request, +`interrupt_receipt`, `still_queued`, `cancel_queued`, `command_lifecycle`, or a +`priority` field on user messages. The nearest entries: + +| Version | Entry (quoted or trimmed) | +| ------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| 2.1.19 | "[SDK] Added replay of `queued_command` attachment messages as `SDKUserMessageReplay` events when `replayUserMessages` is enabled" | +| 2.1.94 | "Fixed SDK/print mode not preserving the partial assistant response in conversation history when interrupted mid-stream" | +| 2.1.162 | "Fixed an interrupt (Esc) sent at the very start of a turn being silently dropped in stream-json/SDK sessions" | +| 2.1.212 | "Fixed SIGTERM during a running Bash tool orphaning the command's process tree in print/SDK mode; the CLI now aborts the turn, kills the tree, and exits 143" | +| 2.1.236 | "SIGTERM in print/SDK mode no longer records an interrupted turn or synthetic tool denials before exiting; running commands are still terminated and the process still exits with code 143" | +| 2.1.246 | "Fixed MCP tool calls interrupted by an incoming message in headless/remote sessions being reported to the model as "completed with no output" instead of an explicit interrupted error" | +| 2.1.261 | "Fixed SDK and cloud sessions ignoring a Stop or interrupt sent just after the first prompt, before the turn had started" | +| 2.1.274 | "Fixed a local `claude -p --resume` started with `CLAUDE_CODE_RESUME_INTERRUPTED_TURN` not reporting background tasks the previous process left unfinished" | + +The SDK reference's own version notes still apply: `interrupt_receipt_v1` from +v2.1.205 and `interrupt_cancel_queued_v1` from v2.1.219 (**Documented**, cited in +[Harness Interrupt and Queued Messages](harness-interrupt-and-queued-messages.md)). +The installed 2.1.284 advertises both, and also `msg_lifecycle_v1`, in +`system/init.capabilities`. + +## Test 7: Windows + +Not tested. No Windows host was available. Signal delivery, `taskkill` tree +behaviour, and the SIGINT and SIGTERM results above are all Linux-only. + +## The Four Original Unknowns + +| Unknown in [Harness Interrupt and Queued Messages](harness-interrupt-and-queued-messages.md) | Result on 2.1.284 | Evidence | +| -------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ------------------------------------------- | +| Does SIGTERM in `-p` mode keep the partial assistant text in the transcript that `--resume` loads? | **No for streamed text:** nothing of the unfinished model call is saved, not even thinking. **Yes for a finished tool round:** the `tool_use` and a real `Exit code 137` result with partial stdout are saved. There is no interrupted marker. Resume adds a synthetic `No response requested.`. | Recorded (2.1.284) | +| Does SIGINT to a `-p --input-format stream-json` process end only the Turn and keep reading stdin? | **No.** It writes one `error_during_execution` `result`, then exits with code 0 about 1 s later, with stdin still open. The partial Turn is kept for `--resume`. A still-queued message is lost. | Recorded (2.1.284) | +| Is a raw `control_request` `interrupt` honoured without the SDK's `initialize`? | **Yes.** `control_response` success with the receipt in 2 to 4 ms, then an `error_during_execution` `result`. The process stays alive, and the next user frame runs in the same Session with the partial Turn in context. `cancel_queued` works raw. | Recorded (2.1.284) | +| Minimum version for headless mid-Turn pickup and for `priority` | **Not named** in the changelog. Pickup between tool rounds is Recorded on 2.1.284. `priority` was not sent. | Documented (none found); Recorded (2.1.284) | + +## Capability Table + +| Stop or send on the raw `-p` stream | Turn ends with | Process | Partial text in context | Tool call in context | Foreground Bash tree | Queued message | +| -------------------------------------------------------- | ------------------------------------------------------------------------ | ------------------ | -------------------------- | ------------------------------------------------------- | -------------------- | ------------------------------------------------------- | +| SIGTERM (group) | No `result` | Exits 143 | No, not even on resume | Yes: `tool_use` plus `Exit code 137` and partial stdout | Killed by the CLI | Not tested | +| SIGINT (group or pid) | `result` `error_during_execution`, `aborted_streaming` / `aborted_tools` | Exits 0 after ~1 s | Yes, on resume | Yes, as a "rejected" result; partial stdout discarded | Killed | Lost; not replayed on resume | +| `control_request` `interrupt` | `control_response` receipt, then the same `result` | Stays alive | Yes, in the next Turn | Yes, as a "rejected" result; partial stdout discarded | Killed | Runs next by itself; listed in `still_queued` | +| `control_request` `interrupt` with `cancel_queued: true` | Same, with `cancelled: [uuid]` | Stays alive | Not tested (mid-tool only) | Same as above | Killed | Dropped; `command_lifecycle` `cancelled` | +| User frame during a tool call | Same Turn continues | Stays alive | n/a | n/a | Runs to completion | Delivered with the tool result; in `user_message_uuids` | +| User frame while text streams, no tool round left | First Turn ends normally | Stays alive | n/a | n/a | n/a | Runs next as its own Turn with its own `result` | + +## Still Unknown + +- Whether a stdin user frame with `priority: "now"` preempts the running Turn on + the raw stream, and whether `"later"` holds a frame past a tool boundary. + `priority` was never sent. +- Why `queued_turn_count` stayed `0` while a message was waiting, and what makes it + non-zero. +- Whether an interrupted mid-text Turn's tokens are counted anywhere. Its `result` + reported `total_cost_usd: 0` and zero usage. +- Whether a user frame is picked up at a boundary between two tool calls when the + tool round has several calls, or with MCP or permission-bridge tools. Only one + foreground Bash call per round was used. +- What any stop does while a tool call waits on a permission prompt through + Secant's `--permission-prompt-tool` bridge. The driver used `--allowedTools`. +- How a SIGTERM that arrives while a user frame is queued treats that frame. It was + not tested, but the SIGINT result suggests it is lost. **Inferred.** +- Whether `CLAUDE_CODE_RESUME_INTERRUPTED_TURN=1` changes what a resume after + SIGTERM or SIGINT sees. It was not set. +- Whether a stronger model follows a mid-Turn message that conflicts with the + original prompt. Haiku ignored two in test 4c. +- Windows: every behaviour in this note. From 9d021e4ed5fcceac3755995a8a9ede9976432a08 Mon Sep 17 00:00:00 2001 From: rohan Date: Tue, 29 Sep 2026 11:18:13 +0530 Subject: [PATCH 02/11] docs(adr): interrupt ends only the Turn and a mid-Turn message is a native Steer (#255) --- .../0022-own-a-truthful-deep-harness-seam.md | 6 +- ...nd-a-mid-turn-message-is-a-native-steer.md | 75 +++++++++++++++++++ docs/glossary/secant-run-lifecycle.md | 39 ++++++---- docs/headless-parity.md | 15 ++++ .../harness-interrupt-and-queued-messages.md | 4 +- src/headless/AGENTS.md | 4 + 6 files changed, 125 insertions(+), 18 deletions(-) create mode 100644 docs/adr/0035-interrupt-ends-only-the-turn-and-a-mid-turn-message-is-a-native-steer.md create mode 100644 docs/headless-parity.md diff --git a/docs/adr/0022-own-a-truthful-deep-harness-seam.md b/docs/adr/0022-own-a-truthful-deep-harness-seam.md index 97c14f6..accb7f8 100644 --- a/docs/adr/0022-own-a-truthful-deep-harness-seam.md +++ b/docs/adr/0022-own-a-truthful-deep-harness-seam.md @@ -36,7 +36,11 @@ clarification; `interrupt` means Harness-confirmed termination of the active Tur and ending an Interactive agent step remains a Crucible control above the Seam. Expected races return an accepted or rejected receipt rather than throwing. Acceptance does not prove final effect: after an accepted interrupt the Adapter rejects new inputs, drains to native terminal evidence, and the Turn result confirms whether interruption occurred. Harness Requests are Turn-scoped, independently keyed, may coexist, and expire when the -Turn ends, is interrupted, or is lost; an ordinary assistant question at a Turn boundary is not a Harness Request. +Turn ends, is interrupted, or is lost; an ordinary assistant question at a Turn boundary is not a Harness Request. (Edited 2026-09-29: +[ADR 0035](./0035-interrupt-ends-only-the-turn-and-a-mid-turn-message-is-a-native-steer.md) widens `steer` to native mid-Turn delivery of the human's text at the Harness's +next boundary, keeps a Turn open until every accepted Steer is delivered, adds a delivered-or-dropped Steer lifecycle to the event stream, drops +undelivered Steers on interrupt, and qualifies Claude Code's `interrupt` control request as a confirmed active-Turn interruption with a process-stop +fallback.) Native conversation identifiers are opaque recovery coordinates, never Run truth. When one is observable before submission, the Adapter awaits a Crucible-owned durable recorder before sending content; recording failure proves `not-started`. When a Harness reveals it only after acceptance, the diff --git a/docs/adr/0035-interrupt-ends-only-the-turn-and-a-mid-turn-message-is-a-native-steer.md b/docs/adr/0035-interrupt-ends-only-the-turn-and-a-mid-turn-message-is-a-native-steer.md new file mode 100644 index 0000000..98423f7 --- /dev/null +++ b/docs/adr/0035-interrupt-ends-only-the-turn-and-a-mid-turn-message-is-a-native-steer.md @@ -0,0 +1,75 @@ +# Interrupt Ends Only the Turn, and a Mid-Turn Message Is a Native Steer + +An **Interrupt** stops the live **Turn** and nothing more: the **Step Attempt** stays open, the **Run** waits on the human, and the human's next +message continues the same **Harness Session**. A message the human sends while a Turn is working is a **Steer** in both agent Step kinds, delivered +natively at the **Harness**'s next boundary, and the Turn does not end until every Steer has been delivered. The need came from the pre-public-release +reports ([#235](https://github.com/secantdev/secant/issues/235)): in the native clients Escape stops the Turn and the human types on in the same +Session, and a message typed mid-Turn is picked up without interrupting, while Secant halted the Run on Interrupt, offered Claude Code no Steer, and +refused a mid-Turn send. This supersedes the Run lifecycle glossary's Interrupt rule (the Attempt ends `cancelled` and the Run rests `halted`) and +its Steer definition ("not a new Turn, and unsupported Harnesses do not emulate it"), [Spec: M3](https://github.com/secantdev/secant/issues/107) +stories 18 and 19 and its deferral of a Session-preserving Claude Code interrupt, and [Spec: M4](https://github.com/secantdev/secant/issues/137) +story 19's Codex-only Steer. It amends [ADR 0022](./0022-own-a-truthful-deep-harness-seam.md)'s `steer` and `interrupt` controls and Claude Code's +interruption mode. Unchanged: closing Secant or Ctrl+C halts the Run ([ADR 0019](./0019-failed-and-halted-runs-are-resumable-resting-states.md)), +Cancel is the only route to `cancelled`, an Agent call made in an interrupted Turn is dropped +([ADR 0032](./0032-let-opted-in-interactive-agent-steps-accept-agent-declared-completion.md)), and headless gains no Interrupt or Steer. + +**Interrupt, per Step kind.** In an **Interactive agent step** the interrupted Turn ends `interrupted` and the Run returns to `blocked` on the human, +whose next message goes into the same Session; no Attempt is published and the Iteration stays open. An interrupted **Entry Turn** is treated the same +way and is never re-sent. In an **Agent step** the interrupted Turn ends `interrupted`, the Attempt stays open, and the Run is `blocked` on the human's +message. That message starts a human-origin follow-up Turn in the same Session and the same Attempt, and the Attempt takes its outcome from its last +Turn: a clean end completes the Step and the Routing advances without the human, exactly as an uninterrupted Agent step does, and a failure takes the +Step's ordinary retry policy. A second Interrupt is the only way to hold the Step again. A `lost` Turn keeps its current meaning. When Secant closes +while a Run waits after an Interrupt, the Run halts as any live Run does, and resuming returns it to waiting on the human's message. + +**Steer.** Steer now means native mid-Turn delivery of the human's own text, landing at the Harness's next boundary; it is offered in Agent steps and +Interactive agent steps alike, wherever the profile's Steer evidence says the Harness supports it. Codex delivers it through `turn/steer` with +`expectedTurnId` and a Secant-minted `clientUserMessageId`, which comes back as the `clientId` of the `userMessage` item written when the text is taken +into history before the next model request. Claude Code delivers it as a stdin `user` frame on the existing `claude -p --input-format stream-json` +process, stamped with a Secant-minted `uuid`: a frame written during a tool call is taken with that tool's result in the same exchange and listed in +`result.user_message_uuids`, and one written while text streams runs as the next native exchange with its own `result`; `command_lifecycle` frames, +advertised as `msg_lifecycle_v1`, say which frames are still queued. Delivery means the Harness put the text in front of the model, never that the +model followed it. + +**The Turn stretches until every Steer is delivered.** A Turn ends at the first Harness boundary after which no accepted Steer is pending, so one Secant +Turn may span several native exchanges. The human's message therefore always reaches the agent before an Agent-declared completion from that Turn is +applied. Codex can take a steered text into history after its last model request, during stop hooks, post-Turn compaction, or a failing Turn, and then +complete without answering it; the Codex Adapter re-delivers that leftover with a native `turn/start` inside the same Secant Turn, using an empty +input if Codex accepts one and the same text otherwise, so a Steer is answered before its Turn ends on both Harnesses. That text is the human's own, so +this is ordinary turn-taking, not emulation. A Steer that arrives after the Turn has ended is rejected and its draft kept: in an Interactive agent +step the human sends it as the next Turn, and in an Agent step that has advanced the human is told it was not delivered. + +**Interrupt drops what has not been delivered.** Every accepted Steer still pending when an Interrupt lands is dropped and its text returned to the +human's compose, so nothing runs after the human said stop. Codex discards pending steered input on `turn/interrupt`; Claude Code is interrupted with +`cancel_queued: true`, which hands the queued frames back instead of running them as the next Turn. + +**How Claude Code stops.** Claude Code is interrupted with a raw `control_request` `interrupt` written to the same stdin, which, recorded on Claude +Code 2.1.284, answers without the SDK's `initialize` in milliseconds, ends the Turn with a `result`, keeps the process and Session live for the next +frame, keeps the partial text in context, and kills the foreground tool tree. That makes Claude Code's interruption a confirmed active-Turn +interruption rather than a process-only stop. Qualification bounds the dependency: a Claude Code that does not answer the request falls back to +today's SIGTERM of the process tree and `--resume`, which loses partial streamed text. This is the same wire and the same degrade rule +[ADR 0034](./0034-choose-and-change-model-and-effort-as-one-run-wide-model-choice.md) adopted for `set_model`; adopting the Agent SDK itself stays +rejected for the reasons in [Establish Claude Code's viable structured transports](https://github.com/secantdev/secant/issues/3). An Interrupt stops +the Turn's foreground tool work; a shell the model moved to the background may outlive it until the Session closes. + +**Recording.** Each Steer is durable against its Turn: the human's text, when it was sent, and how it settled, as delivered within the Turn, +delivered after a native boundary, re-delivered, or dropped by an Interrupt or a lost Turn. An interrupted Turn records the stop used, the control +request or the process stop. A follow-up Turn after an Interrupt is a human-origin Turn inside the Agent step's Attempt. Steered text is shown as the +human's message inside its Turn, durable content kept distinct from live previews under +[ADR 0024](./0024-use-one-deep-projection-port-for-tui-and-headless-clients.md). + +**Harness Interface.** `steer` keeps its `ControlReceipt`, whose acceptance now means the text was handed to the Harness. The Turn's closed event +stream gains one Steer lifecycle, delivered or dropped, beside the Harness Request lifecycle, and on terminal the Adapter publishes a dropped event +for any Steer never delivered before it closes the producer and settles the one result. Which Steers are pending, the native correlation ids, and +where the real Turn boundary falls stay private to each Adapter; execution still sees one result per Turn and learns nothing native. The profile's +Steer evidence states each Harness's delivery point, and Claude Code declares `active-turn` interruption when qualification confirms the control +request. + +Rejected: keeping Interrupt as a halt, which turns stopping one's own conversation into a Run-level stop outside the workflow's logic; resending an +interrupted Agent step's prompt on resume; an interrupted Agent step that waits for the human to end it, which silently turns it into an Interactive +agent step; a client-side queue sent as the next Turn, which is emulation and would race the one live Turn per Prepared Harness and the +Agent-declared completion's clean-end settlement; a separate queue key beside Steer; counting Claude Code's second native exchange as a new Turn the +Harness started, which would let an Agent-declared completion settle the Step while the human's message is still being answered; reporting a Codex +leftover upward as missed, which pushes a native race into two Step kinds; SIGINT, which exits the process and silently loses queued messages; and +`queued_turn_count`, which stayed 0 while a message waited. How the Run Workbench presents Interrupt, Steer, a dropped Steer's restored draft, and +delivered and dropped Steers is left to the Run Workbench prototype. The decision was made on +[Decide how the human interrupts a Turn and sends a message while a Turn is working](https://github.com/secantdev/secant/issues/255). diff --git a/docs/glossary/secant-run-lifecycle.md b/docs/glossary/secant-run-lifecycle.md index 67c259a..60b73ab 100644 --- a/docs/glossary/secant-run-lifecycle.md +++ b/docs/glossary/secant-run-lifecycle.md @@ -14,7 +14,8 @@ This cluster defines the target Secant terms for a **Run** and everything that h a Repeat group is a Step but not a Stage; the group is the Stage. **End Stage** and an **Agent-declared completion**'s stage done complete a human-controlled group's Stage. - **Agent step** — a **Step kind** running one autonomous **Turn** in a named **Harness Session**. It completes without the human, though the human - may **Steer** it while that Turn is live when the selected Harness supports native Steer. + may **Steer** it while a Turn is live when the selected Harness supports Steer. After an **Interrupt** the human's message starts a follow-up Turn + in the same Session and **Step Attempt**, and the Attempt takes its outcome from its last Turn. - **Command step** — the deterministic non-agent **Step kind**. Its attempt succeeds if the command ran to an exit; the exit status becomes a **Verdict** and the captured output a `text` **Run Artifact**. The attempt fails only when the command could not execute. - **Verdict** — a **Run Artifact** type holding `pass` or `fail`, produced from a deterministic **Step**'s exit status. The only thing a **Repeat @@ -91,8 +92,9 @@ This cluster defines the target Secant terms for a **Run** and everything that h - **Harness Session** — a named conversation with the selected **Harness**, owned by exactly one **Run** and never shared across Runs. The routing names the session each agent **Step** runs in; Secant opens it on first use and reuses it after. - **Turn** — one mechanical user-to-**Harness** exchange inside a **Harness Session**: submitted input, model and tool activity, streamed progress, - and the Harness's authoritative turn boundary. It is neither a Session nor a judgement that the **Step** reached its goal. An **Agent step** has - one Turn per attempt; an **Interactive agent step** may have many. + and the Harness's authoritative turn boundary, the first one after which no **Steer** is still pending, so one Turn may span several native + exchanges. It is neither a Session nor a judgement that the **Step** reached its goal. An **Agent step** has one autonomous Turn per attempt, plus + any follow-up Turns the human starts after an **Interrupt**; an **Interactive agent step** may have many. - **Session availability** — whether a **Harness Session** is `open` (a next Turn can be sent now), `detached` (not live, but holding a native recovery coordinate worth reattaching), or `unusable` (native evidence authoritatively says recovery cannot continue). - **Model choice** — the **Run**'s one current model and effort level, each a real value the selected **Harness** names, never "Harness default". It @@ -117,11 +119,14 @@ This cluster defines the target Secant terms for a **Run** and everything that h recognised phrase in a **Turn**. The legacy grill is this shape. A step may opt into an **Entry Turn**. Inside a **Repeat group** each iteration is its own **Step Attempt** with its own **Harness Session**; ending the step advances only that iteration. - **Entry Turn** — an **Interactive agent step**'s optional first **Turn**: its Bundle-authored prompt, rendered with **Launch inputs** and bundled - skill paths, sent once on entry so the human need not retype what the launch already carries. It is never re-sent: after a halt the human - continues the same **Harness Session**. -- **Steer** — sending native same-Turn guidance while a **Turn** is live. It is not a new Turn, and unsupported Harnesses do not emulate it. -- **Interrupt** — asking the **Harness** to stop the current live **Turn**, including its native tool work. Confirmation ends the **Step Attempt** - `cancelled` and leaves the **Run** `halted` and re-attemptable; it does not close the Harness or cancel the Run. + skill paths, sent once on entry so the human need not retype what the launch already carries. It is never re-sent: after an **Interrupt** or a halt the + human continues the same **Harness Session**. +- **Steer** — a message the human sends while a **Turn** is live, delivered natively at the **Harness**'s next boundary and inside that Turn, + which lasts until every Steer is delivered. Delivered means the Harness put it in front of the model, not that the model followed it. An + **Interrupt** drops a Steer not yet delivered and returns its text to the human. _Avoid_: queued message. +- **Interrupt** — asking the **Harness** to stop the current live **Turn** and its foreground tool work. It ends only the Turn: the **Step Attempt** + stays open and the **Run** waits `blocked` on the human, whose next message continues the same **Harness Session**. It does not close the + Harness, halt the Run, or cancel it. - **Cancel** — explicitly ending a **Run**. The only route to the terminal `cancelled` state. - **Preflight** — the precondition check performed before a **Run** exists: the **Composition check**, presence of required **Launch inputs**, the union of authored **Workspace prerequisites**, intrinsic **Step kind** preconditions and Harness capability needs, and resolution of each selected @@ -137,14 +142,14 @@ An Agent-bearing **Run** pins its semantic **Harness** selection with its **Work routing truth, while the Model choice may change between and during Turns; each Agent-step Attempt separately records the Harness executable, version, and Adapter evidence, and each Turn its requested and effective model and effort. -| State | Meaning | Terminal | -| ----------- | --------------------------------------------------------------- | -------- | -| `running` | a **Step Attempt** is executing | no | -| `blocked` | a **Human Gate** or **Harness Request** is waiting on the human | no | -| `halted` | stopped for a reason outside the workflow's logic | no | -| `failed` | the workflow concluded negatively | no | -| `succeeded` | the routing completed | yes | -| `cancelled` | the user explicitly ended the Run | yes | +| State | Meaning | Terminal | +| ----------- | ----------------------------------------------------------------------------- | -------- | +| `running` | a **Step Attempt** is executing | no | +| `blocked` | a **Human Gate**, **Harness Request**, or the human's next message is waiting | no | +| `halted` | stopped for a reason outside the workflow's logic | no | +| `failed` | the workflow concluded negatively | no | +| `succeeded` | the routing completed | yes | +| `cancelled` | the user explicitly ended the Run | yes | `blocked` is durable truth, not computed: since [#108](https://github.com/secantdev/secant/issues/108) the authored **Human Gate**'s `pending_gate` row and the `blocked` state are written in one transaction, and execution also stores `blocked` before a checkpoint pause, so a killed Run reconciles @@ -191,6 +196,8 @@ row and the `blocked` state are written in one transaction, and execution also s to ADR 0020's reasons, and the human-controlled group's **Review checkpoint**. - [ADR 0033](../adr/0033-carry-agent-calls-to-secant-over-a-per-session-loopback-mcp-server.md) owns the **Agent call**'s channel, its attribution to a **Harness Session**'s live **Turn**, and its reply. +- [ADR 0035](../adr/0035-interrupt-ends-only-the-turn-and-a-mid-turn-message-is-a-native-steer.md) owns what an **Interrupt** leaves behind, + **Steer** delivery, and why a Turn lasts until every Steer is delivered. - [ADR 0023](../adr/0023-own-durable-run-truth-in-isolated-run-stores.md) owns durable Run truth, Artifact publication, Workspace materialization, retention, and recovery storage. - [ADR 0031](../adr/0031-own-runs-per-run-not-per-workspace.md) owns Run ownership: many live Runs per Workspace, one owner per Run, and what a diff --git a/docs/headless-parity.md b/docs/headless-parity.md new file mode 100644 index 0000000..351fd06 --- /dev/null +++ b/docs/headless-parity.md @@ -0,0 +1,15 @@ +# Headless Parity + +What the Run Workbench (the TUI) offers that the headless `secant` commands do not. Headless is for scripts and CI, so it drives a Run to a resting +state without a human at the keyboard; anything that needs the human inside a live Turn is TUI-only. A gap listed here is deliberate unless its +decision says otherwise; closing one needs its own decision. + +| TUI capability | Headless | Decided in | +| --------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------ | ----------------------------------------------------------------------------------------------- | +| Interactive agent steps | `run launch` refuses a Bundle whose Routing reaches one ("Interactive agent Steps are TUI-only") | [#122](https://github.com/secantdev/secant/issues/122) | +| **Interrupt** a live Turn, leaving the Run waiting on the human | none; Ctrl+C is a signal that halts the Run (ADR 0019) | [ADR 0035](./adr/0035-interrupt-ends-only-the-turn-and-a-mid-turn-message-is-a-native-steer.md) | +| The follow-up message that continues an Agent step after an Interrupt | none, since headless cannot Interrupt | [ADR 0035](./adr/0035-interrupt-ends-only-the-turn-and-a-mid-turn-message-is-a-native-steer.md) | +| **Steer**: a message sent while a Turn is working | none | [ADR 0035](./adr/0035-interrupt-ends-only-the-turn-and-a-mid-turn-message-is-a-native-steer.md) | + +Wider headless parity for the agent screen, including Turn history that grows during a Turn and its frozen `--json` shapes, is still open on +[Chart the fixes Secant needs before its public release](https://github.com/secantdev/secant/issues/235). diff --git a/docs/research/harness-interrupt-and-queued-messages.md b/docs/research/harness-interrupt-and-queued-messages.md index c17e5f2..de982c4 100644 --- a/docs/research/harness-interrupt-and-queued-messages.md +++ b/docs/research/harness-interrupt-and-queued-messages.md @@ -279,7 +279,9 @@ installed version's tag. T3 Code was read from a local clone at the commit above retried. A `lost` Turn maps to `indeterminate`, and the Run also rests `halted`. `resume-run` re-attempts the Step. Because the Session record is `detached`, the retry sends the Step's prompt again into the same native Session.[^sec-exec-cancelled][^sec-agent-recovery] - **Interactive agent step.** - - Each human Turn runs with `detachAfterTurn`, and the Harness is closed after each Turn.[^sec-agent-interactive] + - Each human Turn runs with `detachAfterTurn`, which records the Session `detached` after each Turn.[^sec-agent-interactive] (Corrected + 2026-09-29: the Harness is not closed after each Turn. The wiring keeps one prepared Harness, and one Claude Code process, across the + Step's Turns; only Codex reissues `thread/resume` before each human Turn because of the `detached` record.) - An interrupted or lost human Turn rests the Run `halted`, not `blocked`, and publishes no Attempt. An interrupted Entry Turn does the same.[^sec-app-interactive][^sec-exec-entry] - The glossary says that "after a halt the human continues the same Harness Session". That takes a resume back to `blocked` before the diff --git a/src/headless/AGENTS.md b/src/headless/AGENTS.md index f264148..7e56844 100644 --- a/src/headless/AGENTS.md +++ b/src/headless/AGENTS.md @@ -69,3 +69,7 @@ Inherits the engineering baseline; records only non-obvious local facts. Ownersh `Bun.main` separator normalisation, the #62 Windows entry quirk (`:48-53`); and the top-level error handler that prints a stack and sets exit 1 (`:55-62`). The file records at `:40-47` why it cannot be unit-tested as written; if a fourth branch appears, revisit a small `tests/cli` rather than widening the smoke. + +## Read next + +- [Headless parity](../../docs/headless-parity.md) lists the TUI capabilities headless deliberately lacks; update it when a decision adds or closes one. From 96a07c1afc2f83bb3a488b1f40d3676d8360dd68 Mon Sep 17 00:00:00 2001 From: rohan Date: Tue, 29 Sep 2026 11:43:33 +0530 Subject: [PATCH 03/11] docs(research): record Claude Code interrupt and steer through the permission bridge on 2.1.284 (#255) Round 2 of the live note: a raw interrupt while a tool waits on a stand-in MCP permission bridge, SIGTERM in the same state, and mid-Turn user frames across multi-call rounds, MCP calls, and a permission wait. --- ...e-code-live-interrupt-and-mid-turn-send.md | 346 ++++++++++++++++-- 1 file changed, 324 insertions(+), 22 deletions(-) diff --git a/docs/research/claude-code-live-interrupt-and-mid-turn-send.md b/docs/research/claude-code-live-interrupt-and-mid-turn-send.md index c85f058..b924939 100644 --- a/docs/research/claude-code-live-interrupt-and-mid-turn-send.md +++ b/docs/research/claude-code-live-interrupt-and-mid-turn-send.md @@ -8,7 +8,9 @@ x64; `claude --version` returned `2.1.284 (Claude Code)`). Ticket: [#255](https://github.com/secantdev/secant/issues/255). This note settles the Claude Code items left **Untested** or **Unknown** in [Harness Interrupt and Queued Messages](harness-interrupt-and-queued-messages.md) -by running small live model sessions against the installed CLI. +by running small live model sessions against the installed CLI. A second round, +[Round 2](#round-2-permission-waits-and-multi-tool-rounds), repeats the key stops and +sends through a stand-in for Secant's MCP permission bridge. ## Answer @@ -80,6 +82,30 @@ The nearest entries are listed under [Version floor](#test-6-version-floor). `control_request` `interrupt`. SIGINT keeps the partial Turn but ends the process. SIGTERM ends the process and loses streamed text. +**Round 2: the same holds while a tool call waits on the permission bridge.** See +[Round 2](#round-2-permission-waits-and-multi-tool-rounds). + +- **Raw `interrupt` during a permission wait.** It was honoured in three of three + runs, with a `control_response` in 2 to 4 ms. Claude Code sent the bridge an MCP + `notifications/cancelled` for the pending `approve` call within 4 ms. The Turn + ended with the same `error_during_execution` / `aborted_tools` `result` as a + mid-tool interrupt, and `permission_denials` named the waiting call. The process + stayed alive and the next user frame ran in the same Session. The model saw the + call as "rejected before it could execute". The bridge's late `allow`, 30 s + later, had no effect: the command never ran. +- **SIGTERM during a permission wait (one run).** The process exited 143 with no `result` + and no `tool_result`. No `notifications/cancelled` was sent. During shutdown, + Claude Code opened a new MCP session to the bridge and sent a second `approve` + call for the same `tool_use_id`. On `--resume`, the transcript gained a synthetic + `tool_result`: "[Tool call interrupted: the session ended before this call's + result was recorded, so its outcome is unknown…]". +- **A user frame is still taken at a tool boundary, but only after the whole tool + round.** This held when two MCP calls ran at once, when a Bash call and an MCP + call ran one after the other in one round, and for a single MCP call. It also + held for a frame written while a Bash call waited on the bridge: the frame was + held through the approval and the command. Each time, the frame reached the model + in the same Turn, and its `uuid` was in that `result`'s `user_message_uuids`. + Windows was not tested. ## Evidence Vocabulary @@ -143,8 +169,8 @@ conversation before this message? Quote the last line you wrote…". Differences from Secant's launch: - Secant adds the loopback MCP permission bridge (`--mcp-config`, - `--permission-prompt-tool`). The driver did not. A tool call parked on a - permission prompt at the moment of a stop was not tested. + `--permission-prompt-tool`). The round 1 driver did not. Round 2 added a stand-in + bridge; see [Round 2](#round-2-permission-waits-and-multi-tool-rounds). - The host's user settings load no-op `SessionStart`, `PreToolUse`, `PostToolUse`, and `Stop` hooks, plus a user `CLAUDE.md`. @@ -521,25 +547,296 @@ The installed 2.1.284 advertises both, and also `msg_lifecycle_v1`, in Not tested. No Windows host was available. Signal delivery, `taskkill` tree behaviour, and the SIGINT and SIGTERM results above are all Linux-only. +## Round 2: permission waits and multi-tool rounds + +Run on 2026-09-29 against the same Claude Code **2.1.284**. Every fact in this +section is Recorded (2.1.284) unless it carries another label. + +### Round 2 method + +The round 2 driver used the same launch flags as `src/harness/claude-code.ts`: +`-p --input-format stream-json --output-format stream-json --verbose +--include-partial-messages --model haiku --session-id ` (or `--resume `), +then the bridge fragment. It passed no `--allowedTools` and no permission-mode flag, +so `system/init` reported `permissionMode: "default"`. + +The stand-in bridge copied `src/harness/permission-bridge.ts`: + +- a Streamable HTTP MCP server on `127.0.0.1` with a random port and a bearer + token, one transport per MCP session, built on the repository's + `@modelcontextprotocol/sdk` 1.29.0; +- server `secant-permissions` with one `approve` tool (`tool_name`, `input`, + `tool_use_id`), returning `{"behavior":"allow","updatedInput":…}` or + `{"behavior":"deny",…}` as JSON text; +- launch fragment `--mcp-config --permission-prompt-tool +mcp__secant-permissions__approve`. + +Unlike Secant, the stand-in let the experiment set a delay for each `approve` +answer, and it did not stop when the call was cancelled. It logged every HTTP +request body, every closed response stream, and the handler's abort signal. For +round 2 it also served a second MCP server, `slowtools`, from the same process. Its +one tool, `slow_echo(text, delay_s)`, waits `delay_s` seconds and returns `echo: +`. It is marked `readOnlyHint: true`. + +Other differences from round 1: + +- `ORCA_*` variables were removed from the child environment, as well as + `CLAUDE*`. The host's user hooks, including a `PermissionRequest` hook, then + printed `{}` and did not decide anything. Every `Bash ./slowjob.sh` and + `./slow15.sh` call reached the stand-in `approve`. A Bash `echo` the model ran + on its own did not: Claude Code allowed it without a prompt. +- The user's other MCP servers (context7 and the claude.ai connectors) also loaded, + as they would under Secant. MCP tools were deferred, so the model sometimes + called `ToolSearch` before `slow_echo`. +- `slow15.sh` is `echo started-15; sleep 15; echo done-15`. + +`HTTP` and `MCP` lines below are the stand-in's log, on the same clock as the +stream frames. + +### R2-A1: raw `interrupt` while Bash waits on the bridge (two runs, plus one with `cancel_queued`) + +The stand-in held the Bash `approve` for 30 s, then answered `allow`. The +interrupt came 3 s into the wait. Run 1: + +```text +t=5832 assistant tool_use toolu_01Y7… Bash {"command":"./slowjob.sh",…} +t=5854 HTTP POST /mcp tools/call#2 name=approve +t=5857 MCP approve#1 CALLED tool=Bash tool_use_id=toolu_01Y7… -> will allow after 30000ms +t=8937 ps: no slowjob or sleep 27 +t=8938 in {"type":"control_request","request_id":"int1","request":{"subtype":"interrupt"}} +t=8940 control_response {"subtype":"success","request_id":"int1","response":{"still_queued":[]}} +t=8942 HTTP POST /mcp notifications/cancelled params={"requestId":2,"reason":"AbortError: remote-cancel"} +t=8944 MCP approve#1 extra.signal ABORTED +t=8946 user tool_result "The user doesn't want to proceed with this tool use. The tool use was + rejected …" is_error +t=8948 user text "[Request interrupted by user for tool use]" +t=8958 result subtype=error_during_execution is_error=true terminal_reason=aborted_tools + stop_reason=tool_use num_turns=3 result_index=0 queued_turn_count=0 + user_message_uuids=[] + permission_denials=[{"tool_name":"Bash","tool_use_id":"toolu_01Y7…",…}] +t=8960 command_lifecycle state=cancelled +t=10495 in user +t=17396 assistant text "I attempted to call the Bash tool to run `./slowjob.sh`, but the tool use + was rejected before it could execute. No command finished and no output was + produced—the tool call was denied by the user." +t=17424 result subtype=success num_turns=1 result_index=1 +t=35858 MCP approve#1 RETURNING {"behavior":"allow",…} (signal.aborted=true) +t=39891 ps: no slowjob or sleep 27; no further stream frames +t=42932 HTTP approve response stream closed (only when the driver closed stdin) +t=43767 EXIT code=0 +``` + +Run 2 (test C, the repeat) gave the same frames: `control_response` in 2 ms, +`notifications/cancelled` 4 ms after the interrupt, the same `result`, and the same +`permission_denials`. Asked afterwards, the model said "My only action was an +attempted Bash tool call to run `./slowjob.sh`, which was rejected before +execution." + +- **Honoured, and the process stays alive.** The next user frame ran as + `result_index: 1` in the same process and Session. +- **Claude Code cancels the bridge call.** The bridge receives a standard MCP + `notifications/cancelled` with the `approve` call's JSON-RPC id and `reason: +"AbortError: remote-cancel"`. It is sent on a new POST. The HTTP response stream + for the cancelled call is not closed. It stayed open until the process exited. +- **A late answer does nothing.** The stand-in's `allow` came 27 s after the + interrupt. The SDK server does not send a response for a cancelled request, so + nothing reached Claude Code. **Inferred** from the SDK, and consistent with what + was seen: `slowjob.sh` never started, and no frame followed. +- **The transcript matches a mid-tool interrupt.** The `tool_result` is stored with + `toolUseResult: "User rejected tool use"` and `toolDenialKind: "user-rejected"`, + followed by `[Request interrupted by user for tool use]`. Nothing marks that the + call was waiting on approval and never ran. Only `permission_denials` in the + `result` names it. + +With `cancel_queued: true` and a user frame written 1.5 s before the interrupt: + +```text +t=5571 MCP approve#1 CALLED tool=Bash … -> will allow after 30000ms +t=7110 in user uuid=a1a1a1a1-…-01 "Queued message: reply with the single word MANGO." +t=7111 command_lifecycle a1a1a1a1-… state=queued +t=8610 in control_request interrupt int1 cancel_queued=true +t=8613 command_lifecycle a1a1a1a1-… state=cancelled +t=8614 control_response {"still_queued":[],"cancelled":["a1a1a1a1-0000-4000-8000-000000000001"]} +t=8615 HTTP POST /mcp notifications/cancelled params={"requestId":2,…} +t=8631 result subtype=error_during_execution terminal_reason=aborted_tools result_index=0 +t=13667 in user (no Turn started before this) +t=18831 result subtype=success result_index=1 +``` + +The queued message was dropped as in test 5b, and the bridge call was cancelled the +same way. + +### R2-A2: SIGTERM while Bash waits on the bridge + +```text +t=4926 assistant tool_use toolu_01Ub… Bash {"command":"./slowjob.sh",…} +t=4975 MCP approve#1 CALLED tool=Bash tool_use_id=toolu_01Ub… reqId=2 (held 30 s) +t=8033 SIGNAL SIGTERM group +t=8037 HTTP response streams closed: both GET streams and the approve#1 stream +t=8055 HTTP POST /mcp server/discover, then initialize → new MCP session cea0a024 +t=8072 HTTP POST /mcp sid=cea0a024 tools/call#1 name=approve +t=8073 MCP approve#2 CALLED tool=Bash tool_use_id=toolu_01Ub… (same call, again) +t=8896 approve#2 response stream closed +t=8897 EXIT code=143 signal=null +``` + +- **The process exits 143 with no `result`**, 864 ms after the signal. No + `notifications/cancelled` was sent. The bridge saw its streams drop. +- **A second `approve` call arrives during shutdown.** Claude Code opened a new MCP + session and asked again for the same `tool_use_id`, 39 ms after the SIGTERM. It + exited without waiting for the answer. The bash command never ran. Seen in the + one run. +- **The original process writes nothing for the waiting call.** The transcript ends + at the `tool_use`, with no `tool_result`. +- **`--resume` fills the gap with a synthetic result.** The resume process + (launched with the bridge flags) appended, before its first Turn: + +```text +user tool_result "[Tool call interrupted: the session ended before this call's result + was recorded, so its outcome is unknown. Check whether it took effect before + relying on it or running it again.]" is_error toolDenialKind="interrupted" +assistant "No response requested." model= +user +``` + +The resumed model answered: "The last line I wrote was "No response requested." The +command did not finish — the session ended before the tool result was recorded, so +I received no output from `./slowjob.sh`." The resume process made no `approve` +call. + +This differs from test 1a. There, a SIGTERM during a running Bash call wrote a real +`Exit code 137` result before the exit. + +### R2-B1: a user frame during a round with two slow tool calls + +**Two `slow_echo` calls in one message (8 s and 20 s).** Both ran at once. Both +`approve` calls came first, and both `slow_echo` calls started within 350 ms. + +```text +t=5514 assistant tool_use toolu_013T… slow_echo {"text":"alpha","delay_s":8} +t=5543 MCP slow_echo#1 START alpha +t=5855 assistant tool_use toolu_012b… slow_echo {"text":"beta","delay_s":20} +t=5890 MCP slow_echo#2 START beta +t=8585 in user uuid=b1a00000-…-01 "Additional instruction: when you reply, also say the word MANGO." +t=8587 command_lifecycle b1a00000-… state=queued +t=13562 user tool_result toolu_013T… "echo: alpha" +t=25904 user tool_result toolu_012b… "echo: beta" +t=25921 command_lifecycle b1a00000-… state=started +t=30394 assistant text "echo: alpha\necho: beta" +t=30430 command_lifecycle b1a00000-… state=completed +t=30434 result subtype=success num_turns=3 result_index=0 + user_message_uuids=[, "b1a00000-0000-4000-8000-000000000001"] +``` + +**Bash `./slow15.sh` and `slow_echo` (5 s) in one message.** These ran one after +the other. The `approve` call for `slow_echo` came only after Bash finished. + +```text +t=10142 assistant tool_use toolu_01E8… Bash {"command":"./slow15.sh",…} +t=10166 MCP approve#1 CALLED tool=Bash (allowed at once) +t=10347 assistant tool_use toolu_01UN… slow_echo {"text":"gamma","delay_s":5} +t=14186 in user uuid=b1b00000-…-01 "Additional instruction: … MANGO." +t=14188 command_lifecycle b1b00000-… state=queued +t=25233 user tool_result toolu_01E8… "started-15\ndone-15" +t=25246 MCP approve#2 CALLED tool=mcp__slowtools__slow_echo +t=25253 MCP slow_echo#1 START gamma +t=30266 user tool_result toolu_01UN… "echo: gamma" +t=30299 command_lifecycle b1b00000-… state=started +t=33981 assistant text "PINEAPPLE" +t=34012 result subtype=success num_turns=4 + user_message_uuids=[, "b1b00000-0000-4000-8000-000000000001"] +``` + +An earlier run wrote the frame during the second call, the MCP one. It was taken +after that call's result, in the same way. + +- **A frame is taken only after the whole round.** It is not taken after the first + call's result, whether the calls run together (first result at 13.6 s, pickup at + 25.9 s) or one after the other (the Bash result, then the whole MCP call, then + pickup). In the transcript, the `queued_command` attachment follows the last + `tool_result` of the round, and the `queue-operation` `remove` has `reason: +"absorbed_mid_turn"`. +- **Same Turn, listed.** Each time there was one `result`, and the picked-up `uuid` + was in its `user_message_uuids`. +- Haiku did not say MANGO in either run. Delivery does not mean compliance, as in + test 4c. + +### R2-B2: a user frame during one MCP call, and during a permission wait + +**During one `slow_echo` (15 s).** + +```text +t=9359 assistant tool_use toolu_01Qq… slow_echo {"text":"delta","delay_s":15} +t=9387 MCP slow_echo#1 START delta +t=12410 in user uuid=b2a00000-…-01 "Additional instruction: … MANGO." +t=12412 command_lifecycle b2a00000-… state=queued +t=24400 user tool_result toolu_01Qq… "echo: delta" +t=24415 command_lifecycle b2a00000-… state=started +t=28050 assistant text "delta" +t=28085 result subtype=success num_turns=4 + user_message_uuids=[, "b2a00000-0000-4000-8000-000000000001"] +``` + +The MCP call was not cut short, and no `notifications/cancelled` was sent. The +frame was taken after the MCP result, as for Bash in test 4a. In a first run, the +frame arrived while the model was streaming the call that became `slow_echo`. It +was held through the whole MCP call and taken after its result, in the same Turn. + +**While Bash waits on the bridge (held 10 s, then `allow`).** + +```text +t=29043 assistant tool_use toolu_01Nn… Bash {"command":"./slow15.sh",…} +t=29068 MCP approve#1 CALLED tool=Bash -> will allow after 10000ms +t=32092 in user uuid=b2b00000-…-01 "Additional instruction: … MANGO." +t=32095 command_lifecycle b2b00000-… state=queued +t=39070 MCP approve#1 RETURNING {"behavior":"allow",…} +t=42120 system/task_started local_bash +t=54146 user tool_result toolu_01Nn… "started-15\ndone-15" +t=54168 command_lifecycle b2b00000-… state=started +t=56342 assistant text "PINEAPPLE MANGO" +t=56369 result subtype=success num_turns=2 + user_message_uuids=[, "b2b00000-0000-4000-8000-000000000001"] +``` + +- **The frame does not end or answer the permission wait.** It was held through + the rest of the wait, the approval, and the 15 s command. It was taken after the + tool result, in the same Turn, and listed. Haiku acted on it this time. + +### `command_lifecycle` in round 2 + +In every run: + +- A written frame got `queued` within 1 to 3 ms. +- A picked-up frame got `started` 15 to 33 ms after the round's last `tool_result` + frame, and `completed` just before the Turn's `result`. The prompt's `completed` + came just after the `result`. +- After an interrupt, the prompt got `cancelled` just after the `result`. +- With `cancel_queued`, the queued frame got `cancelled` just before the + `control_response`. + ## The Four Original Unknowns -| Unknown in [Harness Interrupt and Queued Messages](harness-interrupt-and-queued-messages.md) | Result on 2.1.284 | Evidence | -| -------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ------------------------------------------- | -| Does SIGTERM in `-p` mode keep the partial assistant text in the transcript that `--resume` loads? | **No for streamed text:** nothing of the unfinished model call is saved, not even thinking. **Yes for a finished tool round:** the `tool_use` and a real `Exit code 137` result with partial stdout are saved. There is no interrupted marker. Resume adds a synthetic `No response requested.`. | Recorded (2.1.284) | -| Does SIGINT to a `-p --input-format stream-json` process end only the Turn and keep reading stdin? | **No.** It writes one `error_during_execution` `result`, then exits with code 0 about 1 s later, with stdin still open. The partial Turn is kept for `--resume`. A still-queued message is lost. | Recorded (2.1.284) | -| Is a raw `control_request` `interrupt` honoured without the SDK's `initialize`? | **Yes.** `control_response` success with the receipt in 2 to 4 ms, then an `error_during_execution` `result`. The process stays alive, and the next user frame runs in the same Session with the partial Turn in context. `cancel_queued` works raw. | Recorded (2.1.284) | -| Minimum version for headless mid-Turn pickup and for `priority` | **Not named** in the changelog. Pickup between tool rounds is Recorded on 2.1.284. `priority` was not sent. | Documented (none found); Recorded (2.1.284) | +| Unknown in [Harness Interrupt and Queued Messages](harness-interrupt-and-queued-messages.md) | Result on 2.1.284 | Evidence | +| -------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------- | +| Does SIGTERM in `-p` mode keep the partial assistant text in the transcript that `--resume` loads? | **No for streamed text:** nothing of the unfinished model call is saved, not even thinking. **Yes for a finished tool round:** the `tool_use` and a real `Exit code 137` result with partial stdout are saved. There is no interrupted marker. Resume adds a synthetic `No response requested.`. **Round 2:** for a call waiting on the permission bridge nothing is saved after the `tool_use`; resume adds a synthetic "Tool call interrupted… outcome is unknown" `tool_result`. | Recorded (2.1.284) | +| Does SIGINT to a `-p --input-format stream-json` process end only the Turn and keep reading stdin? | **No.** It writes one `error_during_execution` `result`, then exits with code 0 about 1 s later, with stdin still open. The partial Turn is kept for `--resume`. A still-queued message is lost. | Recorded (2.1.284) | +| Is a raw `control_request` `interrupt` honoured without the SDK's `initialize`? | **Yes.** `control_response` success with the receipt in 2 to 4 ms, then an `error_during_execution` `result`. The process stays alive, and the next user frame runs in the same Session with the partial Turn in context. `cancel_queued` works raw. **Round 2:** the same during a permission-bridge wait, and the bridge call gets MCP `notifications/cancelled`. | Recorded (2.1.284) | +| Minimum version for headless mid-Turn pickup and for `priority` | **Not named** in the changelog. Pickup between tool rounds is Recorded on 2.1.284. `priority` was not sent. | Documented (none found); Recorded (2.1.284) | ## Capability Table -| Stop or send on the raw `-p` stream | Turn ends with | Process | Partial text in context | Tool call in context | Foreground Bash tree | Queued message | -| -------------------------------------------------------- | ------------------------------------------------------------------------ | ------------------ | -------------------------- | ------------------------------------------------------- | -------------------- | ------------------------------------------------------- | -| SIGTERM (group) | No `result` | Exits 143 | No, not even on resume | Yes: `tool_use` plus `Exit code 137` and partial stdout | Killed by the CLI | Not tested | -| SIGINT (group or pid) | `result` `error_during_execution`, `aborted_streaming` / `aborted_tools` | Exits 0 after ~1 s | Yes, on resume | Yes, as a "rejected" result; partial stdout discarded | Killed | Lost; not replayed on resume | -| `control_request` `interrupt` | `control_response` receipt, then the same `result` | Stays alive | Yes, in the next Turn | Yes, as a "rejected" result; partial stdout discarded | Killed | Runs next by itself; listed in `still_queued` | -| `control_request` `interrupt` with `cancel_queued: true` | Same, with `cancelled: [uuid]` | Stays alive | Not tested (mid-tool only) | Same as above | Killed | Dropped; `command_lifecycle` `cancelled` | -| User frame during a tool call | Same Turn continues | Stays alive | n/a | n/a | Runs to completion | Delivered with the tool result; in `user_message_uuids` | -| User frame while text streams, no tool round left | First Turn ends normally | Stays alive | n/a | n/a | n/a | Runs next as its own Turn with its own `result` | +| Stop or send on the raw `-p` stream | Turn ends with | Process | Partial text in context | Tool call in context | Foreground Bash tree | Queued message | +| -------------------------------------------------------------- | ----------------------------------------------------------------------------------------- | ------------------ | -------------------------- | ------------------------------------------------------------------------------ | ------------------------------------------- | ---------------------------------------------------------------------------------------- | +| SIGTERM (group) | No `result` | Exits 143 | No, not even on resume | Yes: `tool_use` plus `Exit code 137` and partial stdout | Killed by the CLI | Not tested | +| SIGINT (group or pid) | `result` `error_during_execution`, `aborted_streaming` / `aborted_tools` | Exits 0 after ~1 s | Yes, on resume | Yes, as a "rejected" result; partial stdout discarded | Killed | Lost; not replayed on resume | +| `control_request` `interrupt` | `control_response` receipt, then the same `result` | Stays alive | Yes, in the next Turn | Yes, as a "rejected" result; partial stdout discarded | Killed | Runs next by itself; listed in `still_queued` | +| `control_request` `interrupt` with `cancel_queued: true` | Same, with `cancelled: [uuid]` | Stays alive | Not tested (mid-tool only) | Same as above | Killed | Dropped; `command_lifecycle` `cancelled` | +| User frame during a tool call | Same Turn continues | Stays alive | n/a | n/a | Runs to completion | Delivered with the tool result; in `user_message_uuids` | +| `interrupt` while a tool call waits on the permission bridge | Same `result`; `permission_denials` names the call; bridge gets `notifications/cancelled` | Stays alive | n/a | Yes, as a "rejected" result; the tool never ran | Never started; a late `allow` has no effect | With `cancel_queued`: dropped, as above | +| SIGTERM while a tool call waits on the permission bridge | No `result`; a second `approve` for the same call during shutdown | Exits 143 | n/a | Only on resume: a synthetic "Tool call interrupted… outcome is unknown" result | Never started | Not tested | +| User frame during an MCP call, or a round of two or more calls | Same Turn continues | Stays alive | n/a | n/a | Runs to completion | Delivered after the round's last tool result, not between calls; in `user_message_uuids` | +| User frame while a tool call waits on the permission bridge | Same Turn continues | Stays alive | n/a | n/a | Runs after approval | Held through approval and the tool; delivered with its result; in `user_message_uuids` | +| User frame while text streams, no tool round left | First Turn ends normally | Stays alive | n/a | n/a | n/a | Runs next as its own Turn with its own `result` | ## Still Unknown @@ -550,11 +847,16 @@ behaviour, and the SIGINT and SIGTERM results above are all Linux-only. non-zero. - Whether an interrupted mid-text Turn's tokens are counted anywhere. Its `result` reported `total_cost_usd: 0` and zero usage. -- Whether a user frame is picked up at a boundary between two tool calls when the - tool round has several calls, or with MCP or permission-bridge tools. Only one - foreground Bash call per round was used. -- What any stop does while a tool call waits on a permission prompt through - Secant's `--permission-prompt-tool` bridge. The driver used `--allowedTools`. +- What SIGINT does while a tool call waits on the permission bridge. Round 2 + tested only the raw `interrupt` and SIGTERM there. +- Whether the second `approve` call that SIGTERM set off (R2-A2, one run) always + happens, and whether a quick `allow` answer to it could start the tool before + the process exits. The stand-in held it. +- Whether a user frame is taken between calls when a round has a call the bridge + denies, or a call that fails. Round 2 allowed every call. +- Whether MCP tools without `readOnlyHint` run at the same time in one round. The + stand-in's `slow_echo` set it, and a Bash call plus an MCP call ran one after the + other. - How a SIGTERM that arrives while a user frame is queued treats that frame. It was not tested, but the SIGINT result suggests it is lost. **Inferred.** - Whether `CLAUDE_CODE_RESUME_INTERRUPTED_TURN=1` changes what a resume after From 12ad2562630b75283fc1a15b1d246d69af9dd837 Mon Sep 17 00:00:00 2001 From: rohan Date: Tue, 29 Sep 2026 11:45:05 +0530 Subject: [PATCH 04/11] docs(adr): pin ADR 0035's Steer boundary and permission-wait stop to the second live recording (#255) --- ...ly-the-turn-and-a-mid-turn-message-is-a-native-steer.md | 7 ++++--- 1 file changed, 4 insertions(+), 3 deletions(-) diff --git a/docs/adr/0035-interrupt-ends-only-the-turn-and-a-mid-turn-message-is-a-native-steer.md b/docs/adr/0035-interrupt-ends-only-the-turn-and-a-mid-turn-message-is-a-native-steer.md index 98423f7..8be8b08 100644 --- a/docs/adr/0035-interrupt-ends-only-the-turn-and-a-mid-turn-message-is-a-native-steer.md +++ b/docs/adr/0035-interrupt-ends-only-the-turn-and-a-mid-turn-message-is-a-native-steer.md @@ -25,8 +25,8 @@ while a Run waits after an Interrupt, the Run halts as any live Run does, and re Interactive agent steps alike, wherever the profile's Steer evidence says the Harness supports it. Codex delivers it through `turn/steer` with `expectedTurnId` and a Secant-minted `clientUserMessageId`, which comes back as the `clientId` of the `userMessage` item written when the text is taken into history before the next model request. Claude Code delivers it as a stdin `user` frame on the existing `claude -p --input-format stream-json` -process, stamped with a Secant-minted `uuid`: a frame written during a tool call is taken with that tool's result in the same exchange and listed in -`result.user_message_uuids`, and one written while text streams runs as the next native exchange with its own `result`; `command_lifecycle` frames, +process, stamped with a Secant-minted `uuid`: a frame written while a tool round runs, including one waiting on a tool approval, is taken with the round's +last tool result in the same exchange and listed in `result.user_message_uuids`, and one written while text streams runs as the next native exchange with its own `result`; `command_lifecycle` frames, advertised as `msg_lifecycle_v1`, say which frames are still queued. Delivery means the Harness put the text in front of the model, never that the model followed it. @@ -44,7 +44,8 @@ human's compose, so nothing runs after the human said stop. Codex discards pendi **How Claude Code stops.** Claude Code is interrupted with a raw `control_request` `interrupt` written to the same stdin, which, recorded on Claude Code 2.1.284, answers without the SDK's `initialize` in milliseconds, ends the Turn with a `result`, keeps the process and Session live for the next -frame, keeps the partial text in context, and kills the foreground tool tree. That makes Claude Code's interruption a confirmed active-Turn +frame, keeps the partial text in context, and kills the foreground tool tree. It holds while a tool waits on Secant's permission bridge too: Claude +Code cancels the pending approval call with an MCP `notifications/cancelled`, which expires the Harness Request, and the tool never runs. That makes Claude Code's interruption a confirmed active-Turn interruption rather than a process-only stop. Qualification bounds the dependency: a Claude Code that does not answer the request falls back to today's SIGTERM of the process tree and `--resume`, which loses partial streamed text. This is the same wire and the same degrade rule [ADR 0034](./0034-choose-and-change-model-and-effort-as-one-run-wide-model-choice.md) adopted for `set_model`; adopting the Agent SDK itself stays From a271c4aa44ab4cbda900d3639036cfccd89ac07b Mon Sep 17 00:00:00 2001 From: rohan Date: Tue, 29 Sep 2026 12:18:41 +0530 Subject: [PATCH 05/11] docs(research): record Claude Code SIGTERM re-approval and steer at a denied tool call on 2.1.284 (#255) --- ...e-code-live-interrupt-and-mid-turn-send.md | 274 ++++++++++++++++-- 1 file changed, 248 insertions(+), 26 deletions(-) diff --git a/docs/research/claude-code-live-interrupt-and-mid-turn-send.md b/docs/research/claude-code-live-interrupt-and-mid-turn-send.md index b924939..95589db 100644 --- a/docs/research/claude-code-live-interrupt-and-mid-turn-send.md +++ b/docs/research/claude-code-live-interrupt-and-mid-turn-send.md @@ -10,7 +10,9 @@ the Claude Code items left **Untested** or **Unknown** in [Harness Interrupt and Queued Messages](harness-interrupt-and-queued-messages.md) by running small live model sessions against the installed CLI. A second round, [Round 2](#round-2-permission-waits-and-multi-tool-rounds), repeats the key stops and -sends through a stand-in for Secant's MCP permission bridge. +sends through a stand-in for Secant's MCP permission bridge. A third round, +[Round 3](#round-3-sigterm-re-approval-and-denied-tool-calls), repeats the SIGTERM +during a permission wait and sends a user frame while the bridge denies a call. ## Answer @@ -93,12 +95,13 @@ SIGTERM ends the process and loses streamed text. stayed alive and the next user frame ran in the same Session. The model saw the call as "rejected before it could execute". The bridge's late `allow`, 30 s later, had no effect: the command never ran. -- **SIGTERM during a permission wait (one run).** The process exited 143 with no `result` +- **SIGTERM during a permission wait.** The process exited 143 with no `result` and no `tool_result`. No `notifications/cancelled` was sent. During shutdown, Claude Code opened a new MCP session to the bridge and sent a second `approve` call for the same `tool_use_id`. On `--resume`, the transcript gained a synthetic `tool_result`: "[Tool call interrupted: the session ended before this call's - result was recorded, so its outcome is unknown…]". + result was recorded, so its outcome is unknown…]". Round 3 repeated this; see + below. - **A user frame is still taken at a tool boundary, but only after the whole tool round.** This held when two MCP calls ran at once, when a Bash call and an MCP call ran one after the other in one round, and for a single MCP call. It also @@ -106,6 +109,29 @@ SIGTERM ends the process and loses streamed text. held through the approval and the command. Each time, the frame reached the model in the same Turn, and its `uuid` was in that `result`'s `user_message_uuids`. +**Round 3: SIGTERM always re-asks, and a denial is a tool boundary.** See +[Round 3](#round-3-sigterm-re-approval-and-denied-tool-calls). + +- **The second `approve` after SIGTERM came every time.** It came in six of six + runs, with the SIGTERM 0.5 s, 3 s, or 10 s into the wait, 29 to 47 ms after the + signal, on a new MCP session, with the same `tool_use_id`. With R2-A2 that is + seven of seven. +- **Answering it did not run the command.** A quick `allow` (two runs) or `deny` + (one run) was delivered in full, but the command's side-effect file never + appeared, no frame was written, and the process still exited 143 about 0.9 s + after the signal. On `--resume` the call got the same synthetic "outcome is + unknown" result whatever the answer. +- **A user frame written during a wait that ends in a denial is taken with the + denial, in the same Turn, and listed.** Two of two runs: the frame reached the + model next to the denial's `tool_result`, and its `uuid` was in the Turn's one + `result`, which also named the call in `permission_denials`. +- **A frame written as the denial lands can miss the Turn.** Written 8 ms before + or 33 ms after the denial's `tool_result` frame, it ran as the next Turn, in + three of three runs. +- **A denial does not end a round early.** With one call denied and one allowed in + the same round, the frame was taken only after the allowed call's result, in the + same Turn (one run). + Windows was not tested. ## Evidence Vocabulary @@ -814,29 +840,220 @@ In every run: - With `cancel_queued`, the queued frame got `cancelled` just before the `control_response`. +## Round 3: SIGTERM re-approval and denied tool calls + +Run on 2026-09-29 against the same Claude Code **2.1.284**. Every fact in this +section is Recorded (2.1.284) unless it carries another label. + +### Round 3 method + +The round 3 driver and stand-in bridge were the round 2 ones, with the same launch +flags, the same `default` permission mode, no `--allowedTools`, and `CLAUDE*` and +`ORCA_*` removed from the child environment. Two changes: + +- The stand-in's answer policy could depend on the call's position, so the first + `approve` for a Bash call could be held and a second one for the same call + answered differently. It also called back when an answer was returned, so the + driver could write a frame at that moment. +- The command was `./mark.sh`, which appends a timestamp to `mark.out` and prints + `marked`. The driver deleted `mark.out` before each run and checked for it + before the signal, 2 s after the exit, and after the resume. `mark.out` was + absent at every check in every round 3 run. + +The B prompts said "run exactly this command: ./mark.sh ." and Haiku sometimes sent +the command as `./mark.sh .`. It was denied every time, so this made no difference. + +### R3-A: SIGTERM while Bash waits on the bridge (six runs) + +Each run held the first Bash `approve` for 120 s, then sent a process-group SIGTERM +after a delay. A second `approve` for the same call was held, answered `allow` after +50 ms, or answered `deny` after 50 ms. The `allow` run with a 3 s delay: + +```text +t=5168 MCP approve#1 CALLED tool=Bash tool_use_id=toolu_01Ct… reqId=2 (held 120 s) +t=8211 mark.out before SIGTERM: ABSENT +t=8212 SIGNAL SIGTERM group +t=8216 HTTP approve#1 response stream closed +t=8242 HTTP POST /mcp server/discover, then initialize → new MCP session bb09799d +t=8257 MCP approve#2 CALLED tool=Bash tool_use_id=toolu_01Ct… reqId=1 (same call, again) +t=8307 MCP approve#2 RETURNING {"behavior":"allow","updatedInput":{"command":"./mark.sh",…}} +t=8309 HTTP approve#2 response written in full (writableFinished=true) +t=9119 EXIT code=143 signal=null +t=11123 mark.out 2s after exit: ABSENT +``` + +| Run | SIGTERM after the first `approve` | Second `approve` after SIGTERM | Answer to it | Exit after SIGTERM | `mark.out` | +| ---------- | --------------------------------- | ------------------------------ | -------------- | ------------------ | ---------- | +| a_hold_05 | 0.5 s | 47 ms | held | 1,306 ms, 143 | absent | +| a_hold_3 | 3 s | 37 ms | held | 933 ms, 143 | absent | +| a_hold_10 | 10 s | 33 ms | held | 920 ms, 143 | absent | +| a_allow_05 | 0.5 s | 29 ms | `allow`, 50 ms | 978 ms, 143 | absent | +| a_allow_3 | 3 s | 45 ms | `allow`, 50 ms | 907 ms, 143 | absent | +| a_deny_3 | 3 s | 38 ms | `deny`, 50 ms | 951 ms, 143 | absent | + +- **The second `approve` came in all six runs.** Each time Claude Code opened a new + MCP session (`server/discover`, then `initialize`) and called `approve` again with + the same `tool_use_id` and input, 29 to 47 ms after the signal. It happened at + every delay tried, 0.5 s to 10 s. With R2-A2 that makes seven of seven runs. +- **Answering it did not run the command.** An `allow` delivered 50 ms after the + call, 80 to 95 ms after the signal, did not start `./mark.sh`: `mark.out` never + appeared, and no `task_started`, `tool_result`, or other stream frame followed + the signal. The process exited 812 to 899 ms after the answer, about as long as + when the call was held. A `deny` changed nothing either. +- **No frame after the signal.** No stream frame of any kind was written after the + SIGTERM, and no `result`. +- **The answer is not recorded.** In all six transcripts nothing follows the + `tool_use` until the resume. The `--resume` process wrote the same synthetic + `tool_result` as in R2-A2, with `toolDenialKind: "interrupted"`, whether the second + `approve` was held, allowed, or denied. It made no `approve` call. Each resumed + model said the command did not finish and its outcome was unknown. +- Claude Code opens the second `approve` but does not act on its answer before it + exits. **Inferred** from the six runs. Whether a slower shutdown could act on it + is Unknown. + +### R3-B1: a user frame while a Bash call waits on a denial (two runs) + +The stand-in denied the Bash `approve` after 8 s. The driver wrote a user frame 3 s +into the wait. Run 1: + +```text +t=4952 assistant tool_use toolu_01XY… Bash {"command":"./mark.sh",…} +t=4977 MCP approve#1 CALLED tool=Bash -> will deny after 8000ms +t=7985 in user uuid=b3100000-…-01 "Additional instruction: when you reply, also say the word MANGO." +t=7987 command_lifecycle b3100000-… state=queued +t=12978 MCP approve#1 RETURNING {"behavior":"deny","message":"Denied by the stand-in bridge."} +t=12987 user tool_result toolu_01XY… "Denied by the stand-in bridge." is_error +t=13003 command_lifecycle b3100000-… state=started +t=14954 assistant text "PINEAPPLE MANGO" +t=14970 command_lifecycle b3100000-… state=completed +t=14974 result subtype=success terminal_reason=completed num_turns=2 result_index=0 + queued_turn_count=0 + permission_denials=[{"tool_name":"Bash","tool_use_id":"toolu_01XY…",…}] + user_message_uuids=[, "b3100000-0000-4000-8000-000000000001"] +t=14975 command_lifecycle state=completed +``` + +Run 2 gave the same sequence: `queued` 2 ms after the write, `started` 15 ms after +the denied `tool_result`, the reply "PINEAPPLE MANGO", and one `result` listing the +frame. + +- **Taken at the denial's boundary, in the same Turn, and listed.** The frame was + held through the rest of the wait. It reached the model next to the denial's + `tool_result`, and its `uuid` was in the Turn's one `result`. Haiku acted on it in + both runs. +- In the transcript, the `tool_result` is stored with `toolUseResult: "Error: Denied +by the stand-in bridge."` and `toolDenialKind: "permission-rule"`. The + `queued_command` attachment follows it, and the `queue-operation` `remove` has + `reason: "absorbed_mid_turn"`, as in rounds 1 and 2. +- The denial's `message` reaches the model as the `tool_result` text. + +### R3-B2: a user frame written as the denial is returned (three runs) + +The stand-in denied after 5 s. The driver wrote the frame from the stand-in's +return callback: at once (two runs) or 40 ms later (one run). + +```text +t=9663 MCP approve#1 RETURNING {"behavior":"deny",…} +t=9663 in user uuid=b3200000-…-01 "Additional instruction: … MANGO." +t=9671 user tool_result toolu_013H… "Denied by the stand-in bridge." is_error +t=9698 command_lifecycle b3200000-… state=queued +t=11285 assistant text "PINEAPPLE" +t=11307 result subtype=success num_turns=2 result_index=0 user_message_uuids=[] +t=11309 command_lifecycle state=completed +t=11310 command_lifecycle b3200000-… state=started +t=15623 assistant text "PINEAPPLE" +t=15673 result subtype=success num_turns=1 result_index=1 + user_message_uuids=["b3200000-0000-4000-8000-000000000001"] +``` + +| Run | Frame written, relative to the denied `tool_result` frame | `queued` after the write | Taken in the Turn? | +| ----- | --------------------------------------------------------- | ------------------------ | ------------------ | +| b2 | 8 ms before | 35 ms | No, next Turn | +| b2-r2 | 8 ms before | 31 ms | No, next Turn | +| b2d | 33 ms after | 2 ms | No, next Turn | + +- **Not taken in the running Turn.** In all three runs the frame was not in the + first `result`'s `user_message_uuids`. It ran by itself as the next Turn + (`result_index: 1`) with no new stdin write, as in test 4b. In each transcript + the `enqueue` comes after the denial's `tool_result`, with a `dequeue` for the + next Turn and no `queued_command` attachment. +- **The window closes very close to the `tool_result`.** A frame written 8 ms + before the `tool_result` frame was acknowledged `queued` 31 to 35 ms after the + write, which was after the pickup point. In R3-B1 the pickup came 15 to 16 ms + after the `tool_result`. So a frame that is not yet `queued` when the denial + lands can miss the Turn. **Inferred** from three runs; the exact cut-off was not + measured. +- In b2d the second Turn's model refused the instruction, saying it had already + answered. Delivery does not mean compliance. + +### R3-B3: one call denied, one allowed, in one round (one run) + +The prompt asked for Bash `./mark.sh` and `slow_echo("gamma", 5)` in one message. +The model first called `ToolSearch` for `slow_echo` in a round of its own. The +stand-in denied the Bash `approve` after 6 s and allowed `slow_echo` at once. The +driver wrote the frame 3 s into the Bash wait. + +```text +t=7098 assistant tool_use toolu_01V2… Bash {"command":"./mark.sh",…} +t=7118 MCP approve#1 CALLED tool=Bash -> will deny after 6000ms +t=7296 assistant tool_use toolu_01Qx… mcp__slowtools__slow_echo {"text":"gamma","delay_s":5} +t=10135 in user uuid=b3300000-…-01 "Additional instruction: … MANGO." +t=10137 command_lifecycle b3300000-… state=queued +t=13119 MCP approve#1 RETURNING {"behavior":"deny",…} +t=13129 user tool_result toolu_01V2… "Denied by the stand-in bridge." is_error +t=13142 MCP approve#2 CALLED tool=mcp__slowtools__slow_echo (allowed at once) +t=13152 MCP slow_echo#1 START gamma +t=18166 user tool_result toolu_01Qx… "echo: gamma" +t=18182 command_lifecycle b3300000-… state=started +t=22955 assistant text "PINEAPPLE" +t=22976 result subtype=success num_turns=4 result_index=0 + permission_denials=[{"tool_name":"Bash",…}] + user_message_uuids=[, "b3300000-0000-4000-8000-000000000001"] +``` + +- **A denied call does not end the round early.** The frame was not taken after the + denial. The `slow_echo` call's `approve` came 13 ms after the denial, and the + frame was taken 16 ms after the `slow_echo` result, the round's last. This is + the same whole-round rule as R2-B1. +- **Same Turn, listed.** One `result`, with the frame's `uuid` in + `user_message_uuids` and the denial in `permission_denials`. Haiku did not say + MANGO. + +### `command_lifecycle` in round 3 + +- A frame written during a wait got `queued` within 2 ms. A frame written within a + few milliseconds of the denial landing got it 31 to 35 ms later. +- A frame taken at a boundary got `started` 15 to 16 ms after the round's last + `tool_result`, `completed` 3 to 4 ms before the `result`, and the prompt's + `completed` 1 to 2 ms after it. +- A frame that missed the Turn got `started` 1 ms after the prompt's + `completed`, then ran as the next Turn. + ## The Four Original Unknowns -| Unknown in [Harness Interrupt and Queued Messages](harness-interrupt-and-queued-messages.md) | Result on 2.1.284 | Evidence | -| -------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------- | -| Does SIGTERM in `-p` mode keep the partial assistant text in the transcript that `--resume` loads? | **No for streamed text:** nothing of the unfinished model call is saved, not even thinking. **Yes for a finished tool round:** the `tool_use` and a real `Exit code 137` result with partial stdout are saved. There is no interrupted marker. Resume adds a synthetic `No response requested.`. **Round 2:** for a call waiting on the permission bridge nothing is saved after the `tool_use`; resume adds a synthetic "Tool call interrupted… outcome is unknown" `tool_result`. | Recorded (2.1.284) | -| Does SIGINT to a `-p --input-format stream-json` process end only the Turn and keep reading stdin? | **No.** It writes one `error_during_execution` `result`, then exits with code 0 about 1 s later, with stdin still open. The partial Turn is kept for `--resume`. A still-queued message is lost. | Recorded (2.1.284) | -| Is a raw `control_request` `interrupt` honoured without the SDK's `initialize`? | **Yes.** `control_response` success with the receipt in 2 to 4 ms, then an `error_during_execution` `result`. The process stays alive, and the next user frame runs in the same Session with the partial Turn in context. `cancel_queued` works raw. **Round 2:** the same during a permission-bridge wait, and the bridge call gets MCP `notifications/cancelled`. | Recorded (2.1.284) | -| Minimum version for headless mid-Turn pickup and for `priority` | **Not named** in the changelog. Pickup between tool rounds is Recorded on 2.1.284. `priority` was not sent. | Documented (none found); Recorded (2.1.284) | +| Unknown in [Harness Interrupt and Queued Messages](harness-interrupt-and-queued-messages.md) | Result on 2.1.284 | Evidence | +| -------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------- | +| Does SIGTERM in `-p` mode keep the partial assistant text in the transcript that `--resume` loads? | **No for streamed text:** nothing of the unfinished model call is saved, not even thinking. **Yes for a finished tool round:** the `tool_use` and a real `Exit code 137` result with partial stdout are saved. There is no interrupted marker. Resume adds a synthetic `No response requested.`. **Round 2:** for a call waiting on the permission bridge nothing is saved after the `tool_use`; resume adds a synthetic "Tool call interrupted… outcome is unknown" `tool_result`. **Round 3:** the same in six more runs, whether the bridge held, allowed, or denied the second `approve` that shutdown sends. | Recorded (2.1.284) | +| Does SIGINT to a `-p --input-format stream-json` process end only the Turn and keep reading stdin? | **No.** It writes one `error_during_execution` `result`, then exits with code 0 about 1 s later, with stdin still open. The partial Turn is kept for `--resume`. A still-queued message is lost. | Recorded (2.1.284) | +| Is a raw `control_request` `interrupt` honoured without the SDK's `initialize`? | **Yes.** `control_response` success with the receipt in 2 to 4 ms, then an `error_during_execution` `result`. The process stays alive, and the next user frame runs in the same Session with the partial Turn in context. `cancel_queued` works raw. **Round 2:** the same during a permission-bridge wait, and the bridge call gets MCP `notifications/cancelled`. | Recorded (2.1.284) | +| Minimum version for headless mid-Turn pickup and for `priority` | **Not named** in the changelog. Pickup between tool rounds is Recorded on 2.1.284. `priority` was not sent. | Documented (none found); Recorded (2.1.284) | ## Capability Table -| Stop or send on the raw `-p` stream | Turn ends with | Process | Partial text in context | Tool call in context | Foreground Bash tree | Queued message | -| -------------------------------------------------------------- | ----------------------------------------------------------------------------------------- | ------------------ | -------------------------- | ------------------------------------------------------------------------------ | ------------------------------------------- | ---------------------------------------------------------------------------------------- | -| SIGTERM (group) | No `result` | Exits 143 | No, not even on resume | Yes: `tool_use` plus `Exit code 137` and partial stdout | Killed by the CLI | Not tested | -| SIGINT (group or pid) | `result` `error_during_execution`, `aborted_streaming` / `aborted_tools` | Exits 0 after ~1 s | Yes, on resume | Yes, as a "rejected" result; partial stdout discarded | Killed | Lost; not replayed on resume | -| `control_request` `interrupt` | `control_response` receipt, then the same `result` | Stays alive | Yes, in the next Turn | Yes, as a "rejected" result; partial stdout discarded | Killed | Runs next by itself; listed in `still_queued` | -| `control_request` `interrupt` with `cancel_queued: true` | Same, with `cancelled: [uuid]` | Stays alive | Not tested (mid-tool only) | Same as above | Killed | Dropped; `command_lifecycle` `cancelled` | -| User frame during a tool call | Same Turn continues | Stays alive | n/a | n/a | Runs to completion | Delivered with the tool result; in `user_message_uuids` | -| `interrupt` while a tool call waits on the permission bridge | Same `result`; `permission_denials` names the call; bridge gets `notifications/cancelled` | Stays alive | n/a | Yes, as a "rejected" result; the tool never ran | Never started; a late `allow` has no effect | With `cancel_queued`: dropped, as above | -| SIGTERM while a tool call waits on the permission bridge | No `result`; a second `approve` for the same call during shutdown | Exits 143 | n/a | Only on resume: a synthetic "Tool call interrupted… outcome is unknown" result | Never started | Not tested | -| User frame during an MCP call, or a round of two or more calls | Same Turn continues | Stays alive | n/a | n/a | Runs to completion | Delivered after the round's last tool result, not between calls; in `user_message_uuids` | -| User frame while a tool call waits on the permission bridge | Same Turn continues | Stays alive | n/a | n/a | Runs after approval | Held through approval and the tool; delivered with its result; in `user_message_uuids` | -| User frame while text streams, no tool round left | First Turn ends normally | Stays alive | n/a | n/a | n/a | Runs next as its own Turn with its own `result` | +| Stop or send on the raw `-p` stream | Turn ends with | Process | Partial text in context | Tool call in context | Foreground Bash tree | Queued message | +| -------------------------------------------------------------- | ----------------------------------------------------------------------------------------- | ------------------ | -------------------------- | ------------------------------------------------------------------------------------- | ------------------------------------------------------------------ | ---------------------------------------------------------------------------------------- | +| SIGTERM (group) | No `result` | Exits 143 | No, not even on resume | Yes: `tool_use` plus `Exit code 137` and partial stdout | Killed by the CLI | Not tested | +| SIGINT (group or pid) | `result` `error_during_execution`, `aborted_streaming` / `aborted_tools` | Exits 0 after ~1 s | Yes, on resume | Yes, as a "rejected" result; partial stdout discarded | Killed | Lost; not replayed on resume | +| `control_request` `interrupt` | `control_response` receipt, then the same `result` | Stays alive | Yes, in the next Turn | Yes, as a "rejected" result; partial stdout discarded | Killed | Runs next by itself; listed in `still_queued` | +| `control_request` `interrupt` with `cancel_queued: true` | Same, with `cancelled: [uuid]` | Stays alive | Not tested (mid-tool only) | Same as above | Killed | Dropped; `command_lifecycle` `cancelled` | +| User frame during a tool call | Same Turn continues | Stays alive | n/a | n/a | Runs to completion | Delivered with the tool result; in `user_message_uuids` | +| `interrupt` while a tool call waits on the permission bridge | Same `result`; `permission_denials` names the call; bridge gets `notifications/cancelled` | Stays alive | n/a | Yes, as a "rejected" result; the tool never ran | Never started; a late `allow` has no effect | With `cancel_queued`: dropped, as above | +| SIGTERM while a tool call waits on the permission bridge | No `result`; a second `approve` for the same call during shutdown, in seven of seven runs | Exits 143 | n/a | Only on resume: a synthetic "Tool call interrupted… outcome is unknown" result | Never started, even when the second `approve` was answered `allow` | Not tested | +| User frame during an MCP call, or a round of two or more calls | Same Turn continues | Stays alive | n/a | n/a | Runs to completion | Delivered after the round's last tool result, not between calls; in `user_message_uuids` | +| User frame while a tool call waits on the permission bridge | Same Turn continues | Stays alive | n/a | n/a | Runs after approval | Held through approval and the tool; delivered with its result; in `user_message_uuids` | +| User frame while a tool call waits on a bridge denial | Same Turn continues | Stays alive | n/a | Yes, as an error `tool_result` with the denial message; `permission_denials` names it | Never started | Held through the wait; delivered with the denial's result; in `user_message_uuids` | +| User frame written as the bridge denial lands | First Turn ends normally | Stays alive | n/a | Same as above | Never started | Missed the boundary in three of three runs; runs next as its own Turn | +| User frame while text streams, no tool round left | First Turn ends normally | Stays alive | n/a | n/a | n/a | Runs next as its own Turn with its own `result` | ## Still Unknown @@ -849,11 +1066,16 @@ In every run: reported `total_cost_usd: 0` and zero usage. - What SIGINT does while a tool call waits on the permission bridge. Round 2 tested only the raw `interrupt` and SIGTERM there. -- Whether the second `approve` call that SIGTERM set off (R2-A2, one run) always - happens, and whether a quick `allow` answer to it could start the tool before - the process exits. The stand-in held it. -- Whether a user frame is taken between calls when a round has a call the bridge - denies, or a call that fails. Round 2 allowed every call. +- Why SIGTERM sends a second `approve` for the waiting call, and whether any + answer to it can take effect. It came in seven of seven runs. Round 3 answered it + within 50 ms in three runs and the command never ran, but a slower shutdown, or + an answer that lands at another moment, was not tested. +- Whether a user frame is taken between calls when a round has a call that fails + on its own, rather than being denied. Round 3 found that a denied call does not + end the round early (one run). +- The exact cut-off for a frame to be taken at a denial's boundary. Frames written + 5 s before the denial were taken; frames written 8 ms before or 33 ms after its + `tool_result` frame were not. Nothing in between was tried. - Whether MCP tools without `readOnlyHint` run at the same time in one round. The stand-in's `slow_echo` set it, and a Bash call plus an MCP call ran one after the other. From 32c8f72d7e0d615d8102cb2af83b54ce22965f4e Mon Sep 17 00:00:00 2001 From: rohan Date: Tue, 29 Sep 2026 12:27:41 +0530 Subject: [PATCH 06/11] docs(research): record Codex live leftover steer and empty turn/start (#255) Live codex-cli 0.157.1 app-server runs: empty-input turn/start is accepted on an idle thread and answers a leftover steer; a steer during a Stop hook is answered in the same Turn; the leftover window after the regular-task pending-input check is hit 6/22 by steering on the Stop hook's hook/completed; same-text re-delivery records the text twice. --- ...ive-leftover-steer-and-empty-turn-start.md | 252 ++++++++++++++++++ 1 file changed, 252 insertions(+) create mode 100644 docs/research/codex-live-leftover-steer-and-empty-turn-start.md diff --git a/docs/research/codex-live-leftover-steer-and-empty-turn-start.md b/docs/research/codex-live-leftover-steer-and-empty-turn-start.md new file mode 100644 index 0000000..ebf61d9 --- /dev/null +++ b/docs/research/codex-live-leftover-steer-and-empty-turn-start.md @@ -0,0 +1,252 @@ +# Codex Live Leftover Steer and Empty Turn Start + +Research date: 2026-09-29 + +Harness version examined: Codex **codex-cli 0.157.1** (installed standalone build, Linux x64, `codex --version` returned `codex-cli 0.157.1`), +source read at upstream tag `rust-v0.157.1`, commit +[`36650394c5b38c2990ccf2a3457165ca3e9d9726`](https://github.com/openai/codex/commit/36650394c5b38c2990ccf2a3457165ca3e9d9726). + +Ticket: [#255](https://github.com/secantdev/secant/issues/255). This note runs live `codex app-server` sessions to settle three questions that +[Harness Interrupt and Queued Messages](harness-interrupt-and-queued-messages.md) left open and that ADR 0035 (on `docs/interrupt-steer`) depends +on. ADR 0035 says that when Codex takes a steered text into history after its last model request and completes the Turn without answering it, +the Codex Adapter re-delivers that leftover with a native `turn/start`. It would use an empty input if Codex accepts one, and the same text +otherwise. + +## Answer + +**1. Codex accepts an empty-input `turn/start` on an idle thread, and the model answers the leftover.** `turn/start` with `input: []` returns a +normal `{ turn }`. Codex then emits `turn/started`, no `userMessage` item, one model response, and `turn/completed` `status: "completed"`. The +model responds to whatever the history already holds. With a leftover steer at the end of the history, it answered the steered question in all +4 of 4 runs, and `thread/read` then shows the steered text exactly once. An input of one empty text item is accepted too, but it records an +empty `userMessage` and the model returned an empty reply. While a Turn is active, an empty `turn/start` is refused with +`-32603 "failed to submit turn input: EmptyInput"`, and an empty `turn/steer` with `-32600 "input must not be empty"`.[^empty-src] + +**2. The leftover window was reproduced, but not during Stop hooks.** A `turn/steer` sent while a 12-second Stop hook runs is accepted, and Codex +answers it **in the same Turn** (2 of 2 runs). After the hooks, it drains the steer as a `userMessage` with the `clientId`, makes another model +request, runs the Stop hooks again, and only then completes. The source shows why: the regular task re-runs the turn loop whenever pending input +remains after it returns.[^regular-loop] The real window is the gap between that last pending-input check and task finish. A `turn/steer` +written the moment the Stop hook's `hook/completed` arrives hit it in **6 of 22** trials (the pilot's 1 of 2 plus 5 of 20 in the main run). In +those trials Codex accepted the steer, emitted its `userMessage` item with the matching `clientId`, and then sent `turn/completed` +`status: "completed"` with no model output after the item. The other 16 were refused with `-32600 "no active turn to steer"`. With a 1 to 3 ms +delay, all 4 trials were refused. + +**3. Re-delivering the same text writes it into history twice.** Given a leftover, `turn/start` with the same text (and the same +`clientUserMessageId`) is answered normally. `thread/read` then lists two `userMessage` items with the same text and the same `clientId`: one in +the leftover Turn and one in the re-delivery Turn. Asked how many times it had been asked the question, the model answered `2` (2 of 2). After an +empty re-delivery it answered `1` (3 of 3; the pilot run asked a different question). + +**4. Nothing experimental was used.** Every method, field, and notification in these runs is in the stable schema that the installed binary +generates (`codex app-server generate-json-schema`, without `--experimental`), and every session initialized with +`capabilities.experimentalApi: false`. `TurnStartParams.input` is a plain array with no `minItems`, both there and in Secant's recorded fixture +`tests/harness/fixtures/codex/codex-qualification/stable-schema.generated.json`. Only the test setup was out of the ordinary: the Stop hook was +added through the per-thread `config` override `bypass_hook_trust: true`, which Codex flags as dangerous (see Method). + +## Evidence Vocabulary + +- **Observed**: seen in these live runs, quoted from the raw frames. +- **Not observed**: looked for in these runs and absent. +- **Source-observed**: read in the Codex source at the commit above. +- **Untested**: implied by the source but not run. + +## Method + +A Python driver spawned `codex app-server` (stdio JSONL) in a throwaway working directory under the session scratchpad, not in the repository. +It sent Secant's handshake from `src/harness/codex/qualification.ts`: + +```json +{"method":"initialize","params":{"clientInfo":{"name":"secant","title":"Secant","version":"0.0.0-dev"},"capabilities":{"experimentalApi":false}}} +{"method":"initialized"} +``` + +It then sent `thread/start {cwd}` and `turn/start {threadId, input:[{type:"text",text}], model}` in the shape `src/harness/codex.ts` uses, and +`turn/steer {threadId, expectedTurnId, input, clientUserMessageId}`, the ADR 0035 shape. It logged every frame with a millisecond offset +(`t`, seconds since spawn below). Any server request would have been declined automatically. + +Differences from Secant's launch: + +- Every `turn/start` also set `effort: "low"` to keep cost down. Secant sets no effort. The model was `gpt-6-luna` ("Fast and affordable model + for easier tasks" in `model/list`), not the user's configured `gpt-6-sol`. +- The user's real `~/.codex` was the Codex home, unmodified, as in Secant's user-compatible launch. Its `hooks.json` registers + Orca command hooks, including `SessionStart`, `UserPromptSubmit`, and `Stop`. They ran on every Turn and appear in the frames. Its MCP servers + (`context7`, `codex_apps`, and two that failed to start) loaded too. +- To keep a Turn open during a Stop hook, experiment 2 started its threads with a per-thread config override: + + ```json + { + "bypass_hook_trust": true, + "hooks.Stop": [ + { "hooks": [{ "type": "command", "command": "sleep 12", "timeout": 60 }] } + ] + } + ``` + + The race variant used the same override with `"command": "true"`. Hooks from a config layer run only when they are trusted, and a + `bypass_hook_trust` request override lifts that gate for the session.[^hook-trust] Codex answered with + `configWarning` "`--dangerously-bypass-hook-trust` is enabled. Enabled hooks may run without review for this invocation." The hook ran + from the `sessionFlags` source (`"sourcePath": "//config.toml"`). No file in `~/.codex` was written by hand, and + `config.toml` and `hooks.json` kept their earlier modification times. Codex still wrote its own session rollouts, as any run does. + +- No temporary `CODEX_HOME` was needed, so none was created and `auth.json` was never touched. +- Every `codex app-server` the driver started exited when the driver closed stdin. The only `codex` processes left afterwards predate the runs: + the user's 0.158.0 app-server daemon, the VS Code extension, and an interactive `codex resume`. + +## Experiment 1: empty-input `turn/start` + +Each run used one thread and three Turns, one after another, each started after the previous `turn/completed`. + +**Observed.** T1 `"Remember the code word ZEBRA-42. Reply with just: OK"` completed with `OK`. T2 then sent `input: []`: + +```text +[4.748] out turn/start {"threadId":"…d1ad6","input":[],"model":"gpt-6-luna","effort":"low"} +[4.778] in response {"turn":{"id":"…d1d46","items":[],"status":"inProgress",…}} +[4.781] in turn/started {"turn":{"id":"…d1d46",…}} +[6.187] in item/completed agentMessage text="OK" +[6.224] in hook/started stop (user hooks.json) +[6.247] in turn/completed {"turn":{"id":"…d1d46","status":"completed","error":null,…}} +``` + +- **Not observed** in T2: a `userMessage` item, or a `UserPromptSubmit` hook run. Both appear in T1 and in T3. +- The model's reply repeated T1's answer: with no new input, it answered the last user message in history again. +- `thread/read` (`includeTurns: true`) lists T2 as `completed` with only `agentMessage "OK"`. + +T3 sent `input: [{"type":"text","text":""}]`. Codex recorded +`userMessage content=[{"type":"text","text":"","text_elements":[]}]`, ran the `UserPromptSubmit` hook, and the model returned +`agentMessage text=""`. `turn/completed` had `status: "completed"` and `items: []`. + +**Addendum, fresh thread (1 run).** An empty `turn/start` as the first Turn of a new thread was also accepted. With no user message in history, +the model made up a task from the injected context. It wrote a commentary message ("I'll check the current Codex docs and your local config +guidance first…"), called two `context7` MCP tools, tried a shell command (which failed), and gave a final answer about running Codex subagents in +parallel. + +**Observed, empty input during an active Turn** (experiment 2 mode `c`, sent while the Stop hook ran): + +```text +[7.440] turn/steer {"expectedTurnId":"…dd47","input":[]} -> {"error":{"code":-32600,"message":"input must not be empty"}} +[7.442] turn/start {"input":[]} -> {"error":{"code":-32603,"message":"failed to submit turn input: EmptyInput"}} +``` + +The Turn then completed normally with `ALPHA`. + +**Source-observed.** `start_or_steer` treats an empty `UserInput` as `has_explicit_input = false`. With no active Turn, it spawns the regular +task without pushing any input, so the model request is built from the existing history alone. With an active Turn, `steer_input` returns +`EmptyInput`, which `turn/start` reports as `failed to submit turn input: EmptyInput`.[^empty-src] + +## Experiment 2: reproducing a leftover steer + +### 2a. Steer during a Stop hook (answered in the same Turn) + +Prompt `"Reply with just: ALPHA"`. The driver waited for the session-flags Stop hook's `hook/started`, then 1 s, then sent +`turn/steer` with `clientUserMessageId: "secant-steer-n2"`. Run `n2` (run `n1` matched it): + +```text +[ 6.108] item/completed agentMessage text="ALPHA" +[ 6.144] hook/started stop sourcePath=//config.toml (sleep 12) +[ 7.146] turn/steer -> {"result":{"turnId":"…b256"}} +[18.146] hook/completed stop +[18.202] item/completed userMessage clientId="secant-steer-n2" "New question: what is 17+25? Reply with just the number." +[19.944] item/completed agentMessage text="42" +[19.977] hook/started stop (second run of the Stop hooks) +[31.985] hook/completed stop +[31.993] turn/completed {"turn":{"id":"…b256","status":"completed",…}} +``` + +**Observed.** A steer sent during a Stop hook is accepted, drained after the hooks, and answered within the same Turn, and the Stop hooks run +again. `thread/read` shows one Turn holding `userMessage ALPHA`, `agentMessage ALPHA`, `userMessage (clientId secant-steer-n2) 17+25`, +`agentMessage 42`. So a Stop hook does not open a leftover window at this version. The pending-input check that runs before the Stop hooks sees +nothing, but the regular task checks again after `run_turn` returns and re-enters it with the pending input.[^regular-loop] + +### 2b. Steer at the Stop hook's `hook/completed` (the leftover) + +The window left is the gap between that final check (`regular.rs` line 120) and `on_task_finished`. There Codex takes the active task and then +records any pending input to history, running `UserPromptSubmit` hooks on it and emitting its `userMessage` item.[^task-finish] To aim at it, the +driver used a fresh thread per trial with a no-op session-flags Stop hook (`true`). The reader thread wrote `turn/steer` as soon as it parsed +the first Stop `hook/completed` of the Turn (0 ms delay), or after a fixed delay. + +| Delay after Stop `hook/completed` | Trials | Leftover (accepted, item, no answer) | Refused `no active turn to steer` | Answered in the Turn | +| --------------------------------- | -----: | -----------------------------------: | --------------------------------: | -------------------: | +| 0 ms (pilot and main run) | 22 | 6 | 16 | 0 | +| 1 ms | 2 | 0 | 2 | 0 | +| 3 ms | 2 | 0 | 2 | 0 | + +Hit rate at 0 ms: 6 of 22 (27%). The steer response always came back within 1 to 2 ms. By the server's own `emittedAtMs` stamps, +`turn/completed` followed the Stop `hook/completed` by 3 to 6 ms in refused trials and by 43 to 58 ms in leftover trials. The extra time is +spent recording the leftover, including the user's `UserPromptSubmit` hook. + +**Observed, decisive frames** (trial `r0-0`, emitted in this order within about 4 ms): + +```json +{"method":"item/started","params":{"item":{"type":"userMessage","id":"01a0ebee-5b21-…","clientId":"secant-steer-r0-0","content":[{"type":"text","text":"New question: what is 17+25? Reply with just the number.","text_elements":[]}]},"turnId":"01a0ebee-3d03-…"}} +{"method":"item/completed","params":{"item":{"type":"userMessage","id":"01a0ebee-5b21-…","clientId":"secant-steer-r0-0",…},"turnId":"01a0ebee-3d03-…"}} +{"method":"turn/completed","params":{"turn":{"id":"01a0ebee-3d03-…","items":[{"type":"agentMessage","text":"OK","phase":"final_answer",…}],"itemsView":"summary","status":"completed","error":null,…}}} +``` + +- The `turn/steer` response was `{"result":{"turnId":"01a0ebee-3d03-…"}}`, the same success shape as a delivered steer. +- **Not observed** after the item: any `agentMessage`, any further model request, or any `error` notification. +- The `turn/completed` summary lists only the final `agentMessage`. It never lists `userMessage` items in any Turn. +- `thread/read` places the leftover `userMessage` inside the completed Turn, after `agentMessage "OK"`. +- In the pilot trial, the user's `UserPromptSubmit` hook ran between the steer and the leftover item. + +**Source-observed, not reproduced.** The regular task also returns early, without re-checking pending input, when the Turn has a terminal +error. Pending input then reaches `on_task_finished` the same way, which is the "failing Turn" case.[^regular-loop] Post-turn compaction runs +inside `run_turn` before it returns, so the same re-check covers it. That case is Untested. + +## Experiment 3: re-delivery of a leftover + +Each leftover in 2b was followed on the same thread by a re-delivery `turn/start`, then an account Turn. The account Turn asked: "Ignore any +AGENTS.md or environment context. Counting only messages I typed in this chat before this one: how many times did I ask you what 17+25 is? Reply +with just the number." After that came a `thread/read`. The pilot asked for the messages listed verbatim instead, and the model listed its +injected AGENTS.md context, so only its `thread/read` counts. + +### 3a. Empty input (4 leftovers: pilot, `r0-0`, `r0-14`, `r0-19`) + +```text +out turn/start {"threadId":"…20fa68","input":[],"model":"gpt-6-luna","effort":"low"} +in response {"turn":{"id":"01a0ebee-5b2d-…","status":"inProgress",…}} +in item/completed agentMessage text="42" phase=final_answer +in turn/completed {"turn":{"id":"01a0ebee-5b2d-…","status":"completed",…}} +account Turn -> "1" +thread/read: + TURN …3d03 completed: userMessage "Reply with just: OK" | agentMessage "OK" | userMessage clientId=secant-steer-r0-0 "New question: what is 17+25? …" + TURN …5b2d completed: agentMessage "42" +``` + +**Observed.** In 4 of 4 runs, the empty `turn/start` produced an answer to the leftover (`42`) with no new `userMessage` item. The steered text +appears once in history, and the model counted it once in 3 of 3 account Turns. + +### 3b. Same text and same `clientUserMessageId` (2 leftovers: `r0-2`, `r0-17`) + +```text +out turn/start {"input":[{"type":"text","text":"New question: what is 17+25? Reply with just the number."}],"clientUserMessageId":"secant-steer-r0-2",…} +in item/completed userMessage clientId="secant-steer-r0-2" "New question: what is 17+25? …" +in item/completed agentMessage text="42" +in turn/completed status=completed +account Turn -> "2" +thread/read: + TURN …8b66 completed: userMessage "Reply with just: OK" | agentMessage "OK" | userMessage clientId=secant-steer-r0-2 "New question: what is 17+25? …" + TURN …9b50 completed: userMessage clientId=secant-steer-r0-2 "New question: what is 17+25? …" | agentMessage "42" +``` + +**Observed.** In 2 of 2 runs, the same-text re-delivery is answered. History then holds the text twice, as two `userMessage` items with the same +`clientId` in two Turns, and Codex did not de-duplicate on `clientUserMessageId`. The model counted the question twice in both account Turns. + +## Unknowns and caveats + +- The race hit rate depends on this host and on the user's hooks. The Orca `UserPromptSubmit` hook lengthens the finishing path, and a client + that is not reacting to `hook/completed` will land in the window only by chance. The rate is not a property of Codex. +- The failing-Turn path (terminal error, then pending input recorded at task finish) was read in source and not reproduced. +- Only one model (`gpt-6-luna`, effort `low`) was used, with trivial prompts. Whether larger models answer the leftover as reliably after an + empty `turn/start` is Untested. In a longer history, an empty start makes the model respond to "whatever history holds", which was a + made-up task on a thread with no user message. +- Timestamps are the client's read times, rounded to 1 ms, except where `emittedAtMs` is named. +- The Stop hook in experiment 2 depended on `bypass_hook_trust` and the session-flags hook layer. Secant uses neither. They were only a way to + open the window on purpose. + +## Primary Sources + +[^empty-src]: OpenAI Codex source at `36650394`: [`core/src/session/turn_input.rs` lines 289-300](https://github.com/openai/codex/blob/36650394c5b38c2990ccf2a3457165ca3e9d9726/codex-rs/core/src/session/turn_input.rs#L289-L300) (`has_explicit_input`), [lines 356-362](https://github.com/openai/codex/blob/36650394c5b38c2990ccf2a3457165ca3e9d9726/codex-rs/core/src/session/turn_input.rs#L356-L362) (an empty input spawns the task with no pushed input), [lines 632-638 and 665-667](https://github.com/openai/codex/blob/36650394c5b38c2990ccf2a3457165ca3e9d9726/codex-rs/core/src/session/turn_input.rs#L632-L667) (`NoActiveTurn`, `EmptyInput` in `steer_input`); [`app-server/src/request_processors/turn_processor.rs` lines 596-620 and 672-682](https://github.com/openai/codex/blob/36650394c5b38c2990ccf2a3457165ca3e9d9726/codex-rs/app-server/src/request_processors/turn_processor.rs#L596-L682) (`turn/start` input mapping and `failed to submit turn input: {reason:?}`), [lines 1084-1085 and 1133-1134](https://github.com/openai/codex/blob/36650394c5b38c2990ccf2a3457165ca3e9d9726/codex-rs/app-server/src/request_processors/turn_processor.rs#L1084-L1134) (`turn/steer` errors `no active turn to steer`, `input must not be empty`). + +[^regular-loop]: OpenAI Codex source, [`core/src/tasks/regular.rs` lines 104-122](https://github.com/openai/codex/blob/36650394c5b38c2990ccf2a3457165ca3e9d9726/codex-rs/core/src/tasks/regular.rs#L104-L122) (loop over `run_turn`, early return on `terminal_error`, re-run while `has_pending_input`); [`core/src/session/turn.rs` lines 551-563 and 640-733](https://github.com/openai/codex/blob/36650394c5b38c2990ccf2a3457165ca3e9d9726/codex-rs/core/src/session/turn.rs#L551-L733) (post-sampling pending-input check, then Stop hooks and post-turn compaction before `break`). + +[^task-finish]: OpenAI Codex source, [`core/src/tasks/mod.rs` lines 621-684](https://github.com/openai/codex/blob/36650394c5b38c2990ccf2a3457165ca3e9d9726/codex-rs/core/src/tasks/mod.rs#L621-L684) (`on_task_finished` takes the task, takes pending input, and records it via `run_hooks_and_record_inputs`); [`core/src/session/turn.rs` lines 839-887](https://github.com/openai/codex/blob/36650394c5b38c2990ccf2a3457165ca3e9d9726/codex-rs/core/src/session/turn.rs#L839-L887) (`UserPromptSubmit` inspection, then `record_pending_input`). + +[^hook-trust]: OpenAI Codex source, [`app-server/src/config_manager.rs` lines 438-445](https://github.com/openai/codex/blob/36650394c5b38c2990ccf2a3457165ca3e9d9726/codex-rs/app-server/src/config_manager.rs#L438-L445) (`bypass_hook_trust` request override), [`hooks/src/engine/discovery.rs` lines 714-719 and 831](https://github.com/openai/codex/blob/36650394c5b38c2990ccf2a3457165ca3e9d9726/codex-rs/hooks/src/engine/discovery.rs#L714-L831) (only trusted, managed, or bypassed handlers run; `SessionFlags` hook source), [`config/src/hook_config.rs` lines 19-25 and 57-58](https://github.com/openai/codex/blob/36650394c5b38c2990ccf2a3457165ca3e9d9726/codex-rs/config/src/hook_config.rs#L19-L58) (`hooks.` and `hooks.state` in config), [`features/src/lib.rs` lines 1211-1216](https://github.com/openai/codex/blob/36650394c5b38c2990ccf2a3457165ca3e9d9726/codex-rs/features/src/lib.rs#L1211-L1216) (`hooks` feature stable, on by default). From 9dc6f21aa72ff95507a99917b74602abb79ff024 Mon Sep 17 00:00:00 2001 From: rohan Date: Tue, 29 Sep 2026 12:29:16 +0530 Subject: [PATCH 07/11] docs(adr): settle ADR 0035's Codex leftover re-delivery on an empty turn/start from the live recording (#255) --- ...urn-and-a-mid-turn-message-is-a-native-steer.md | 14 +++++++++----- 1 file changed, 9 insertions(+), 5 deletions(-) diff --git a/docs/adr/0035-interrupt-ends-only-the-turn-and-a-mid-turn-message-is-a-native-steer.md b/docs/adr/0035-interrupt-ends-only-the-turn-and-a-mid-turn-message-is-a-native-steer.md index 8be8b08..0ff939f 100644 --- a/docs/adr/0035-interrupt-ends-only-the-turn-and-a-mid-turn-message-is-a-native-steer.md +++ b/docs/adr/0035-interrupt-ends-only-the-turn-and-a-mid-turn-message-is-a-native-steer.md @@ -32,11 +32,15 @@ model followed it. **The Turn stretches until every Steer is delivered.** A Turn ends at the first Harness boundary after which no accepted Steer is pending, so one Secant Turn may span several native exchanges. The human's message therefore always reaches the agent before an Agent-declared completion from that Turn is -applied. Codex can take a steered text into history after its last model request, during stop hooks, post-Turn compaction, or a failing Turn, and then -complete without answering it; the Codex Adapter re-delivers that leftover with a native `turn/start` inside the same Secant Turn, using an empty -input if Codex accepts one and the same text otherwise, so a Steer is answered before its Turn ends on both Harnesses. That text is the human's own, so -this is ordinary turn-taking, not emulation. A Steer that arrives after the Turn has ended is rejected and its draft kept: in an Interactive agent -step the human sends it as the next Turn, and in an Agent step that has advanced the human is told it was not delivered. +applied. Codex can accept a Steer in the few milliseconds between its last check for pending input and the end of the Turn, or on a failing Turn, +write it into history, and complete without answering it; a Steer sent during its Stop hooks is answered, because pending input re-runs the Turn. +The Codex Adapter detects that leftover, a `userMessage` item with the Steer's `clientId` followed by `turn/completed` with no model output, and +re-delivers it inside the same Secant Turn with a native `turn/start` whose input is empty. Recorded on codex-cli 0.157.1, that start is accepted on +the idle thread and the model answers the leftover from history, which holds the text once. It is sent only on a detected leftover, since on a thread +with nothing pending it makes the model repeat itself or invent work. A Codex that refuses an empty input falls back to re-sending the same text, +which answers too but leaves the message in history twice. Either way the input is the human's own message, so a Steer is answered before its Turn +ends on both Harnesses and this is ordinary turn-taking, not emulation. A Steer that arrives after the Turn has ended is rejected and its draft +kept: in an Interactive agent step the human sends it as the next Turn, and in an Agent step that has advanced the human is told it was not delivered. **Interrupt drops what has not been delivered.** Every accepted Steer still pending when an Interrupt lands is dropped and its text returned to the human's compose, so nothing runs after the human said stop. Codex discards pending steered input on `turn/interrupt`; Claude Code is interrupted with From e3d43917335eb0d75da495f6b5f245d7d8d84038 Mon Sep 17 00:00:00 2001 From: rohan Date: Tue, 29 Sep 2026 12:29:42 +0530 Subject: [PATCH 08/11] docs(adr): rewrap ADR 0035 (#255) --- ...y-the-turn-and-a-mid-turn-message-is-a-native-steer.md | 8 ++++---- 1 file changed, 4 insertions(+), 4 deletions(-) diff --git a/docs/adr/0035-interrupt-ends-only-the-turn-and-a-mid-turn-message-is-a-native-steer.md b/docs/adr/0035-interrupt-ends-only-the-turn-and-a-mid-turn-message-is-a-native-steer.md index 0ff939f..81356ed 100644 --- a/docs/adr/0035-interrupt-ends-only-the-turn-and-a-mid-turn-message-is-a-native-steer.md +++ b/docs/adr/0035-interrupt-ends-only-the-turn-and-a-mid-turn-message-is-a-native-steer.md @@ -26,8 +26,8 @@ Interactive agent steps alike, wherever the profile's Steer evidence says the Ha `expectedTurnId` and a Secant-minted `clientUserMessageId`, which comes back as the `clientId` of the `userMessage` item written when the text is taken into history before the next model request. Claude Code delivers it as a stdin `user` frame on the existing `claude -p --input-format stream-json` process, stamped with a Secant-minted `uuid`: a frame written while a tool round runs, including one waiting on a tool approval, is taken with the round's -last tool result in the same exchange and listed in `result.user_message_uuids`, and one written while text streams runs as the next native exchange with its own `result`; `command_lifecycle` frames, -advertised as `msg_lifecycle_v1`, say which frames are still queued. Delivery means the Harness put the text in front of the model, never that the +last tool result in the same exchange and listed in `result.user_message_uuids`, and one written while text streams runs as the next native +exchange with its own `result`; `command_lifecycle` frames, advertised as `msg_lifecycle_v1`, say which frames are still queued. Delivery means the Harness put the text in front of the model, never that the model followed it. **The Turn stretches until every Steer is delivered.** A Turn ends at the first Harness boundary after which no accepted Steer is pending, so one Secant @@ -49,8 +49,8 @@ human's compose, so nothing runs after the human said stop. Codex discards pendi **How Claude Code stops.** Claude Code is interrupted with a raw `control_request` `interrupt` written to the same stdin, which, recorded on Claude Code 2.1.284, answers without the SDK's `initialize` in milliseconds, ends the Turn with a `result`, keeps the process and Session live for the next frame, keeps the partial text in context, and kills the foreground tool tree. It holds while a tool waits on Secant's permission bridge too: Claude -Code cancels the pending approval call with an MCP `notifications/cancelled`, which expires the Harness Request, and the tool never runs. That makes Claude Code's interruption a confirmed active-Turn -interruption rather than a process-only stop. Qualification bounds the dependency: a Claude Code that does not answer the request falls back to +Code cancels the pending approval call with an MCP `notifications/cancelled`, which expires the Harness Request, and the tool never runs. That +makes Claude Code's interruption a confirmed active-Turn interruption rather than a process-only stop. Qualification bounds the dependency: a Claude Code that does not answer the request falls back to today's SIGTERM of the process tree and `--resume`, which loses partial streamed text. This is the same wire and the same degrade rule [ADR 0034](./0034-choose-and-change-model-and-effort-as-one-run-wide-model-choice.md) adopted for `set_model`; adopting the Agent SDK itself stays rejected for the reasons in [Establish Claude Code's viable structured transports](https://github.com/secantdev/secant/issues/3). An Interrupt stops From 1174b550e6227f046b61f03dc2d4dc60af2b6796 Mon Sep 17 00:00:00 2001 From: rohan Date: Tue, 29 Sep 2026 12:30:03 +0530 Subject: [PATCH 09/11] docs(adr): rewrap ADR 0035 (#255) --- ...nd-a-mid-turn-message-is-a-native-steer.md | 87 ++++++++++--------- 1 file changed, 44 insertions(+), 43 deletions(-) diff --git a/docs/adr/0035-interrupt-ends-only-the-turn-and-a-mid-turn-message-is-a-native-steer.md b/docs/adr/0035-interrupt-ends-only-the-turn-and-a-mid-turn-message-is-a-native-steer.md index 81356ed..c1da853 100644 --- a/docs/adr/0035-interrupt-ends-only-the-turn-and-a-mid-turn-message-is-a-native-steer.md +++ b/docs/adr/0035-interrupt-ends-only-the-turn-and-a-mid-turn-message-is-a-native-steer.md @@ -5,42 +5,42 @@ message continues the same **Harness Session**. A message the human sends while natively at the **Harness**'s next boundary, and the Turn does not end until every Steer has been delivered. The need came from the pre-public-release reports ([#235](https://github.com/secantdev/secant/issues/235)): in the native clients Escape stops the Turn and the human types on in the same Session, and a message typed mid-Turn is picked up without interrupting, while Secant halted the Run on Interrupt, offered Claude Code no Steer, and -refused a mid-Turn send. This supersedes the Run lifecycle glossary's Interrupt rule (the Attempt ends `cancelled` and the Run rests `halted`) and -its Steer definition ("not a new Turn, and unsupported Harnesses do not emulate it"), [Spec: M3](https://github.com/secantdev/secant/issues/107) -stories 18 and 19 and its deferral of a Session-preserving Claude Code interrupt, and [Spec: M4](https://github.com/secantdev/secant/issues/137) -story 19's Codex-only Steer. It amends [ADR 0022](./0022-own-a-truthful-deep-harness-seam.md)'s `steer` and `interrupt` controls and Claude Code's -interruption mode. Unchanged: closing Secant or Ctrl+C halts the Run ([ADR 0019](./0019-failed-and-halted-runs-are-resumable-resting-states.md)), -Cancel is the only route to `cancelled`, an Agent call made in an interrupted Turn is dropped -([ADR 0032](./0032-let-opted-in-interactive-agent-steps-accept-agent-declared-completion.md)), and headless gains no Interrupt or Steer. +refused a mid-Turn send. This supersedes the Run lifecycle glossary's Interrupt rule (the Attempt ends `cancelled` and the Run rests `halted`) and its +Steer definition ("not a new Turn, and unsupported Harnesses do not emulate it"), [Spec: M3](https://github.com/secantdev/secant/issues/107) stories +18 and 19 and its deferral of a Session-preserving Claude Code interrupt, and [Spec: M4](https://github.com/secantdev/secant/issues/137) story 19's +Codex-only Steer. It amends [ADR 0022](./0022-own-a-truthful-deep-harness-seam.md)'s `steer` and `interrupt` controls and Claude Code's interruption +mode. Unchanged: closing Secant or Ctrl+C halts the Run ([ADR 0019](./0019-failed-and-halted-runs-are-resumable-resting-states.md)), Cancel is the +only route to `cancelled`, an Agent call made in an interrupted Turn is dropped ([ADR +0032](./0032-let-opted-in-interactive-agent-steps-accept-agent-declared-completion.md)), and headless gains no Interrupt or Steer. **Interrupt, per Step kind.** In an **Interactive agent step** the interrupted Turn ends `interrupted` and the Run returns to `blocked` on the human, whose next message goes into the same Session; no Attempt is published and the Iteration stays open. An interrupted **Entry Turn** is treated the same -way and is never re-sent. In an **Agent step** the interrupted Turn ends `interrupted`, the Attempt stays open, and the Run is `blocked` on the human's -message. That message starts a human-origin follow-up Turn in the same Session and the same Attempt, and the Attempt takes its outcome from its last -Turn: a clean end completes the Step and the Routing advances without the human, exactly as an uninterrupted Agent step does, and a failure takes the -Step's ordinary retry policy. A second Interrupt is the only way to hold the Step again. A `lost` Turn keeps its current meaning. When Secant closes -while a Run waits after an Interrupt, the Run halts as any live Run does, and resuming returns it to waiting on the human's message. +way and is never re-sent. In an **Agent step** the interrupted Turn ends `interrupted`, the Attempt stays open, and the Run is `blocked` on the +human's message. That message starts a human-origin follow-up Turn in the same Session and the same Attempt, and the Attempt takes its outcome from +its last Turn: a clean end completes the Step and the Routing advances without the human, exactly as an uninterrupted Agent step does, and a failure +takes the Step's ordinary retry policy. A second Interrupt is the only way to hold the Step again. A `lost` Turn keeps its current meaning. When +Secant closes while a Run waits after an Interrupt, the Run halts as any live Run does, and resuming returns it to waiting on the human's message. **Steer.** Steer now means native mid-Turn delivery of the human's own text, landing at the Harness's next boundary; it is offered in Agent steps and Interactive agent steps alike, wherever the profile's Steer evidence says the Harness supports it. Codex delivers it through `turn/steer` with -`expectedTurnId` and a Secant-minted `clientUserMessageId`, which comes back as the `clientId` of the `userMessage` item written when the text is taken -into history before the next model request. Claude Code delivers it as a stdin `user` frame on the existing `claude -p --input-format stream-json` -process, stamped with a Secant-minted `uuid`: a frame written while a tool round runs, including one waiting on a tool approval, is taken with the round's -last tool result in the same exchange and listed in `result.user_message_uuids`, and one written while text streams runs as the next native -exchange with its own `result`; `command_lifecycle` frames, advertised as `msg_lifecycle_v1`, say which frames are still queued. Delivery means the Harness put the text in front of the model, never that the -model followed it. +`expectedTurnId` and a Secant-minted `clientUserMessageId`, which comes back as the `clientId` of the `userMessage` item written when the text is +taken into history before the next model request. Claude Code delivers it as a stdin `user` frame on the existing `claude -p --input-format +stream-json` process, stamped with a Secant-minted `uuid`: a frame written while a tool round runs, including one waiting on a tool approval, is taken +with the round's last tool result in the same exchange and listed in `result.user_message_uuids`, and one written while text streams runs as the next +native exchange with its own `result`; `command_lifecycle` frames, advertised as `msg_lifecycle_v1`, say which frames are still queued. Delivery means +the Harness put the text in front of the model, never that the model followed it. -**The Turn stretches until every Steer is delivered.** A Turn ends at the first Harness boundary after which no accepted Steer is pending, so one Secant -Turn may span several native exchanges. The human's message therefore always reaches the agent before an Agent-declared completion from that Turn is -applied. Codex can accept a Steer in the few milliseconds between its last check for pending input and the end of the Turn, or on a failing Turn, -write it into history, and complete without answering it; a Steer sent during its Stop hooks is answered, because pending input re-runs the Turn. -The Codex Adapter detects that leftover, a `userMessage` item with the Steer's `clientId` followed by `turn/completed` with no model output, and +**The Turn stretches until every Steer is delivered.** A Turn ends at the first Harness boundary after which no accepted Steer is pending, so one +Secant Turn may span several native exchanges. The human's message therefore always reaches the agent before an Agent-declared completion from that +Turn is applied. Codex can accept a Steer in the few milliseconds between its last check for pending input and the end of the Turn, or on a failing +Turn, write it into history, and complete without answering it; a Steer sent during its Stop hooks is answered, because pending input re-runs the +Turn. The Codex Adapter detects that leftover, a `userMessage` item with the Steer's `clientId` followed by `turn/completed` with no model output, and re-delivers it inside the same Secant Turn with a native `turn/start` whose input is empty. Recorded on codex-cli 0.157.1, that start is accepted on the idle thread and the model answers the leftover from history, which holds the text once. It is sent only on a detected leftover, since on a thread with nothing pending it makes the model repeat itself or invent work. A Codex that refuses an empty input falls back to re-sending the same text, which answers too but leaves the message in history twice. Either way the input is the human's own message, so a Steer is answered before its Turn -ends on both Harnesses and this is ordinary turn-taking, not emulation. A Steer that arrives after the Turn has ended is rejected and its draft -kept: in an Interactive agent step the human sends it as the next Turn, and in an Agent step that has advanced the human is told it was not delivered. +ends on both Harnesses and this is ordinary turn-taking, not emulation. A Steer that arrives after the Turn has ended is rejected and its draft kept: +in an Interactive agent step the human sends it as the next Turn, and in an Agent step that has advanced the human is told it was not delivered. **Interrupt drops what has not been delivered.** Every accepted Steer still pending when an Interrupt lands is dropped and its text returned to the human's compose, so nothing runs after the human said stop. Codex discards pending steered input on `turn/interrupt`; Claude Code is interrupted with @@ -49,18 +49,19 @@ human's compose, so nothing runs after the human said stop. Codex discards pendi **How Claude Code stops.** Claude Code is interrupted with a raw `control_request` `interrupt` written to the same stdin, which, recorded on Claude Code 2.1.284, answers without the SDK's `initialize` in milliseconds, ends the Turn with a `result`, keeps the process and Session live for the next frame, keeps the partial text in context, and kills the foreground tool tree. It holds while a tool waits on Secant's permission bridge too: Claude -Code cancels the pending approval call with an MCP `notifications/cancelled`, which expires the Harness Request, and the tool never runs. That -makes Claude Code's interruption a confirmed active-Turn interruption rather than a process-only stop. Qualification bounds the dependency: a Claude Code that does not answer the request falls back to -today's SIGTERM of the process tree and `--resume`, which loses partial streamed text. This is the same wire and the same degrade rule -[ADR 0034](./0034-choose-and-change-model-and-effort-as-one-run-wide-model-choice.md) adopted for `set_model`; adopting the Agent SDK itself stays -rejected for the reasons in [Establish Claude Code's viable structured transports](https://github.com/secantdev/secant/issues/3). An Interrupt stops -the Turn's foreground tool work; a shell the model moved to the background may outlive it until the Session closes. +Code cancels the pending approval call with an MCP `notifications/cancelled`, which expires the Harness Request, and the tool never runs. That makes +Claude Code's interruption a confirmed active-Turn interruption rather than a process-only stop. Qualification bounds the dependency: a Claude Code +that does not answer the request falls back to today's SIGTERM of the process tree and `--resume`, which loses partial streamed text. This is the same +wire and the same degrade rule [ADR 0034](./0034-choose-and-change-model-and-effort-as-one-run-wide-model-choice.md) adopted for `set_model`; adopting +the Agent SDK itself stays rejected for the reasons in [Establish Claude Code's viable structured +transports](https://github.com/secantdev/secant/issues/3). An Interrupt stops the Turn's foreground tool work; a shell the model moved to the +background may outlive it until the Session closes. -**Recording.** Each Steer is durable against its Turn: the human's text, when it was sent, and how it settled, as delivered within the Turn, -delivered after a native boundary, re-delivered, or dropped by an Interrupt or a lost Turn. An interrupted Turn records the stop used, the control -request or the process stop. A follow-up Turn after an Interrupt is a human-origin Turn inside the Agent step's Attempt. Steered text is shown as the -human's message inside its Turn, durable content kept distinct from live previews under -[ADR 0024](./0024-use-one-deep-projection-port-for-tui-and-headless-clients.md). +**Recording.** Each Steer is durable against its Turn: the human's text, when it was sent, and how it settled, as delivered within the Turn, delivered +after a native boundary, re-delivered, or dropped by an Interrupt or a lost Turn. An interrupted Turn records the stop used, the control request or +the process stop. A follow-up Turn after an Interrupt is a human-origin Turn inside the Agent step's Attempt. Steered text is shown as the human's +message inside its Turn, durable content kept distinct from live previews under [ADR +0024](./0024-use-one-deep-projection-port-for-tui-and-headless-clients.md). **Harness Interface.** `steer` keeps its `ControlReceipt`, whose acceptance now means the text was handed to the Harness. The Turn's closed event stream gains one Steer lifecycle, delivered or dropped, beside the Harness Request lifecycle, and on terminal the Adapter publishes a dropped event @@ -71,10 +72,10 @@ request. Rejected: keeping Interrupt as a halt, which turns stopping one's own conversation into a Run-level stop outside the workflow's logic; resending an interrupted Agent step's prompt on resume; an interrupted Agent step that waits for the human to end it, which silently turns it into an Interactive -agent step; a client-side queue sent as the next Turn, which is emulation and would race the one live Turn per Prepared Harness and the -Agent-declared completion's clean-end settlement; a separate queue key beside Steer; counting Claude Code's second native exchange as a new Turn the -Harness started, which would let an Agent-declared completion settle the Step while the human's message is still being answered; reporting a Codex -leftover upward as missed, which pushes a native race into two Step kinds; SIGINT, which exits the process and silently loses queued messages; and -`queued_turn_count`, which stayed 0 while a message waited. How the Run Workbench presents Interrupt, Steer, a dropped Steer's restored draft, and -delivered and dropped Steers is left to the Run Workbench prototype. The decision was made on -[Decide how the human interrupts a Turn and sends a message while a Turn is working](https://github.com/secantdev/secant/issues/255). +agent step; a client-side queue sent as the next Turn, which is emulation and would race the one live Turn per Prepared Harness and the Agent-declared +completion's clean-end settlement; a separate queue key beside Steer; counting Claude Code's second native exchange as a new Turn the Harness started, +which would let an Agent-declared completion settle the Step while the human's message is still being answered; reporting a Codex leftover upward as +missed, which pushes a native race into two Step kinds; SIGINT, which exits the process and silently loses queued messages; and `queued_turn_count`, +which stayed 0 while a message waited. How the Run Workbench presents Interrupt, Steer, a dropped Steer's restored draft, and delivered and dropped +Steers is left to the Run Workbench prototype. The decision was made on [Decide how the human interrupts a Turn and sends a message while a Turn is +working](https://github.com/secantdev/secant/issues/255). From db84242c859eef60f766b1d9a42ed38f264fc075 Mon Sep 17 00:00:00 2001 From: rohan Date: Tue, 29 Sep 2026 13:32:17 +0530 Subject: [PATCH 10/11] docs(research): record Windows live Claude Code interrupt and steer and Codex controls (#255) --- .../windows-live-interrupt-and-steer.md | 634 ++++++++++++++++++ 1 file changed, 634 insertions(+) create mode 100644 docs/research/windows-live-interrupt-and-steer.md diff --git a/docs/research/windows-live-interrupt-and-steer.md b/docs/research/windows-live-interrupt-and-steer.md new file mode 100644 index 0000000..9139e29 --- /dev/null +++ b/docs/research/windows-live-interrupt-and-steer.md @@ -0,0 +1,634 @@ +# Windows Live Interrupt and Steer + +Research date: 2026-09-29 + +Harness versions examined, on Windows 11 Home 10.0.26200 (x64): + +- Claude Code **2.1.283**, the WinGet native executable (`claude --version` returned + `2.1.283 (Claude Code)`), resolved on PATH as + `%LOCALAPPDATA%\Microsoft\WinGet\Links\claude.exe`, a symbolic link into the WinGet package. +- Codex **codex-cli 0.155.0**, the standalone build at + `%LOCALAPPDATA%\Programs\OpenAI\Codex\bin\codex.exe`. + +Ticket: [#255](https://github.com/secantdev/secant/issues/255). The Linux rounds in +[Claude Code Live Interrupt and Mid-Turn Send](claude-code-live-interrupt-and-mid-turn-send.md) +(on `research/255-claude-live-interrupt`, Claude Code 2.1.284) and +[Codex Live Leftover Steer and Empty Turn Start](codex-live-leftover-steer-and-empty-turn-start.md) +(codex-cli 0.157.1) left Windows untested. This note repeats the stops and sends that +[ADR 0035](../adr/0035-interrupt-ends-only-the-turn-and-a-mid-turn-message-is-a-native-steer.md) +relies on, against the installed Windows CLIs. The Windows builds are one patch release (Claude +Code) and two minor releases (Codex) older than the Linux ones. + +## Answer + +**W1. The raw `control_request` `interrupt` works on Windows without `initialize`. The Turn +stops and the process and Session stay alive. But it does not kill a script that the Bash tool +runs.** The `control_response` came back 6 to 30 ms after the write, in 12 of 12 runs, with the +receipt `{"still_queued":[…]}`. The Turn then ended with the same `result` as on Linux: +`subtype: "error_during_execution"`, `terminal_reason` `aborted_tools` or `aborted_streaming`. The +next stdin user frame ran as `result_index: 1` in the same process and `session_id`. Mid-text, +the partial text stayed in context (`aborted: true`, `isAbortedMidStream: true`), and the model +quoted its last partial line. Mid-tool, the `tool_result` was the standard "The user doesn't want +to proceed with this tool use… rejected" text, and the model said the command "never executed". +To stop the tool, Claude Code runs `taskkill /PID /T /F` (seen in the process +table). That reached a native command run directly by the Bash tool (`ping`) and a PowerShell-tool +script, and both were killed. It did not reach `./slowjob.sh` run by the Bash tool, which runs +through Git Bash. The script's `bash.exe` had a Windows parent pid that no longer existed, so it +was outside the tree. It ran to completion in all 6 runs where the raw interrupt stopped it, and in +5 of them it was still running after Claude Code had exited. A backgrounded Bash task was not stopped by the interrupt. Claude Code +reported it `killed` when the process exited, but its script also ran on. + +**W2. `cancel_queued: true` behaves as on Linux.** The queued frame was listed in `cancelled`, +with a `command_lifecycle` `cancelled` frame just before the `control_response`. No Turn ran in +the next 10 s, and the process stayed alive (2 of 2 runs). The transcript holds `enqueue`, then +`remove`, and no `queued_command` attachment. A plain `interrupt` listed the frame in +`still_queued` and ran it by itself as the next Turn, 10 ms after the interrupted `result` +(1 run). + +**W3. Mid-Turn pickup behaves as on Linux.** A frame written during a Bash call was held until +the call finished. It then reached the model as a `queued_command` attachment, next to the tool +result, in the same Turn. Its `uuid` was in the one `result`'s `user_message_uuids`, `num_turns` +was 2, and the transcript's `remove` had `reason: "absorbed_mid_turn"` (1 run). A frame written +while text streamed ran as the next native exchange, with its own `system/init` and `result` +(`result_index: 1`) (1 run). `system/init.capabilities` advertised `msg_lifecycle_v1`, and +`command_lifecycle` frames (`queued`, `started`, `completed`, `cancelled`) came on the raw stream +in every run. `queued_turn_count` was `0` on every `result`. + +**W4. An interrupt during a permission-bridge wait behaves as on Linux.** The stand-in bridge +held the `approve` answer for 30 s. After the interrupt, the stand-in got MCP +`notifications/cancelled` for the pending call (`"reason":"AbortError: remote-cancel"`) 3 ms after +the `control_response`, and its handler's abort signal fired. The `result` named the +call in `permission_denials`. The process stayed alive and the next frame ran in the same +Session. The late `allow` had no effect: the script never started (2 of 2 runs, one with +`cancel_queued`). + +**W5. Secant's Windows stop today (`taskkill /pid /T /F`) kills Claude Code outright. +The Bash tool's script still survives.** Claude Code exited with code 1, 29 to 41 ms after +`taskkill` returned, and wrote no `result` and no frame after the kill. Mid-text, nothing of the +partial answer was saved. On `--resume`, Claude Code added a synthetic `No response requested.`, +and the resumed model said it had "declined" the task, as after a Linux SIGTERM. Mid-tool, the +transcript ended at the `tool_use`. `--resume` added a synthetic `tool_result`, "[Tool call +interrupted: the session ended before this call's result was recorded, so its outcome is +unknown…]" (`toolDenialKind: "interrupted"`), plus the synthetic reply. That is the Linux result +for a SIGTERM during a bridge wait, not for a SIGTERM during a running Bash call. `taskkill /T` +listed the Claude Code process, its two tool shells, and their consoles as killed. The script and +its `ping` were not in that list, and the script finished 24 s after the kill (1 run each). + +**W6. Codex on Windows behaves as on Linux.** `turn/interrupt` answered `{}` in 92 ms, and +`turn/completed` `status: "interrupted"` followed 3 ms later. The same thread took the next +`turn/start` (2 of 2). A `turn/steer` with `clientUserMessageId` was accepted at once while text +streamed. It was taken into history only after the whole answer, as a `userMessage` item whose +`clientId` echoed the id, and it was answered in the same Turn (2 of 2). An empty-input +`turn/start` on an idle thread after a completed Turn was accepted. It wrote no `userMessage` +item, and the model answered from history by repeating its last reply (2 of 2). The leftover race +was not attempted. + +## Evidence Vocabulary + +- **Observed**: seen in the live frames, process tables, marker files, or session transcripts + captured for this note on 2026-09-29 against the versions above. Every fact below without + another label is Observed. +- **Source-observed**: read in the Secant source at this branch's base, `docs/interrupt-steer` + (`1174b55`). +- **Inferred**: a consequence drawn from Observed facts that needs a check before it becomes a + compatibility promise. +- **Not observed**: looked for in these runs and absent. + +## Method + +### Driver + +A small Bun 1.4.2 driver spawned each CLI the way `src/process/process.ts` spawns an owned +process on win32: `node:child_process` `spawn` with `overlapped` stdio pipes, +`windowsHide: true`, and `detached: false`, since Windows has no process groups +(**Source-observed**, `spawnOwnedProcessWithNode`, lines 399-414). The driver: + +- wrote stdin frames and logged every stdout line with a millisecond offset from spawn (`t` + below), and logged the exit code and signal; +- read the whole Windows process table (`Get-CimInstance Win32_Process`: pid, parent pid, name, + command line) before and after each stop, and walked Claude Code's descendants by parent pid; +- copied the session transcript from `%USERPROFILE%\.claude\projects\\.jsonl` + after each run. + +Each run used its own throwaway directory under the session scratchpad as cwd. Scripts and raw +logs stayed there, outside the repository. + +### Claude Code launch + +The flags were those of `src/harness/claude-code.ts` (**Source-observed**, lines 764-780): + +```text +claude.exe -p --input-format stream-json --output-format stream-json --verbose \ + --include-partial-messages --model haiku --session-id (or --resume ) +``` + +- `system/init` reported `permissionMode: "default"` and + `capabilities: ["interrupt_receipt_v1","interrupt_cancel_queued_v1","msg_lifecycle_v1","mcp_read_resource_v1","mcp_tool_ui_meta_v1"]` + in every run. The tool list included both `Bash` and `PowerShell`. +- Every `CLAUDE*` and `ORCA_*` variable was removed from the child environment. The host's user + settings still load Orca hook commands (`claude-hook.cmd || echo {}`), a user `CLAUDE.md`, and + the user's MCP servers, as they would under Secant. +- W1 to W3 and W5 pre-allowed only the slow script, with `--allowedTools "Bash(./slowjob.sh)" +"Bash(./slowjob.sh:*)"`. The PowerShell and direct-`ping` variants pre-allowed only their own + command. W4 passed no `--allowedTools`. +- User frames were in the Adapter's shape plus a `uuid`: + `{"type":"user","message":{"role":"user","content":""},"parent_tool_use_id":null,"uuid":""}`. + The interrupt was `{"type":"control_request","request_id":"int1","request":{"subtype":"interrupt"}}`, + with `"cancel_queued":true` added in W2. No `initialize` was ever sent. + +### The slow tool + +`slowjob.sh` waits about 27 s on a native Windows child and leaves marker files, so that its +survival shows as a side effect: + +```bash +#!/bin/bash +echo started-27 +date +%s%3N > started.marker +ping -n 28 127.0.0.1 > /dev/null +echo done-27 +date +%s%3N > done.marker +``` + +The prompt asked for `./slowjob.sh` in the foreground with the Bash tool, then the word +`PINEAPPLE`. Claude Code's Bash tool ran it through Git for Windows: `Git\bin\bash.exe -c "source +…shell-snapshots\snapshot-bash-….sh …"`, then `Git\usr\bin\bash.exe`, then the script's own +`bash.exe`. Two variants checked other tool paths: + +- `slowjob.ps1`, the same steps in PowerShell, run by the PowerShell tool + (`cmd.exe /d /s /c "chcp 65001 & pwsh.exe -NoProfile -NonInteractive … -Command …"`); +- `ping -n 28 127.0.0.1` typed directly as the Bash tool's command. + +In one early W1a run, Haiku sent the command as `./slowjob.sh .`, which the pre-allow rule did not +match. Claude Code denied it ("This command requires approval"), and the interrupt landed while +the model streamed its next step (`terminal_reason: aborted_streaming`). That run counts toward +the receipt timings only. The prompt was then reworded to name the command in backticks with "no +arguments". + +The mid-text prompt was "count from one to three hundred in English words, one number per line". +The stop or send came 2.5 s after the first `text_delta`, or 3 to 5 s after the `tool_use` frame. +After each stop, the same process (or a `--resume` process) was asked, "Without running any +tools: what was the last thing you wrote or did…? Quote the last line you wrote…". + +### Permission-bridge stand-in (W4) + +The stand-in copied `src/harness/permission-bridge.ts` (**Source-observed**): a Streamable HTTP +MCP server on `127.0.0.1` with a random port and a 256-bit bearer token, one transport per MCP +session, built on the repository's `@modelcontextprotocol/sdk` 1.29.0. It served +`secant-permissions` with one `approve` tool (`tool_name`, `input`, `tool_use_id`) returning +`{"behavior":"allow","updatedInput":…}` as JSON text. The launch fragment was +`--mcp-config --permission-prompt-tool mcp__secant-permissions__approve`. Unlike +Secant, it held each answer for 30 s, did not stop when cancelled, and logged every HTTP request +body and the handler's abort signal. + +### Secant's Windows stop (W5) + +Secant's process Module has no graceful stage on Windows. `interrupt` calls `killGroup`, which +spawns `taskkill /pid /T /F` at once and reports `escalated: true` for a live child +(**Source-observed**, `safeInterrupt` lines 561-595, `killGroup` lines 756-776, and +[`src/process/AGENTS.md`](../../src/process/AGENTS.md)). The Claude Code Adapter then settles the +Turn `lost` with `interruption-unknown` (`src/harness/claude-code.ts` lines 545-580), and the +profile's Windows interruption evidence says so (lines 1542-1548). The driver ran the same +`taskkill` command against the Claude Code pid. + +### Codex launch (W6) + +The driver spawned `codex.exe app-server` over JSONL with the handshake of +`src/harness/codex/qualification.ts`: + +```json +{"id":1,"method":"initialize","params":{"clientInfo":{"name":"secant","title":"Secant","version":"0.0.0-dev"},"capabilities":{"experimentalApi":false}}} +{"method":"initialized"} +``` + +- `ORCA_*`, `CLAUDE*`, and the inherited `CODEX_HOME` (which pointed at an Orca runtime home) + were removed. The `initialize` response then reported + `"codexHome":"C:\\Users\\rg\\.codex"`, the user's real home. Nothing in it was changed. +- `model/list` offered `gpt-6-luna` ("Fast and affordable model for easier tasks"), which every + `turn/start` used, with `effort: "low"`. Secant sets no effort. +- `thread/start {cwd}` and `turn/start {threadId, input, model}` followed the shape in + `src/harness/codex.ts`. `turn/steer` carried `threadId`, `expectedTurnId`, `input`, and + `clientUserMessageId`, the ADR 0035 shape. The 0.155.0 stable schema + (`codex app-server generate-json-schema`, without `--experimental`) lists + `clientUserMessageId` on both `TurnSteerParams` and `TurnStartParams`. +- No hook ran in these sessions. No server request arrived. + +### Process hygiene + +Every process started for this note exited. A final process-table check against a snapshot taken +before the first run found no `claude`, `codex`, `bash`, or `ping` left from the runs. The +scripts that outlived Claude Code (W1, W5) ended on their own when their `ping` finished. + +## W1: raw `interrupt` without `initialize` + +### W1a: during a Bash-tool script (`w1-tool2`, run 2) + +```text +t=5772 assistant tool_use Bash {"command":"./slowjob.sh"} +t=7355 started.marker written +t=10483 ps: 7880 claude.exe + └ 18100 Git\bin\bash.exe -c "source …snapshot-bash-….sh …" + └ 5132 Git\usr\bin\bash.exe -c "source …" + 8180 Git\usr\bin\bash.exe ./slowjob.sh (parent 17956: no such process) + └ 2920 Git\usr\bin\bash.exe ./slowjob.sh + └ 9792 PING.EXE -n 28 127.0.0.1 +t=10484 in control_request interrupt int1 +t=10502 control_response {"subtype":"success","request_id":"int1","response":{"still_queued":[]}} +t=10506 system/task_notification status=stopped +t=10526 user tool_result "The user doesn't want to proceed with this tool use. The tool use was + rejected …" is_error +t=10569 result subtype=error_during_execution is_error=true terminal_reason=aborted_tools + stop_reason=tool_use num_turns=3 result_index=0 queued_turn_count=0 +t=12232 ps: 7880 claude.exe + └ 11088 taskkill.exe /PID 18100 /T /F + 8180, 2920, 9792 PING.EXE still running +t=12233 in user +t=16819 assistant text "…I immediately attempted to run the Bash tool with the command + `./slowjob.sh`, but that tool use was rejected…, so the command never executed and I + received no output from it." +t=17017 result subtype=success result_index=1 (same session_id) +t=18803 EXIT code=0 (after the driver closed stdin) +t=29681 ps: 8180, 2920, 9792 PING.EXE still running +t=34803 done.marker written: the script ran to completion +``` + +- **The interrupt is honoured, and the process and Session stay alive.** The receipt came + 18 ms after the write, before the `result`. The next frame ran in the same process and Session. +- **Claude Code stops the tool with `taskkill /T /F` on the outer shell.** The `taskkill.exe` + child of `claude.exe` names pid 18100, the Bash tool's `Git\bin\bash.exe`. +- **That kill does not reach the script.** The script's first `bash.exe` had a parent pid that no + process held, even before the interrupt. The same held for every Bash-tool script whose tree was + captured (four scripts in three runs: the parent pids 17956, 2176, 20276, and 15548 were never + live). `taskkill /T` walks + live parent links only, so it could not find the script's subtree. **Inferred**: Git Bash's + Cygwin-style fork and exec leave the child's Windows parent pointing at a process that has + already exited. +- **The script ran to completion in all 6 runs where the raw interrupt stopped it:** W1a runs 1 + and 2, W1c, W2 runs 1 and 2, and the plain interrupt in W2. It also did in W5a, Secant's own + kill. `done.marker` was written about 27.5 s after `started.marker`, the script's full length. + In every raw-interrupt run except W1a run 1, Claude Code had exited before then. (In W1c the two + scripts share one marker file, and both were seen running after the exit.) +- **What the model is told does not match what happened.** The stored `tool_result` is the + rejection text (`toolUseResult: "User rejected tool use"`, `toolDenialKind: "user-rejected"`), + followed by `[Request interrupted by user for tool use]`. The model said the command never ran. + Its side effects all happened. +- The Linux run's "foreground Bash tree was killed" does not carry over to this Windows tool path. + +### W1a': a PowerShell-tool script and a direct native command (one run each) + +```text +PowerShell tool, ./slowjob.ps1: +t=9441 ps: claude.exe └ cmd.exe /d /s /c "chcp 65001 & pwsh.exe …" └ pwsh.exe -Command … └ PING.EXE +t=9442 in control_request interrupt +t=9461 control_response {"still_queued":[]} +t=9526 result error_during_execution terminal_reason=aborted_tools +t=11185 ps: claude.exe └ taskkill.exe /PID 17180 /T /F (cmd, pwsh, and PING gone) + done.marker absent 32 s after exit + +Bash tool, `ping -n 28 127.0.0.1` typed directly: +t=11969 ps: claude.exe └ Git\bin\bash.exe └ Git\usr\bin\bash.exe └ Git\usr\bin\bash.exe └ PING.EXE +t=11991 control_response {"still_queued":[]} +t=13750 ps: claude.exe └ taskkill.exe /PID 17844 /T /F (bash chain and PING gone) +``` + +- **A tool tree with unbroken parent links is killed.** That covers the PowerShell tool's + `cmd.exe`, `pwsh.exe`, and `ping`, and a native command that the Bash tool starts itself. What + escapes is a process that Git Bash starts as a script interpreter. + +### W1b: while text streams (one run) + +````text +t=10817 stream first text_delta +t=13320 in control_request interrupt int1 +t=13329 control_response {"subtype":"success","request_id":"int1","response":{"still_queued":[]}} +t=13369 assistant text "one\ntwo\n…one hundred forty\none hundred forty-" (1896 chars) aborted=true +t=13373 user text "[Request interrupted by user]" +t=13382 result subtype=error_during_execution terminal_reason=aborted_streaming stop_reason=null + total_cost_usd=0 num_turns=2 +t=14883 alive=true; in user +t=19805 assistant text "The last line I wrote was:\n\n```\none hundred forty-\n```\n\nMy response + was cut off mid-word…" +t=20022 result subtype=success result_index=1 +```` + +The transcript stores the partial text with `isAbortedMidStream: true`, then +`[Request interrupted by user]`. This is the same as Linux test 3b, down to `total_cost_usd: 0` +on the interrupted `result`. + +### W1c: a background task and a foreground call in one Turn (one run) + +The model started `./slowjob.sh` with `run_in_background: true`, then again in the foreground. +The interrupt came 4 s into the foreground call. + +```text +t=5837 system/task_started task_id=but651f8j is_backgrounded=true +t=9558 control_response {"still_queued":[]} +t=9618 result error_during_execution terminal_reason=aborted_tools +t=11864 ps: background shell (Git\bin\bash.exe 20356 └ 132) still a child of claude.exe; + both scripts and both PING.EXE running outside the tree +t=20633 system/task_updated task_id=but651f8j status=killed (after the driver closed stdin) +t=22958 EXIT code=1 +t=25187 ps: both scripts and both PING.EXE still running +``` + +- **A background task survives the interrupt**, as on Linux. Claude Code killed its shell only as + the process exited. +- **On Windows, its script outlived the exit too.** So did the killed foreground call's script. + Both ran to completion. + +### Exit code at stdin close + +After the driver closed stdin, Claude Code exited with code 0 when the last `result` was a +success. It exited with code 1 in the two runs where the last `result` was the interrupted one +(W1a', direct `ping`; W1c). This was not compared on Linux. + +## W2: `cancel_queued: true` with a queued frame + +Two runs. Each wrote a user frame 3 s into the Bash call and interrupted 1.5 s later. Run 1: + +```text +t=6726 in user uuid=eeeeeeee-…-01 "Queued message: reply with the single word MANGO." +t=6730 command_lifecycle eeeeeeee-… state=queued +t=8229 in {"type":"control_request","request_id":"int1","request":{"subtype":"interrupt","cancel_queued":true}} +t=8257 command_lifecycle eeeeeeee-… state=cancelled +t=8259 control_response {"subtype":"success","request_id":"int1", + "response":{"still_queued":[],"cancelled":["eeeeeeee-0000-4000-8000-000000000001"]}} +t=8349 result subtype=error_during_execution terminal_reason=aborted_tools result_index=0 +t=8353 command_lifecycle state=cancelled +t=18351 after 10 s: one result, process alive; in user +t=24246 result subtype=success result_index=1 +``` + +- Run 2 gave the same sequence, with the receipt 28 ms after the write. +- The transcript has `queue-operation` `enqueue`, then `remove`, for the queued frame, and no + `queued_command` attachment. The model never saw it. + +**Plain `interrupt` with a queued frame (one run):** + +```text +t=10119 in control_request interrupt int1 +t=10140 control_response {"still_queued":["dddddddd-0000-4000-8000-000000000001"]} +t=10208 result subtype=error_during_execution terminal_reason=aborted_tools result_index=0 +t=10218 command_lifecycle dddddddd-… state=started +t=11968 assistant text "MANGO" +t=12193 result subtype=success result_index=1 user_message_uuids=["dddddddd-…"] +``` + +The queued frame ran by itself as the next Turn, with no new stdin write, as in Linux test 5a. + +## W3: mid-Turn pickup + +### W3a: one frame during a Bash call (one run) + +```text +t=4061 assistant tool_use Bash {"command":"./slowjob.sh"} +t=7063 in user uuid=aaaaaaaa-…-01 "Additional instruction: after the command finishes, also say the word MANGO." +t=7067 command_lifecycle aaaaaaaa-… state=queued +t=33192 user tool_result "started-27\ndone-27" +t=33386 command_lifecycle aaaaaaaa-… state=started +t=35944 assistant text "PINEAPPLE MANGO" +t=36161 command_lifecycle aaaaaaaa-… state=completed +t=36171 result subtype=success num_turns=2 result_index=0 queued_turn_count=0 + user_message_uuids=[, "aaaaaaaa-0000-4000-8000-000000000001"] +``` + +- The transcript has the `queued_command` attachment (`source_uuid: aaaaaaaa-…`, + `commandMode: "prompt"`) after the tool result, then `queue-operation` `remove` with + `reason: "absorbed_mid_turn"`. +- `started` came 194 ms after the `tool_result` frame. The Linux runs measured 15 to 33 ms. This + is one run. + +### W3b: one frame while text streams (one run) + +```text +t=12969 stream first text_delta +t=15471 in user uuid=bbbbbbbb-…-01 "Now reply with the single word MANGO." +t=15478 command_lifecycle bbbbbbbb-… state=queued +t=21606 assistant text "One\nTwo\n…Three Hundred" (5488 chars) +t=21804 result subtype=success num_turns=1 result_index=0 user_message_uuids=[] +t=21809 command_lifecycle bbbbbbbb-… state=started +t=22004 system/init +t=23314 assistant text "MANGO" +t=23524 result subtype=success result_index=1 user_message_uuids=["bbbbbbbb-…"] +``` + +The frame was not taken into the running Turn. It ran as the next native exchange, with no new +stdin write, as in Linux test 4b. `queued_turn_count` was `0` on the first `result` while the +frame waited. + +## W4: raw `interrupt` while Bash waits on the permission bridge + +Run 1. The stand-in held the Bash `approve` for 30 s. The interrupt came 3.7 s into the wait. + +```text +t=8169 assistant tool_use toolu_01Wr… Bash {"command":"./slowjob.sh",…} +t=8411 HTTP POST /mcp tools/call#2 name=approve +t=8417 MCP approve#1 CALLED tool=Bash tool_use_id=toolu_01Wr… reqId=2 -> will allow after 30000ms +t=12130 in control_request interrupt int1 +t=12136 control_response {"subtype":"success","request_id":"int1","response":{"still_queued":[]}} +t=12139 HTTP POST /mcp notifications/cancelled params={"requestId":2,"reason":"AbortError: remote-cancel"} +t=12149 MCP approve#1 extra.signal ABORTED +t=12152 user tool_result "The user doesn't want to proceed with this tool use. …" is_error +t=12154 user text "[Request interrupted by user for tool use]" +t=12196 result subtype=error_during_execution terminal_reason=aborted_tools result_index=0 + permission_denials=[{"tool_name":"Bash","tool_use_id":"toolu_01Wr…",…}] +t=20092 result subtype=success result_index=1 (the question, same Session) +t=38420 MCP approve#1 RETURNING {"behavior":"allow",…} (signal.aborted=true) +t=53799 started.marker absent, done.marker absent; process alive +t=54877 EXIT code=0 (after the driver closed stdin) +``` + +Run 2 wrote a user frame 1.5 s into the wait and interrupted with `cancel_queued: true` 2.3 s +later. The queued frame got `cancelled`, the receipt listed it in `cancelled`, and +`notifications/cancelled` for the `approve` call followed 3 ms after the receipt. The same +`result` and `permission_denials` followed. No Turn ran until the next stdin frame, and the +script never started. + +- **Same as Linux R2-A1.** The pending approval is cancelled over MCP, and the late `allow` does + nothing. The tool never ran, so the Git Bash survival in W1 cannot arise here. + +## W5: Secant's Windows stop (`taskkill /T /F`), then `--resume` + +### W5a: during a Bash-tool script (one run) + +```text +t=3735 assistant tool_use Bash {"command":"./slowjob.sh"} +t=5279 started.marker written +t=8472 ps: 13320 claude.exe └ 13336 Git\bin\bash.exe └ 7064 Git\usr\bin\bash.exe + 7880 bash.exe ./slowjob.sh (parent 2176: no such process) └ 18776 └ PING.EXE 3352 +t=8473 SIGNAL taskkill /pid 13320 /T /F +t=8690 taskkill exit=0 "SUCCESS: … PID 6052 (child process of PID 7064) … PID 7064 (child + process of PID 13336) … PID 18888 (child process of PID 13320) … PID 13336 (child + process of PID 13320) … PID 13320 (child process of PID 15680) has been terminated." +t=8731 EXIT code=1 signal=null (no result frame, no frame after the kill) +t=10947 ps: 7880, 18776, PING.EXE still running +t=32704 done.marker written: the script ran to completion +``` + +The transcript ended at the `tool_use`. The `--resume` process appended, before its first Turn: + +```text +user tool_result "[Tool call interrupted: the session ended before this call's result + was recorded, so its outcome is unknown. Check whether it took effect before relying + on it or running it again.]" is_error toolDenialKind="interrupted" +assistant "No response requested." model= stop_reason=stop_sequence +user +``` + +The resumed model answered: "The last line I wrote was: "No response requested." The command I +ran (`./slowjob.sh`) did not finish. It was interrupted…". The script had in fact finished. + +- **A forced kill writes nothing.** On Linux, a SIGTERM during a running Bash call let Claude Code + kill the tree and write a real `Exit code 137` result (Linux test 1a). `taskkill /F` gives it no + chance, so the resume sees the same "outcome is unknown" result as after a Linux SIGTERM during + a bridge wait (R2-A2). +- **Secant's `taskkill /T` misses the script for the same reason Claude Code's own kill does.** + The script was not among the processes `taskkill` listed. + +### W5b: while text streams (one run) + +```text +t=4222 stream first text_delta +t=7373 SIGNAL taskkill /pid 2768 /T /F +t=7575 EXIT code=1 signal=null (no result frame) +``` + +The transcript held only the prompt. The `--resume` process added the synthetic +`No response requested.`, and the resumed model answered: "The last line I wrote was "No response +requested." No commands were run — you explicitly asked me to do that task "without using any +tools," so I declined rather than execute it." This is the Linux SIGTERM result (test 1b): the +partial text and its thinking are lost. + +## W6: Codex app-server + +Two runs of the same three-part session, each on fresh threads. + +### W6a: `turn/interrupt` mid-Turn + +```text +[ 7.672] item/agentMessage/delta "one" (86 deltas before the interrupt) +[ 9.174] out turn/interrupt {"threadId":"…0ff8","turnId":"…3532"} +[ 9.266] in response {} +[ 9.270] in turn/completed {"turn":{"id":"…3532","status":"interrupted","error":null,"items":[]}} +[ 9.271] out turn/start {"threadId":"…0ff8","input":[{"type":"text","text":"Without using tools: quote the last line you wrote…"}],…} +[12.593] in turn/completed status=completed + agentMessage "I didn't write a count before your message, so I can't quote a last line. + I didn't finish the count." +``` + +- **Confirmed interrupt, and the thread is reused.** The response came in 92 ms in both runs, + and `turn/completed` `status: "interrupted"` 3 to 4 ms later. The same thread took the next + `turn/start`, which completed. +- **The partial answer was not in context.** No `item/completed` `agentMessage` was emitted for + the interrupted Turn, and in both runs the next Turn's model said it had written no count. + Whether Codex keeps any of an interrupted message in history was not checked with `thread/read`. + +### W6b: `turn/steer` with `clientUserMessageId` while text streams + +```text +[17.068] item/started agentMessage +[18.072] out turn/steer {"threadId":"…40c9","expectedTurnId":"…94be","input":[{"type":"text", + "text":"New instruction: stop counting now and reply with just the word MANGO."}], + "clientUserMessageId":"secant-steer-r1"} +[18.081] in response {"turnId":"…94be"} +[40.667] item/completed agentMessage (5488 chars, "one … three hundred") +[40.754] item/started userMessage clientId="secant-steer-r1" "New instruction: stop counting now…" +[40.779] item/completed userMessage clientId="secant-steer-r1" +[42.204] item/completed agentMessage "MANGO" +[42.299] turn/completed {"turn":{"id":"…94be","status":"completed",…}} +``` + +`thread/read` (`includeTurns: true`) shows one Turn holding the prompt, the full count, the +steered `userMessage` with `clientId: "secant-steer-r1"`, and `MANGO`. Run 2 matched. + +- **The steer is accepted at once but taken only at the next model-request boundary.** Here that + was after the whole streamed answer. It is answered in the same Turn, and the `clientId` + echoes the Secant-minted id. +- 0.155.0 also printed a `deprecationNotice` for `thread/read` with `includeTurns: true` on a + paginated thread, pointing at `thread/turns/list` and `thread/items/list`. + +### W6c: empty-input `turn/start` on an idle thread + +```text +T1 turn/start "Remember the code word ZEBRA-42. Reply with just: OK" -> agentMessage "OK", completed +T2 turn/start {"threadId":"…","input":[],"model":"gpt-6-luna","effort":"low"} + -> response {"turn":{"id":"…","status":"inProgress",…}} + -> turn/started, agentMessage "OK", turn/completed status=completed + (no userMessage item in T2) +T3 turn/start "What code word did I ask you to remember?" -> agentMessage "ZEBRA-42" +``` + +The same held in both runs. As on Linux, with no new input the model answered the last user +message in history again. The leftover race (a steer landing after Codex's last pending-input +check) was not attempted. + +## Windows Compared With Linux + +| Case | Linux (Claude Code 2.1.284, codex-cli 0.157.1) | Windows (Claude Code 2.1.283, codex-cli 0.155.0) | +| --------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| Raw `interrupt` without `initialize`: receipt | `control_response` success in 2 to 4 ms, before the `result` | Same, in 6 to 30 ms (12 of 12) | +| Raw `interrupt`: Turn end | `error_during_execution`, `aborted_streaming` / `aborted_tools` | Same | +| Raw `interrupt`: process and Session | Alive; next frame runs in the same Session | Same | +| Raw `interrupt` mid-text: partial text | Kept (`aborted: true`); model quotes the last partial line | Same | +| Raw `interrupt` mid-tool: what the model sees | "rejected" `tool_result`; partial stdout dropped | Same | +| Raw `interrupt` mid-tool: foreground tool processes | Bash tree killed | Claude Code runs `taskkill /PID /T /F`. PowerShell tool and a direct native command: killed. A script run by the Bash tool (Git Bash): **not killed, ran to completion** (6 of 6) | +| Background Bash task | Survives the interrupt; killed at process exit | Survives the interrupt; its shell is killed at exit, but **its script outlives the exit** | +| Plain `interrupt` with a queued frame | Listed in `still_queued`; runs next by itself | Same | +| `interrupt` with `cancel_queued: true` | Listed in `cancelled`; lifecycle `cancelled`; no next Turn; alive | Same (2 of 2) | +| Frame during a tool call | Taken after the round's last result, same Turn, in `user_message_uuids` | Same (1 run) | +| Frame while text streams | Runs as the next native exchange, own `result` | Same (1 run) | +| `command_lifecycle` / `msg_lifecycle_v1` | Present by default on the raw stream | Same | +| `queued_turn_count` | Always `0` | Always `0` | +| `interrupt` during a bridge wait | `notifications/cancelled` within 4 ms; `permission_denials` names the call; late `allow` ignored; tool never runs | Same (2 of 2) | +| Today's stop, mid-text | SIGTERM: exit 143, no `result`, partial text lost; resume adds `No response requested.`; model says it "declined" | `taskkill /T /F`: exit 1, no `result`, partial text lost; same resume and same "declined" answer | +| Today's stop, mid-tool | SIGTERM: Claude Code kills the tree and records `Exit code 137` with partial stdout; resume shows it | `taskkill /T /F`: nothing recorded after the `tool_use`; resume adds "Tool call interrupted… outcome is unknown"; the Git Bash script survives and completes | +| Codex `turn/interrupt` | Not re-run in the leftover note; ADR 0035 relies on `turn/completed` `status: "interrupted"` | Response `{}` in 92 ms, then `turn/completed` `status: "interrupted"`; thread reused (2 of 2) | +| Codex `turn/steer` with `clientUserMessageId` | Accepted; `userMessage` item carries the `clientId` | Same; taken after the streamed answer, answered in the same Turn (2 of 2) | +| Codex empty-input `turn/start` on an idle thread | Accepted; no `userMessage` item; model answers from history | Same (2 of 2) | +| Codex leftover steer race | 6 of 22 at 0 ms after the Stop hook's `hook/completed` | Not attempted | + +## Against ADR 0035 + +These Observed results differ from what ADR 0035 states. The note makes no recommendation. + +- **"kills the foreground tool tree."** On Windows, the raw `interrupt` did not kill a script that + Claude Code's Bash tool ran through Git Bash. The script ran to completion in 6 of 6 runs, while + the model was told the call was rejected and said it never ran. It did kill the PowerShell + tool's tree and a native command that the Bash tool started directly. +- **"a shell the model moved to the background may outlive it until the Session closes."** On + Windows, a background Bash task's script outlived the Session's process as well. +- **The SIGTERM fallback.** On Windows, today's fallback is Secant's `taskkill /T /F`, not SIGTERM. + It also misses the Git Bash script. After it, the resumed model sees "outcome is unknown" rather + than a killed tool's `Exit code 137`. The ADR's statement that the fallback loses partial + streamed text holds on Windows. + +Everything else the ADR relies on held on Windows. The raw `interrupt` answered without +`initialize`, ended the Turn with a `result`, kept the process, Session, and partial text, and +cancelled a pending bridge approval with `notifications/cancelled`. `cancel_queued` handed the +queued frame back. Pickup at the tool round's end and next-exchange delivery during text both +held, as did `command_lifecycle` frames. For Codex, `turn/interrupt`, `turn/steer` with +`clientUserMessageId`, and an empty `turn/start` all held. + +## Still Unknown + +- Whether Claude Code 2.1.284, the Linux version, behaves differently on Windows. Only 2.1.283 was + installed here. +- Why the Git Bash script's Windows parent pid is dead. The Cygwin fork and exec explanation is + **Inferred**. It was not traced, and neither was whether a `CLAUDE_CODE_GIT_BASH_PATH` or + other shell setting changes it. +- Which other Bash-tool commands escape the tree: a pipeline, `bash -c`, `npm`, `bun`, or a + command that runs a `.sh` through `sh`. Only a script run as `./slowjob.sh` escaped, and only + one direct native command was tried. +- Whether a Windows Job Object or a later Claude Code release would make `taskkill /T` reach + these processes. +- SIGINT, or `GenerateConsoleCtrlEvent`, to a hidden Windows Claude Code child. Neither was + tried, since ADR 0035 uses neither. +- Whether a user frame is taken between calls in a multi-call round on Windows, and whether the + denial-boundary timing of Linux round 3 holds. Neither was repeated. +- Why W3a's `started` came 194 ms after the `tool_result`, against 15 to 33 ms on Linux, and + whether that widens the window in which a frame written near a boundary misses the Turn. +- The Codex leftover race on Windows, and whether its hit rate differs from Linux. +- Whether Codex keeps any of an interrupted Turn's partial agent message in history. The next + Turn's model said it had written nothing, in 2 of 2 runs. +- Whether Codex's `turn/interrupt` stops a running shell command's process tree on Windows. W6 + interrupted only streamed text. From 09b02bc61ffe4efad0cae6081014b15b401299e7 Mon Sep 17 00:00:00 2001 From: rohan Date: Tue, 29 Sep 2026 13:32:46 +0530 Subject: [PATCH 11/11] docs(adr): record ADR 0035's Windows tool-process limit and process-stop fallback (#255) --- ...nd-a-mid-turn-message-is-a-native-steer.md | 20 +++++++++++-------- 1 file changed, 12 insertions(+), 8 deletions(-) diff --git a/docs/adr/0035-interrupt-ends-only-the-turn-and-a-mid-turn-message-is-a-native-steer.md b/docs/adr/0035-interrupt-ends-only-the-turn-and-a-mid-turn-message-is-a-native-steer.md index c1da853..870df23 100644 --- a/docs/adr/0035-interrupt-ends-only-the-turn-and-a-mid-turn-message-is-a-native-steer.md +++ b/docs/adr/0035-interrupt-ends-only-the-turn-and-a-mid-turn-message-is-a-native-steer.md @@ -48,14 +48,18 @@ human's compose, so nothing runs after the human said stop. Codex discards pendi **How Claude Code stops.** Claude Code is interrupted with a raw `control_request` `interrupt` written to the same stdin, which, recorded on Claude Code 2.1.284, answers without the SDK's `initialize` in milliseconds, ends the Turn with a `result`, keeps the process and Session live for the next -frame, keeps the partial text in context, and kills the foreground tool tree. It holds while a tool waits on Secant's permission bridge too: Claude -Code cancels the pending approval call with an MCP `notifications/cancelled`, which expires the Harness Request, and the tool never runs. That makes -Claude Code's interruption a confirmed active-Turn interruption rather than a process-only stop. Qualification bounds the dependency: a Claude Code -that does not answer the request falls back to today's SIGTERM of the process tree and `--resume`, which loses partial streamed text. This is the same -wire and the same degrade rule [ADR 0034](./0034-choose-and-change-model-and-effort-as-one-run-wide-model-choice.md) adopted for `set_model`; adopting -the Agent SDK itself stays rejected for the reasons in [Establish Claude Code's viable structured -transports](https://github.com/secantdev/secant/issues/3). An Interrupt stops the Turn's foreground tool work; a shell the model moved to the -background may outlive it until the Session closes. +frame, keeps the partial text in context, and stops the foreground tool. It holds while a tool waits on Secant's permission bridge too: Claude Code +cancels the pending approval call with an MCP `notifications/cancelled`, which expires the Harness Request, and the tool never runs. That makes Claude +Code's interruption a confirmed active-Turn interruption rather than a process-only stop. Qualification bounds the dependency: a Claude Code that does +not answer the request falls back to today's process stop, SIGTERM of the process group on POSIX and `taskkill /T /F` on Windows, and `--resume`, +which loses partial streamed text. This is the same wire and the same degrade rule [ADR +0034](./0034-choose-and-change-model-and-effort-as-one-run-wide-model-choice.md) adopted for `set_model`; adopting the Agent SDK itself stays rejected +for the reasons in [Establish Claude Code's viable structured transports](https://github.com/secantdev/secant/issues/3). An Interrupt stops the Turn's +foreground tool work; a shell the model moved to the background may outlive it until the Session closes. The same held on Windows +([recorded](../research/windows-live-interrupt-and-steer.md) on Claude Code 2.1.283 and codex-cli 0.155.0), with one limit: a script the Bash tool +runs through Git Bash leaves the process tree, so neither Claude Code's stop nor Secant's `taskkill /T` reaches it, and it can outlive the Interrupt +and the Session. The Turn stop itself is still confirmed, Claude Code's Windows interruption evidence states the limit, and containment is decided +separately on [Decide how Secant contains a Harness's descendant processes on Windows](https://github.com/secantdev/secant/issues/259). **Recording.** Each Steer is durable against its Turn: the human's text, when it was sent, and how it settled, as delivered within the Turn, delivered after a native boundary, re-delivered, or dropped by an Interrupt or a lost Turn. An interrupted Turn records the stop used, the control request or