diff --git a/docs/adr/0022-own-a-truthful-deep-harness-seam.md b/docs/adr/0022-own-a-truthful-deep-harness-seam.md index b4c5e84..0ed3931 100644 --- a/docs/adr/0022-own-a-truthful-deep-harness-seam.md +++ b/docs/adr/0022-own-a-truthful-deep-harness-seam.md @@ -25,10 +25,13 @@ may have started but no terminal truth survived qualified, safe, read-only recov interruption is unknown and the last authoritative observation. The Step kind, above this Seam, decides the Attempt outcome and retry policy. The event stream preserves every user-meaningful assistant, tool, command, file-edit, subagent, request, retry, failure, model, session, and recovery -fact; unknown but displayable work becomes generic activity. Context-window pressure is prominent when observed or honestly calculable, while usage, +fact; unknown but displayable work becomes generic activity. Context-window pressure is prominent when reported by the Harness, while usage, cost, and rate facts are optional and estimates stay labelled. Raw protocol frames, private reasoning, telemetry, and ordinary stderr remain private. (Edited 2026-09-29: [ADR 0036](./0036-the-run-workbench-mirrors-the-agent.md) settles that a provider-written reasoning summary is not private reasoning and may cross the Seam; the raw chain of thought stays private.) +(Edited 2026-09-29: [ADR 0038](./0038-carry-observed-tool-and-summary-facts-through-the-harness-seam.md) defines identified typed tool observations, +native command previews, patches, and summary rows, inherits Harness summary settings, and limits context to reported facts rather than calculations. +Unknown protocol methods and accounting notices do not become generic activity; meaningful tool work remains displayable.) The Adapter drains the native transport independently of a slow TUI, coalesces only replaceable previews, closes the producer after all final facts are queued, and only then settles the result; no event can follow it. The result alone carries terminal status, authoritative final assistant content when available, the effective-model observation, post-Turn Session availability, and structured failure. diff --git a/docs/adr/0036-the-run-workbench-mirrors-the-agent.md b/docs/adr/0036-the-run-workbench-mirrors-the-agent.md index 78c405c..661149e 100644 --- a/docs/adr/0036-the-run-workbench-mirrors-the-agent.md +++ b/docs/adr/0036-the-run-workbench-mirrors-the-agent.md @@ -24,7 +24,9 @@ appears once the Run leaves an active state. agent, or not delivered. - An **Entry Turn**'s Bundle prompt is one muted line saying Secant started the Step with it, since the human did not write it. - Assistant text streams as markdown and is never truncated, so an agent's question is always read in full. -- A reasoning summary is one collapsed `Thought: · <duration>` row, with a spinner and `Thinking` while it streams. +- A supplied reasoning summary is one collapsed Thought row, with its first nonempty line as a shortened label, a spinner and `Thinking` while it + streams, and a duration only when the Harness reports a trustworthy reasoning duration. Existing Harness summary settings and unset defaults are + inherited (ADR 0038). - Each tool call is one row that changes in place: muted once settled, a spinner while running, the error colour when it fails. Shell output and file diffs are panels. - An **Agent call** shows its call, the agent's reason, and what Secant did with it. An answered Human Gate shows the answer. @@ -33,6 +35,8 @@ appears once the Run leaves an active state. **Truncation.** Only shell output collapses, at 10 lines, ending with how many lines are hidden. Reasoning bodies and the Entry Turn prompt collapse to their line. `ctrl+o` or a click expands everything collapsed. Assistant text, questions, the human's messages, and diffs are never cut; they wrap. +The shell bound applies to live output and completed output alike, fits the available width, and expands only on human action. Native output updates +replace the same call's preview; streaming never automatically opens a panel (ADR 0038). **Colour.** Colour comes only from the vendored theme roles, and everforest is the default theme. The agent colour marks the human's messages, the prompt bar, the Turn line, and the working indicator. Muted text marks settled work, `warning` marks reasoning rows and Harness Requests, `error` @@ -76,7 +80,8 @@ raw reasoning text, stays private under ADR 0022. - typed tool rows (a kind, the main input, and a result count), command output and exit code as fields, and file diffs as data; - a reasoning-summary event; - Turn history that grows during a Turn, with message identity so streaming text settles in place and durable rows arrive mid-Turn; -- noise kept out of `activity`, and `context` filled by both Adapters; +- noise kept out of `activity`, and context/usage facts supplied by both Adapters only where reported; Secant calculates no context occupancy, + token counts, or percentages (ADR 0038); - Steer state, Turn duration, Entry Turn authorship, and Agent-call rows in the Projection; - which of these facts headless `--json` gains. diff --git a/docs/adr/0038-carry-observed-tool-and-summary-facts-through-the-harness-seam.md b/docs/adr/0038-carry-observed-tool-and-summary-facts-through-the-harness-seam.md new file mode 100644 index 0000000..d7d7ac0 --- /dev/null +++ b/docs/adr/0038-carry-observed-tool-and-summary-facts-through-the-harness-seam.md @@ -0,0 +1,77 @@ +# Carry observed tool and summary facts through the Harness Seam + +The Run Workbench needs one tool row that changes in place, command output, file diffs, and collapsed reasoning summaries. The existing Harness +Interface discards native call identity and flattens useful fields into summary strings. [Decide the Harness Seam's tool, command, diff, and +reasoning-summary events](https://github.com/secantdev/secant/issues/262) resolves the semantic contract: the Harness Adapter carries observed, +normalized facts, and the Projection Module owns their reconciliation for clients. Secant shows what the Harness supplies; it does not manufacture +missing evidence or change the user's Harness settings to make an optional display feature available. + +**Tool identity and lifecycle.** Every observed tool call has one stable opaque identity within its Turn. Both Adapters privately translate native +identities; repeated or concurrent calls never pair by name, summary, or adjacency. An observed parent-call relationship uses the same identity +vocabulary. Start, updates, and result refer to the same call. The Interface carries running, completed, failed, and declined outcomes with useful +error or refusal text where reported. Failure belongs to the tool and does not itself decide the Turn or Run outcome. A Harness Request remains a +separate interaction, and an Agent call retains ADR 0033's separate declaration and disposition. + +An authoritative Turn settlement stops live indicators. A tool with no terminal result remains unconfirmed: Turn success cannot prove tool success, +and Turn interruption or loss cannot prove the tool failed or stopped. Clients can show the missing result in the context of the settled Turn +without a synthetic tool completion or a second cancellation controller. This matters because Secant observes an external Harness; OpenCode owns +its tools and can settle its own aborted executions, while T3 Code's Claude Adapter synthesizes outcomes for unmatched calls. + +**Typed facts.** The normalized kinds are read, search, command, file change, web, MCP tool, subagent, and other. They carry their meaningful main +input, with an optional reported result count and its unit, such as lines, matches, or files. Unknown counts remain absent, not zero. A command +remains a command unless native evidence supports a more specific classification. Other preserves meaningful unfamiliar work without exposing raw +protocol frames. Native names, arguments, results, and errors are translated behind the Harness Interface; clients do not parse tool-summary prose +or infer a kind from arbitrary command text. Exact type and event names remain implementation choices. + +**Command output.** Command text, working directory when observed, output, and an observed exit code are separate facts. When the native transport +supplies output during execution, it updates a replaceable preview for the same call. If only completed output is supplied, it appears on completion. +Final native output reconciles with the preview without duplication. A Harness-supplied omission or truncation is preserved as evidence; it is not +confused with the Workbench's collapsed display. Missing output and exit codes remain unknown. No polling, process inspection, or Harness-setting +override is introduced to obtain live output. + +Live and completed shell panels remain bounded by default: ADR 0036's ten-line collapse applies while running as well as after completion, text fits +the available width, and expansion is an explicit human action. Streaming never automatically expands a panel or consumes the screen. Preserve the +available output for expansion or inspection; a short summary alone does not replace the supplied output. Preview updates are coalescible, while +terminal facts are not. The Projection decision fixes retained preview budgets, interrupted partial-output retention, final reconciliation, durable +publication, and headless exposure before implementation; routing each output chunk through today's durable tool-event append is excluded. + +**File changes.** Preserve observed paths, change kinds, and supplied patches as data. A per-call patch belongs to that call. A cumulative Turn diff +remains Turn-scoped when the Harness supplies no call association. If only changed paths are reported, those are displayed without a manufactured +patch. Requested edit input is not evidence that the edit happened. Secant does not run Git, scan the Workspace, or reconstruct changes to fill a +missing Harness diff. ADR 0036's rule that supplied diffs are not cut still applies. + +**Reasoning summaries.** Inherit the Harness's existing summary behavior, including its defaults when no preference is configured. Secant does not +enable summaries, change effort or thinking mode, or rewrite global settings for Thought rows. A supplied provider-written Reasoning summary gains +opaque identity and streaming/completion facts so its text settles in one row without appending the final snapshot to its own deltas. Its first +nonempty line supplies a shortened collapsed label; expansion reveals the complete body. While a summary streams the row shows Thinking and a +spinner. Duration appears only when the Harness reports a trustworthy reasoning duration; summary-delivery timing is not substituted. Missing, +empty, or unsupported summaries produce no Thought body. Raw reasoning, signatures, and encrypted or redacted payloads stay private under ADR 0022. +Recognizing a summary must be grounded in the qualified transport/model/provider semantics, not an arbitrary historical thinking field. + +**Context and noise.** Carry only reported context and usage facts, keeping their meanings distinct. Input/output usage, cached-token counters, +cumulative usage, a model's reported window capacity, and an explicitly reported context measurement are not interchangeable. Secant never tokenizes, +adds counters into context occupancy, calculates a percentage, guesses a model limit, or repairs inconsistent Harness figures. A percentage is shown +only if reported by the Harness. Unavailable fields stay absent. Both Adapters consume the relevant native observations without implying that both +can populate every context field. Usage and account/rate observations remain separate from conversation activity; a genuine failure or pause uses +its appropriate semantic failure or request rather than a protocol-name row. Unknown methods, telemetry, and raw reasoning do not become generic +activity; meaningful known work, including an unfamiliar tool represented as other, is preserved. + +**Ownership and qualification.** Native schemas, classification, call correlation, and source-specific output assembly stay behind the Harness +Interface. The application Projection Module owns bounded per-call reconciliation, preview replacement, and Turn-derived liveness through the +existing Projection Port; presentation owns wrapping, collapse, expansion, and colour. No new public Module, raw-provider escape hatch, or execution +control is introduced. The Interface is the test surface: both Adapters and the fake exercise interleaved identities, reported failure/refusal, +preview/final reconciliation, unknown tool outcomes, and terminal ordering; native fixtures qualify the fields actually consumed. Recorded Codex +evidence establishes command output and exit codes. Current recorded Claude shell cases do not establish its structured exit-code shape; that field +must be qualified before extraction, with absence represented honestly. + +This narrows ADR 0022's permission to calculate context into a reported-facts-only policy, and clarifies ADR 0036's Thought duration, bounded shell +panels, and context requirement. It leaves [Decide how a Turn's history grows mid-Turn in the Run Projection and headless +--json](https://github.com/secantdev/secant/issues/263) to settle persistence, live history, resource retention, and the frozen client contracts. +Implementation still follows the milestone loop. + +The [OpenCode investigation](../research/opencode-tool-output-lifecycle.md) establishes live bounded shell previews and owned tool cleanup in its +existing TUI; its newer Core Bash progress remains unfinished. The [T3 Code investigation](../research/t3code-tool-output-lifecycle.md) establishes +completed output summaries, stable call collapse, and the cost of fabricated unmatched outcomes. These are source and existing-test inspections, +not live qualification runs. Completed-only output was considered; optional native live output earns its incremental complexity by sharing the +already-required identified row and replaceable-preview mechanism. Automatic summary opt-in, calculated context, provider-specific client reducers, +and synthetic tool completion were rejected because they change configuration or redistribute interpretation and invented truth into callers. diff --git a/docs/research/opencode-tool-output-lifecycle.md b/docs/research/opencode-tool-output-lifecycle.md new file mode 100644 index 0000000..1ed7d76 --- /dev/null +++ b/docs/research/opencode-tool-output-lifecycle.md @@ -0,0 +1,143 @@ +# OpenCode Tool Output and Lifecycle + +Research date: 2026-09-29. Decision context: Secant wayfinder #262, Q3 (live shell output) and Q4 (tool identity and lifecycle). + +Source: the user-supplied sibling checkout `../opencode`, HEAD +[`b3f1a96c6dd7adeb28b36dd11add1998fc84d67b`](https://github.com/anomalyco/opencode/commit/b3f1a96c6dd7adeb28b36dd11add1998fc84d67b). +`git status --short` was empty before and after inspection. No sibling files were changed. Root and relevant package/tool/test `AGENTS.md` files +were read. This is source inspection, not a running-terminal observation; referenced tests were read, not executed. + +## Answer + +**Source-observed:** OpenCode's existing TUI shows shell output during execution through replaceable, bounded **metadata snapshots** on an +identified tool part. The panel remains the same tool after completion; its spinner stops. Completion does not switch that terminal panel to +the full model-facing output. Failure retains the latest progress metadata. Tool identities and state changes are owned by its execution +runtime, then reconciled into the terminal store.[^shell-producer][^shell-view][^lifecycle][^sync] + +**Inferred for Secant:** adopt the presentation pattern and the distinction between replaceable previews and final facts. Keep lifecycle +translation inside the Harness Adapter and the per-call view behind the Projection Port. OpenCode owns execution, permissions, process +termination, and settlement; Secant observes an external Harness. OpenCode's synthetic aborted-tool cleanup therefore does not establish that +Secant can declare an external tool interrupted when its transport disappears. Secant's ADR 0022 already requires independent transport drain, +coalescing only replaceable previews, and authoritative terminal results.[^secant] + +## Q3: running and completed shell output + +| Aspect | Source-observed behavior | +| -------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| Before output | The producer sets `metadata.output = ""` before spawning. The TUI tests whether that field exists, so an empty-output panel can appear before the first chunk. Before the field exists, it uses the inline `$ command` / `~ Writing command…` row.[^shell-producer][^shell-view] | +| While running | Every decoded chunk updates a cumulative preview snapshot. `preview()` retains the last 30,000 JavaScript string units, prefixed by an omission marker once trimmed. The TUI strips ANSI, trims whitespace, shows a command spinner only while status is `running`, and renders that preview.[^preview][^shell-producer][^shell-view] | +| Collapsed | The panel shows the first 10 lines of the current preview with a width-dependent character budget `10 * max(20, width - 6)`. Clicking expands/collapses. This is a plain text panel, not a terminal emulator.[^shell-view][^collapse] | +| After completion | The same metadata-backed panel remains, with `$ command` replacing the spinner. The shell result contains separate model-facing `output`, bounded by line/byte limits with a saved full-output path when truncated; metadata keeps its bounded preview, exit code, and truncation facts. The TUI shell renderer does not show exit code or read `props.output`.[^shell-producer][^shell-view] | +| Failed / interrupted | Generic tool failure preserves streamed metadata. A panel that already exists shows that preview and the error at its bottom. An inline failed row turns red and allows click-to-reveal error text.[^lifecycle][^tool-errors] | + +The exposed legacy shell tool id remains `bash`, even for other shell implementations, explicitly for compatibility.[^shell-id] +The web shell renderer differs: it allows opening while pending and reads `props.output || props.metadata.output`, so completion can reveal +the model-facing result instead of only the preview. That web rule is not the terminal rule.[^web] + +### Capture, batching, and backpressure + +**Source-observed:** the shell stream consumer awaits `ctx.metadata(...)` per decoded chunk. That context calls the processor's +`updateToolCall`, which reads and writes the identified part through the Session implementation. The producer retains a bounded chunk tail; +after the full-output byte threshold it starts an append file sink. Its `sink.write(chunk)` return value is ignored: the inspected loop has no +explicit wait for Node writable `drain`.[^shell-producer][^metadata][^lifecycle] + +The TUI SDK queues events arriving within a 16 ms window and emits **all** queued events inside one Solid `batch`. This reduces rendering +work but does not deduplicate snapshots or impose a queue capacity. The terminal sync store replaces/reconciles a part by message id + part id; +text deltas use a separate append path.[^batch][^sync] + +**Inferred:** the awaited metadata callback offers local producer pacing, but neither that nor rendering batches demonstrates end-to-end +backpressure from a slow terminal to the native subprocess. Snapshot length bounds retained display content, not total event volume or sink +buffering. For Secant, a cadence/latest-per-call preview path should be decided before routing stdout chunks into its durable event append path. +Terminal status and final output must survive preview replacement. + +## Q4: identity, updates, and terminal meaning + +**Source-observed:** a legacy tool part has a generated part id, message id, session id, native `callID`, tool name, and a four-variant state: +`pending`, `running`, `completed`, `error`. The processor maintains a map keyed by call id; input start creates one pending part, a call makes +it running, and success/failure updates that part. Permission presentation correlates pending requests using `callID`.[^identities][^lifecycle][^tool-errors] + +| Observation | Runtime / display behavior | +| --------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| Success | Only a currently running tool can settle completed; final output, metadata, and end time replace its running state.[^lifecycle] | +| Failure | Error state stores the input, error string, end time, and latest metadata. Provider error results and `tool-error` use this same failure path.[^lifecycle] | +| Declined permission or question | These are error settlements at the processor. The inline TUI detects denial from error-string substrings (`QuestionRejectedError`, `rejected permission`, `specified a rule`, `user dismissed`) and strikes the row through. This is presentation classification, not a typed declined lifecycle variant.[^lifecycle][^tool-errors] | +| Interruption / unfinished cleanup | Cleanup waits up to 250 ms for each outstanding tool, then writes `error: "Tool execution aborted"` and `metadata.interrupted = true`, preserving previous metadata. There is no separate interrupted tool state. The shell itself can also return a result containing `User aborted the command`; that lower-level return is not independently proof of a failed tool settlement.[^cleanup][^shell-producer] | +| Lost | Neither inspected tool-state schema has a `lost` variant. Absence of a final result is not modeled as a successful tool. The legacy cleanup rule is owned by execution, not evidence about an externally disconnected Harness.[^identities][^current-schema][^cleanup] | + +**Inferred:** Secant can preserve native call correlation and observed states without granting the TUI a new execution-control Interface. +Scope correlation by Turn/session/message where necessary; a native string call id is not shown here to be globally unique. A lost Turn should +leave unresolved tool completion explicitly unknown unless the Adapter has native terminal evidence. Do not translate Turn loss into tool +success, failure, or confirmed interruption merely to produce a tidy row.[^secant] + +## Current Core migration: evidence limit + +This snapshot contains both the existing compatibility TUI/runtime and newer Core contracts. The newer tool Context contains session, +assistant-message, agent, and tool-call identities but no progress callback. Its Bash implementation uses `AppProcess.run`, a 1 MiB capture +limit, and explicitly keeps a TODO to wire live progress. The Core process collector consumes output while bounding retained bytes; it does +not provide live shell preview events to that Bash caller.[^current-bash][^current-context][^current-process] + +The newer event contract does define durable `Tool.Progress` snapshots, with an explicit comment to checkpoint semantic changes or bounded +cadence **rather than persist every stdout/stderr chunk**. Its projector replaces running structured/content progress, replaces it on success, +and preserves it on failure. Its runner marks pending/running tools failed with `Tool execution interrupted` when interrupted. These are +useful owned-lifecycle examples, not evidence that the new Bash implementation already streams shell output.[^current-progress][^current-update][^current-interrupt] + +## Test evidence and complexity placement + +Read tests cover progressive metadata updates, preserving output on shell abort, truncation with full-output retention, processor cleanup +marking pending tools aborted, and durable Core progress retained on failure.[^tests][^cleanup-test][^current-progress-test] +No inspected test establishes slow-consumer boundedness, writable-sink backpressure, or a rendered live-shell frame across interruption. + +**Inferred Module assessment:** the substantial complexity is lifecycle and evidence ownership, not drawing the shell box. OpenCode localizes +execution/capture in the shell implementation, settlement in the Session processor, and rendering/reconciliation in the TUI. For Secant the +deep Module is the Harness Adapter: native ids, status translation, output replacement, transport draining, and uncertainty belong behind its +small Interface. The application Projection Module should own bounded per-call view reconciliation; the TUI should receive those facts and +own collapse/expand only. Passing provider-specific metadata and denial-string matching into Secant's TUI would make that Interface shallow +and redistribute correctness into callers. This is a design assessment, not an implementation proposal or completed dependency audit. + +[^shell-producer]: OpenCode [`tool/shell.ts` lines 435–594](https://github.com/anomalyco/opencode/blob/b3f1a96c6dd7adeb28b36dd11add1998fc84d67b/packages/opencode/src/tool/shell.ts#L435-L594). + +[^preview]: OpenCode [`tool/shell.ts` lines 220–223](https://github.com/anomalyco/opencode/blob/b3f1a96c6dd7adeb28b36dd11add1998fc84d67b/packages/opencode/src/tool/shell.ts#L220-L223), with limit at line 27. + +[^shell-view]: OpenCode [`routes/session/index.tsx` lines 2046–2103](https://github.com/anomalyco/opencode/blob/b3f1a96c6dd7adeb28b36dd11add1998fc84d67b/packages/tui/src/routes/session/index.tsx#L2046-L2103). + +[^collapse]: OpenCode [`collapse-tool-output.ts` lines 1–19](https://github.com/anomalyco/opencode/blob/b3f1a96c6dd7adeb28b36dd11add1998fc84d67b/packages/tui/src/util/collapse-tool-output.ts#L1-L19). + +[^shell-id]: OpenCode [`shell/id.ts` lines 14–17](https://github.com/anomalyco/opencode/blob/b3f1a96c6dd7adeb28b36dd11add1998fc84d67b/packages/opencode/src/tool/shell/id.ts#L14-L17). + +[^web]: OpenCode [`message-part.tsx` lines 2085–2125](https://github.com/anomalyco/opencode/blob/b3f1a96c6dd7adeb28b36dd11add1998fc84d67b/packages/session-ui/src/components/message-part.tsx#L2085-L2125). + +[^metadata]: OpenCode [`session/tools.ts` lines 59–86](https://github.com/anomalyco/opencode/blob/b3f1a96c6dd7adeb28b36dd11add1998fc84d67b/packages/opencode/src/session/tools.ts#L59-L86). + +[^batch]: OpenCode [`context/sdk.tsx` lines 48–80](https://github.com/anomalyco/opencode/blob/b3f1a96c6dd7adeb28b36dd11add1998fc84d67b/packages/tui/src/context/sdk.tsx#L48-L80). + +[^sync]: OpenCode [`context/sync.tsx` lines 376–414](https://github.com/anomalyco/opencode/blob/b3f1a96c6dd7adeb28b36dd11add1998fc84d67b/packages/tui/src/context/sync.tsx#L376-L414). + +[^identities]: OpenCode [`v1/session.ts` lines 259–324](https://github.com/anomalyco/opencode/blob/b3f1a96c6dd7adeb28b36dd11add1998fc84d67b/packages/schema/src/v1/session.ts#L259-L324). + +[^lifecycle]: OpenCode [`session/processor.ts` lines 123–250](https://github.com/anomalyco/opencode/blob/b3f1a96c6dd7adeb28b36dd11add1998fc84d67b/packages/opencode/src/session/processor.ts#L123-L250), [call/result/error lines 331–418](https://github.com/anomalyco/opencode/blob/b3f1a96c6dd7adeb28b36dd11add1998fc84d67b/packages/opencode/src/session/processor.ts#L331-L418). + +[^tool-errors]: OpenCode [`routes/session/index.tsx` lines 1856–1906](https://github.com/anomalyco/opencode/blob/b3f1a96c6dd7adeb28b36dd11add1998fc84d67b/packages/tui/src/routes/session/index.tsx#L1856-L1906), [block error lines 1994–2041](https://github.com/anomalyco/opencode/blob/b3f1a96c6dd7adeb28b36dd11add1998fc84d67b/packages/tui/src/routes/session/index.tsx#L1994-L2041). + +[^cleanup]: OpenCode [`session/processor.ts` lines 585–608](https://github.com/anomalyco/opencode/blob/b3f1a96c6dd7adeb28b36dd11add1998fc84d67b/packages/opencode/src/session/processor.ts#L585-L608). + +[^current-schema]: OpenCode [`session-message.ts` lines 81–139](https://github.com/anomalyco/opencode/blob/b3f1a96c6dd7adeb28b36dd11add1998fc84d67b/packages/schema/src/session-message.ts#L81-L139). + +[^current-bash]: OpenCode [`core/tool/bash.ts` lines 19–81](https://github.com/anomalyco/opencode/blob/b3f1a96c6dd7adeb28b36dd11add1998fc84d67b/packages/core/src/tool/bash.ts#L19-L81), [execution lines 160–210](https://github.com/anomalyco/opencode/blob/b3f1a96c6dd7adeb28b36dd11add1998fc84d67b/packages/core/src/tool/bash.ts#L160-L210). + +[^current-context]: OpenCode [`core/tool/tool.ts` lines 10–15](https://github.com/anomalyco/opencode/blob/b3f1a96c6dd7adeb28b36dd11add1998fc84d67b/packages/core/src/tool/tool.ts#L10-L15). + +[^current-process]: OpenCode [`core/process.ts` lines 121–160](https://github.com/anomalyco/opencode/blob/b3f1a96c6dd7adeb28b36dd11add1998fc84d67b/packages/core/src/process.ts#L121-L160). + +[^current-progress]: OpenCode [`session-event.ts` lines 273–372](https://github.com/anomalyco/opencode/blob/b3f1a96c6dd7adeb28b36dd11add1998fc84d67b/packages/schema/src/session-event.ts#L273-L372). + +[^current-update]: OpenCode [`message-updater.ts` lines 250–342](https://github.com/anomalyco/opencode/blob/b3f1a96c6dd7adeb28b36dd11add1998fc84d67b/packages/core/src/session/message-updater.ts#L250-L342). + +[^current-interrupt]: OpenCode [`runner/llm.ts` lines 119–150](https://github.com/anomalyco/opencode/blob/b3f1a96c6dd7adeb28b36dd11add1998fc84d67b/packages/core/src/session/runner/llm.ts#L119-L150). + +[^tests]: OpenCode [`shell.test.ts` lines 1009–1041](https://github.com/anomalyco/opencode/blob/b3f1a96c6dd7adeb28b36dd11add1998fc84d67b/packages/opencode/test/tool/shell.test.ts#L1009-L1041), [lines 1108–1162](https://github.com/anomalyco/opencode/blob/b3f1a96c6dd7adeb28b36dd11add1998fc84d67b/packages/opencode/test/tool/shell.test.ts#L1108-L1162). + +[^cleanup-test]: OpenCode [`processor-effect.test.ts` lines 873–938](https://github.com/anomalyco/opencode/blob/b3f1a96c6dd7adeb28b36dd11add1998fc84d67b/packages/opencode/test/session/processor-effect.test.ts#L873-L938). + +[^current-progress-test]: OpenCode [`session-tool-progress.test.ts` lines 29–160](https://github.com/anomalyco/opencode/blob/b3f1a96c6dd7adeb28b36dd11add1998fc84d67b/packages/core/test/session-tool-progress.test.ts#L29-L160). + +[^secant]: Secant [ADR 0022](https://github.com/secantdev/secant/blob/199212652826a33b9f3b98969eaed6423647f0da/docs/adr/0022-own-a-truthful-deep-harness-seam.md), authoritative Turn results and event-drain paragraphs. Domain distinction also follows [`CONTEXT.md`](https://github.com/secantdev/secant/blob/199212652826a33b9f3b98969eaed6423647f0da/CONTEXT.md). diff --git a/docs/research/t3code-tool-output-lifecycle.md b/docs/research/t3code-tool-output-lifecycle.md new file mode 100644 index 0000000..8430581 --- /dev/null +++ b/docs/research/t3code-tool-output-lifecycle.md @@ -0,0 +1,33 @@ +# T3 Code: shell output and tool-call lifecycle + +Research date: 2026-09-29. Decision context: Secant [wayfinder #262](https://github.com/secantdev/secant/issues/262), Q3 and Q4. Read-only local source: `../t3code`, HEAD [`d2c9281b8112dc3b2991642c4bdb985e4b08b9bb`](https://github.com/pingdotgg/t3code/tree/d2c9281b8112dc3b2991642c4bdb985e4b08b9bb); `git status --porcelain=v1` was empty before and after investigation. Root `AGENTS.md` was the only tracked guidance file found. No provider, server, browser, or tests were run; observations below follow source and existing test assertions, not measured behavior. + +## Q3: live shell output does not reach this web conversation + +**Source-observed.** Codex maps `item/commandExecution/outputDelta` to `content.delta` with `streamKind: command_output`. Ingestion immediately discards every content stream except assistant text and reasoning. Claude's analogous command-output event is emitted from the final `tool_result`, between its update and completion; it is not a stream of shell stdout while execution is underway. Parent `tool_progress` is also discarded; only task-owned heartbeats become activities. [Codex mapping](https://github.com/pingdotgg/t3code/blob/d2c9281b8112dc3b2991642c4bdb985e4b08b9bb/apps/server/src/provider/Layers/CodexAdapter.ts#L1798-L1817), [Claude result mapping](https://github.com/pingdotgg/t3code/blob/d2c9281b8112dc3b2991642c4bdb985e4b08b9bb/apps/server/src/provider/Layers/ClaudeAdapter.ts#L3181-L3278), [ingestion filter](https://github.com/pingdotgg/t3code/blob/d2c9281b8112dc3b2991642c4bdb985e4b08b9bb/apps/server/src/orchestration/Layers/ProviderRuntimeIngestion.ts#L1781-L1790), [progress filter](https://github.com/pingdotgg/t3code/blob/d2c9281b8112dc3b2991642c4bdb985e4b08b9bb/apps/server/src/orchestration/Layers/ProviderRuntimeIngestion.ts#L813-L845). + +Completed command output takes a different route: Codex's `data.item.aggregatedOutput` and Claude's `data.result` content survive in the completed activity, then read projection reduces them to the first meaningful line, capped at 84 characters. Nonterminal updates persist only this projected payload to avoid repeated accumulated-output storage; completion persists the full payload. Projection preserves command and a few other display fields, but does not preserve a structured command exit code. This is a compact summary pipeline, not an output panel pipeline. [Persistence policy](https://github.com/pingdotgg/t3code/blob/d2c9281b8112dc3b2991642c4bdb985e4b08b9bb/apps/server/src/orchestration/Layers/ProviderRuntimeIngestion.ts#L932-L1033), [Codex output projection and summary](https://github.com/pingdotgg/t3code/blob/d2c9281b8112dc3b2991642c4bdb985e4b08b9bb/apps/server/src/orchestration/ActivityPayloadProjection.ts#L86-L187), [Claude and generic projection](https://github.com/pingdotgg/t3code/blob/d2c9281b8112dc3b2991642c4bdb985e4b08b9bb/apps/server/src/orchestration/ActivityPayloadProjection.ts#L425-L501). + +Actual web rendering uses the projected output as work-entry `detail`, strips a trailing textual exit marker, and expands command/detail/files into a scrolling `<pre>`. It has no full-output fetch in this renderer. While active, rows show labels such as `Running <program>`; starts alone are hidden, and tool-input updates can make a row visible before output exists. Thus an active status label or expandable command row is not evidence of live shell output. [Detail extraction](https://github.com/pingdotgg/t3code/blob/d2c9281b8112dc3b2991642c4bdb985e4b08b9bb/apps/web/src/session-logic.ts#L1217-L1302), [expanded-body builder](https://github.com/pingdotgg/t3code/blob/d2c9281b8112dc3b2991642c4bdb985e4b08b9bb/apps/web/src/components/chat/MessagesTimeline.tsx#L4413-L4462), [rendered body](https://github.com/pingdotgg/t3code/blob/d2c9281b8112dc3b2991642c4bdb985e4b08b9bb/apps/web/src/components/chat/MessagesTimeline.tsx#L4727-L4793), [actual `<pre>`](https://github.com/pingdotgg/t3code/blob/d2c9281b8112dc3b2991642c4bdb985e4b08b9bb/apps/web/src/components/chat/MessagesTimeline.tsx#L4930-L4937), [active label](https://github.com/pingdotgg/t3code/blob/d2c9281b8112dc3b2991642c4bdb985e4b08b9bb/apps/web/src/components/chat/MessagesTimeline.logic.ts#L72-L98). + +**Decision inference.** T3 Code supplies evidence for bounded completed summaries and stable activity rows. It does not validate Q3's live shell panels. Passing command deltas through Secant would additionally require bounded per-call accumulation/replacement, final-output reconciliation, and a retention policy; otherwise each chunk becomes another retained event or repeated accumulated payload. + +## Q4: identity is useful; terminal outcomes require care + +**Source-observed.** Claude uses `tool_use.id` as item identity, remembers streamed input, matches results by `tool_use_id`, and reports `failed` when `is_error` is true. Codex carries its native item identity and explicitly preserves `failed` and `declined` on completion; other completed items default to `completed`. Ingestion promotes item identity to `toolCallId` and retains native status. [Claude start](https://github.com/pingdotgg/t3code/blob/d2c9281b8112dc3b2991642c4bdb985e4b08b9bb/apps/server/src/provider/Layers/ClaudeAdapter.ts#L3073-L3145), [Claude results](https://github.com/pingdotgg/t3code/blob/d2c9281b8112dc3b2991642c4bdb985e4b08b9bb/apps/server/src/provider/Layers/ClaudeAdapter.ts#L3181-L3278), [Codex lifecycle mapping](https://github.com/pingdotgg/t3code/blob/d2c9281b8112dc3b2991642c4bdb985e4b08b9bb/apps/server/src/provider/Layers/CodexAdapter.ts#L1000-L1041), [activity payloads](https://github.com/pingdotgg/t3code/blob/d2c9281b8112dc3b2991642c4bdb985e4b08b9bb/apps/server/src/orchestration/Layers/ProviderRuntimeIngestion.ts#L932-L1033). + +The client skips `tool.started`, collapses updates/completion by `(turnId, toolCallId)` even when other calls interleave, and refuses to merge further updates into an already completed row. Legacy anonymous rows fall back to label/type/detail heuristics. Snapshots remove preceding updates superseded by later same-turn completion; live delivery coalesces identified updates within 50 ms and preserves completion ordering. These are related but separately owned policies. [Client filter](https://github.com/pingdotgg/t3code/blob/d2c9281b8112dc3b2991642c4bdb985e4b08b9bb/apps/web/src/session-logic.ts#L451-L514), [collapse and merge](https://github.com/pingdotgg/t3code/blob/d2c9281b8112dc3b2991642c4bdb985e4b08b9bb/apps/web/src/session-logic.ts#L708-L885), [snapshot pruning](https://github.com/pingdotgg/t3code/blob/d2c9281b8112dc3b2991642c4bdb985e4b08b9bb/apps/server/src/orchestration/ActivityPayloadProjection.ts#L549-L663), [live coalescing](https://github.com/pingdotgg/t3code/blob/d2c9281b8112dc3b2991642c4bdb985e4b08b9bb/apps/server/src/orchestration/ThreadLiveEventCoalescer.ts#L18-L90). + +Unmatched terminal cleanup is asymmetric. Claude's `completeTurn` fabricates completion for every remaining in-flight tool: a successful Turn marks it `completed`; interruption/failure marks it `failed`, without a result block. Its stream-ending and stop paths call this routine. Codex maps Turn completion/abort separately and has no corresponding per-tool synthesis in the inspected Adapter. Shared ingestion clears session/assistant/request state, but does not settle all unmatched tool rows. **Inference:** a lost tool result can therefore produce an invented success/failure in Claude, or an unfinished stored row in Codex. Neither should be treated as confirmed tool outcome. [Claude cleanup](https://github.com/pingdotgg/t3code/blob/d2c9281b8112dc3b2991642c4bdb985e4b08b9bb/apps/server/src/provider/Layers/ClaudeAdapter.ts#L2743-L2829), [stream-ending paths](https://github.com/pingdotgg/t3code/blob/d2c9281b8112dc3b2991642c4bdb985e4b08b9bb/apps/server/src/provider/Layers/ClaudeAdapter.ts#L4230-L4252), [Codex terminal mapping](https://github.com/pingdotgg/t3code/blob/d2c9281b8112dc3b2991642c4bdb985e4b08b9bb/apps/server/src/provider/Layers/CodexAdapter.ts#L1616-L1643), [shared cleanup](https://github.com/pingdotgg/t3code/blob/d2c9281b8112dc3b2991642c4bdb985e4b08b9bb/apps/server/src/orchestration/Layers/ProviderRuntimeIngestion.ts#L2349-L2458). + +Stale animation is chiefly suppressed by Turn/session truth: accepted terminal events clear active Turn and set ready/error/interrupted/stopped; visual active tools require `isWorking` and the current unsettled Turn. That stops active indicators without requiring every tool to receive a terminal event. It does not prove the stored tool completed. The client additionally recognizes native interrupted/cancelled status as `stopped`, but Claude's synthesized tool cleanup emits `failed` instead. [Session lifecycle](https://github.com/pingdotgg/t3code/blob/d2c9281b8112dc3b2991642c4bdb985e4b08b9bb/apps/server/src/orchestration/Layers/ProviderRuntimeIngestion.ts#L1832-L1916), [active-row gating](https://github.com/pingdotgg/t3code/blob/d2c9281b8112dc3b2991642c4bdb985e4b08b9bb/apps/web/src/components/chat/MessagesTimeline.logic.ts#L1022-L1095), [status normalization](https://github.com/pingdotgg/t3code/blob/d2c9281b8112dc3b2991642c4bdb985e4b08b9bb/packages/client-runtime/src/work-log/presentation.ts#L370-L394). + +**Decision inference.** Bound Q4 to call identity, observed native lifecycle/status/error facts, and one row changing in place. Let authoritative Turn settlement stop activity indicators; preserve unconfirmed calls as unconfirmed. Do not infer tool success from Turn success or implement a shared synthetic completion manager just to remove spinners. + +## Snapshot changes, verification evidence, and ownership + +Compared with [the existing conversation-mapping note](./t3code-conversation-mapping.md)'s `de251fc` snapshot, inspected output/projection/collapse paths are unchanged. The material lifecycle change is Claude Stop: it now asks the SDK to interrupt and waits up to three seconds for Turn settlement before closing the query. The in-flight-tool synthesis predates that change and was absent from the older note's summary. The projection comment claiming client collapse is adjacency-only is stale: executable code and its interleaved-call test already contradict it. [New Stop path](https://github.com/pingdotgg/t3code/blob/d2c9281b8112dc3b2991642c4bdb985e4b08b9bb/apps/server/src/provider/Layers/ClaudeAdapter.ts#L5286-L5313), [interleaved-call test](https://github.com/pingdotgg/t3code/blob/d2c9281b8112dc3b2991642c4bdb985e4b08b9bb/apps/web/src/session-logic.test.ts#L974-L1045). + +Existing focused tests assert dropped command/file output deltas, compact Codex output projection, completed Codex/Claude output derivation, failed/declined native statuses, and identified live coalescing. They were inspected, not executed. No integrated test was found proving live shell stdout reaches this web renderer; the ingestion test explicitly asserts the opposite. [Dropped-delta test](https://github.com/pingdotgg/t3code/blob/d2c9281b8112dc3b2991642c4bdb985e4b08b9bb/apps/server/src/orchestration/Layers/ProviderRuntimeIngestion.test.ts#L1362-L1383), [projection test](https://github.com/pingdotgg/t3code/blob/d2c9281b8112dc3b2991642c4bdb985e4b08b9bb/apps/server/src/orchestration/ActivityPayloadProjection.test.ts#L33-L64), [command derivation tests](https://github.com/pingdotgg/t3code/blob/d2c9281b8112dc3b2991642c4bdb985e4b08b9bb/apps/web/src/session-logic.command-output.test.ts#L20-L73), [native outcome tests](https://github.com/pingdotgg/t3code/blob/d2c9281b8112dc3b2991642c4bdb985e4b08b9bb/apps/server/src/provider/Layers/CodexAdapter.test.ts#L1537-L1641), [coalescing tests](https://github.com/pingdotgg/t3code/blob/d2c9281b8112dc3b2991642c4bdb985e4b08b9bb/apps/server/src/orchestration/ThreadLiveEventCoalescer.test.ts#L93-L131). + +**Design assessment, not upstream fact.** Stable call identity belongs behind the Harness Interface; Claude and Codex are two real Adapters at that Seam. Bounded accumulation, row identity, final reconciliation, and Turn-derived liveness belong behind Secant's Projection Port, yielding Leverage to both TUI and headless callers and Locality for tests. Presentation should choose expansion and colours, not recover missing tool truth. T3's split server/client identity fallbacks, output failure heuristics, and Adapter-specific fabricated outcomes show complexity worth avoiding. The deletion test supports a deep Projection Module that owns the row invariant; a pass-through normalizer plus client lifecycle reducers would redistribute it. This matches [ADR 0024](../adr/0024-use-one-deep-projection-port-for-tui-and-headless-clients.md) and [ADR 0036](../adr/0036-the-run-workbench-mirrors-the-agent.md); this note makes no new domain decision.