feat(conversation): reply steps as ai_* attachments, live reasoning windows, client tools and pacing (port of #683, #704) - #706
Merged
Conversation
Nash0x7E2
force-pushed
the
nash/conversation-ai-parts
branch
from
October 1, 2026 19:09
c5a2cca to
64d473c
Compare
Nash0x7E2
force-pushed
the
nash/conversation-artifacts
branch
from
October 1, 2026 19:45
9b4dcf5 to
77b77b6
Compare
…683) * feat(conversation): follow Stream's AI protocol for streamed replies Assistant replies are created with ai_generated: true, which is how Stream's AI components (StreamingMessageView and friends) tell a streamed reply from other messages. User messages are written as the agent too, so the field is set on the assistant's message only. Live updates go out every 100 ms instead of 200 ms, within Stream's guidance for streamed text, so clients animate the reply smoothly. The conversation now sends ai_indicator.update while a reply works (AI_STATE_THINKING, AI_STATE_EXTERNAL_SOURCES while a search runs, AI_STATE_GENERATING once the answer streams) and ai_indicator.clear once the final text is stored. The events are sent from the outbox loop after every pending write succeeds, only when the state changes, and through the SDK's typed SendEvent so Stream flattens ai_state and message_id to the top level where clients read them. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * docs(changelog): note Stream's AI protocol for streamed replies Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Restore CI on accelerate (#676) Brings the CI repair from accelerate onto the Athena runtime line: the dashboard job skips when there is no dashboard, Python CI runs on 3.12.13 with the locked wheels, the Python-versions check syncs only the dev tools, and the lint, typing and Go test fixes that apply here. Adapted to this line: - uv.lock is kept as it is here. The upstream change dropped the lemonslice and liveavatar plugins, which this line still has. - The shared-conversation session test keeps this line's expected author; the upstream expectation matches accelerate's author handling, not this line's. - The simulation config gains llm-flow only; the other aliases upstream belong to accelerate's newer routing. - aa_stats.py and sdks/kotlin/generate.py do not exist here and stay absent. (cherry picked from commit a518c0d) Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * style(stream): format dispatch.py for the CI ruff check Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Co-authored-by: Marcelo Pires <thesyncim@gmail.com>
…thinking Ports the parts of the Athena runtime line's 2d40a96 that #704 builds on. Tool starts and finishes are checkpointed to the local ledger, which restart recovery reads, and reach watchers through ephemeral updates, so Stream Chat is written only when a reply is created and when it settles. A written reply's reasoning (agent.ReasoningDelta) rides on the ephemeral updates and is never stored.
…eltas, client tools and paced live updates (#704) * feat(conversation): stream live thinking in windows clients append Every live update used to repeat the latest 4,000 bytes of the model's thinking, so a long thought cost the same bytes ten times a second and watchers lost everything before its tail. The ephemeral "reasoning" field is now a window, {id, offset, text, length, duration_ms}: the thinking from where the last delivered window ended, in Unicode scalars like answer_start, so a client only appends and never counts what it has. - A window holds at most 2,000 bytes, leaving room within Stream's custom data limit; a longer backlog goes out over several ticks, and one past the 32 KB buffer skips ahead so watchers see a gap. - Every 3 s a window repeats the last 1,500 bytes, so someone who opens the conversation midway or reconnects catches up. - Thinking alone is sent at most every 200 ms; the answer keeps the 100 ms cadence and carries any new thinking with it. Thinking no longer bumps the reply's sequence or republishes it to session watchers. - Each model round starts a new paragraph, and a finished reply's last thoughts go out once after its final text is stored. Thinking is still never written to the outbox, the ledger or the settled message. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * feat(conversation): show a reply's steps as ai_reasoning and ai_tool_call attachments A persistent reply's steps are now its Stream attachments, in order: each round of the model's thinking as ai_reasoning and each tool call people may see as ai_tool_call, followed by the reply's artifacts. The order is the timeline and the answer stays in the message text. The runtime is the only writer: the steps ride on the live updates while a reply works and are stored with the settled reply. - A reasoning step stores its first sentence as its summary and its first 500 characters as its preview; while it streams it carries its latest 200. The whole of the thinking is still only shown live, through the reasoning windows, which now name the step they belong to. - Stream allows 30 attachments on a message, 5 KB together. Artifacts are kept whole; older steps' previews shrink and then go, then summaries, then the oldest finished steps. - Tool steps use the provider's call ID. Only tools already shown to people (Athena's own and web search) and client tools become steps, and a server tool's arguments and results are never shown. A caller can now declare client tools, which run on a person's device: SessionTool gains executor (server or client) and display_title, and RespondRequest gains client_id, the install a command came from, written on the person's message. A client tool's call is shown as awaiting_client, addressed to the command's initiator and install, with its arguments (at most 512 bytes, since every member can read them). The caller answers it over the events socket once the device reports, and a summary in its result is shown on the step. The Go and JavaScript SDKs are regenerated from the spec; the Go client had not yet picked up the image generation endpoint either. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(conversation): pace live updates and back off when Stream refuses them A failed live (ephemeral) update was retried on the next 100 ms tick, so a rate-limited Stream was asked again ten times a second. Now the next live update waits: Stream's Retry-After, or the end of its rate-limit window, on a 429 (at most a minute), and otherwise one second doubling to eight, reset by a success. Stored writes keep their own two-second retry, and nothing held back is lost: the next update carries the reply as it is by then, and thinking counts as delivered only once Stream accepts it. Stream throttles message.updated to 10 a second per channel, and the answer went out every 100 ms, right at that limit. It now goes out at most every 150 ms, leaving room for the channel's other updates; thinking alone keeps its 200 ms pace. The loop ticks every 50 ms so both paces hold closely. Sending a reply's steps raced with updating them: the copy taken for the send outside the lock, and the one published to session watchers, shared the live steps' array. Both, and the stored-write queue, now copy the steps. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
SessionTool gains executor and display_title and RespondRequest gains client_id. Go and JavaScript are regenerated; the rest follow.
Nash0x7E2
force-pushed
the
nash/conversation-ai-parts
branch
from
October 1, 2026 19:52
64d473c to
8d3f8ad
Compare
Nash0x7E2
marked this pull request as ready for review
October 1, 2026 19:52
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Why
Athena's clients on main now read a reply's steps as
ai_reasoning/ai_tool_callattachments, follow live reasoning as windows, and run device tools (Athena'sdocs/contracts/AI_MESSAGE_PARTS.md). That contract was built by Martin on the Athena runtime line (athena/live-reasoning-ephemeral-runtime) in #683 and #704, which are not on accelerate. This ports both onto the accelerate stack, on top of #700 (nash/conversation-artifacts), so accelerate's router serves what Athena reads. Thanks to Martin Mitrevski, whose commits are cherry-picked with his authorship.It supersedes #705, which streamed an older 4 KB reasoning tail behind a
show_reasoningconfig.Changes
ai_generated: true, and the outbox loop sendsai_indicator.update(thinking, external sources, generating) andai_indicator.clearonce the final text is stored.agent.ReasoningDelta. feat(conversation): reply steps as ai_* attachments, live reasoning deltas, client tools and paced live updates #704 builds on both.ai_reasoningattachment and each visible tool call anai_tool_call, in order, followed by the artifacts, within Stream's 30 attachments / 5 KB (rendered within 4,800 bytes; artifacts kept whole, older steps give way).reasoningfield is{id, offset, text, length}for the streaming step, at most 2,000 bytes, a 1,500-byte keyframe every 3 s, a 32 KB buffer that skips ahead, and the last thoughts sent once after the reply settles. Reasoning is always on for written replies, as Athena expects; there is no config gate.SessionToolgainsexecutor(serverorclient) anddisplay_title;RespondRequestgainsclient_id, written on the person's message. A client tool's call isawaiting_client, addressed to the command's initiator and install, with its arguments (at most 512 bytes) and asummaryfrom its result.Retry-Afteror the rate-limit reset on a 429 (at most a minute), otherwise 1 s doubling to 8 s. Includes the fix for the race on the live steps' array.visible_tools(plus client tools) instead of feat(conversation): reply steps as ai_* attachments, live reasoning deltas, client tools and paced live updates #704's hard-codedathena_*/ search names; artifacts keep the base'sChatAttachmentsshape; the user's message stays stored under the command ID (feat(acceleration): store a command's user message under the command ID #687); command receipts andRespondCommand's response ID are unchanged.legacy.yaml(sessions still live there), and the Go and JavaScript SDKs are regenerated. The sdk skill notes what the other SDKs need. No changelog on accelerate.PartsSuite,ReasoningSuiteandLiveSuite(windows, keyframes, backlog, gaps, steps, client-tool targeting, attachment limits, pacing, a 429 backoff against the fake Stream server).🤖 Generated with Claude Code
https://claude.ai/code/session_0198ttpps8iUcinspwFJg8rF