feat(conversation): follow Stream's AI protocol for streamed replies - #683
Merged
martinmitrevski merged 4 commits intoSep 30, 2026
Conversation
Assistant replies are created with ai_generated: true, which is how Stream's AI components (StreamingMessageView and friends) tell a streamed reply from other messages. User messages are written as the agent too, so the field is set on the assistant's message only. Live updates go out every 100 ms instead of 200 ms, within Stream's guidance for streamed text, so clients animate the reply smoothly. The conversation now sends ai_indicator.update while a reply works (AI_STATE_THINKING, AI_STATE_EXTERNAL_SOURCES while a search runs, AI_STATE_GENERATING once the answer streams) and ai_indicator.clear once the final text is stored. The events are sent from the outbox loop after every pending write succeeds, only when the state changes, and through the SDK's typed SendEvent so Stream flattens ai_state and message_id to the top level where clients read them. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Brings the CI repair from accelerate onto the Athena runtime line: the dashboard job skips when there is no dashboard, Python CI runs on 3.12.13 with the locked wheels, the Python-versions check syncs only the dev tools, and the lint, typing and Go test fixes that apply here. Adapted to this line: - uv.lock is kept as it is here. The upstream change dropped the lemonslice and liveavatar plugins, which this line still has. - The shared-conversation session test keeps this line's expected author; the upstream expectation matches accelerate's author handling, not this line's. - The simulation config gains llm-flow only; the other aliases upstream belong to accelerate's newer routing. - aa_stats.py and sdks/kotlin/generate.py do not exist here and stay absent. (cherry picked from commit a518c0d) Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
martinmitrevski
marked this pull request as ready for review
September 29, 2026 15:54
martinmitrevski
merged commit Sep 30, 2026
f768457
into
athena/live-reasoning-ephemeral-runtime
15 of 17 checks passed
Nash0x7E2
pushed a commit
that referenced
this pull request
Oct 1, 2026
…683) * feat(conversation): follow Stream's AI protocol for streamed replies Assistant replies are created with ai_generated: true, which is how Stream's AI components (StreamingMessageView and friends) tell a streamed reply from other messages. User messages are written as the agent too, so the field is set on the assistant's message only. Live updates go out every 100 ms instead of 200 ms, within Stream's guidance for streamed text, so clients animate the reply smoothly. The conversation now sends ai_indicator.update while a reply works (AI_STATE_THINKING, AI_STATE_EXTERNAL_SOURCES while a search runs, AI_STATE_GENERATING once the answer streams) and ai_indicator.clear once the final text is stored. The events are sent from the outbox loop after every pending write succeeds, only when the state changes, and through the SDK's typed SendEvent so Stream flattens ai_state and message_id to the top level where clients read them. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * docs(changelog): note Stream's AI protocol for streamed replies Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Restore CI on accelerate (#676) Brings the CI repair from accelerate onto the Athena runtime line: the dashboard job skips when there is no dashboard, Python CI runs on 3.12.13 with the locked wheels, the Python-versions check syncs only the dev tools, and the lint, typing and Go test fixes that apply here. Adapted to this line: - uv.lock is kept as it is here. The upstream change dropped the lemonslice and liveavatar plugins, which this line still has. - The shared-conversation session test keeps this line's expected author; the upstream expectation matches accelerate's author handling, not this line's. - The simulation config gains llm-flow only; the other aliases upstream belong to accelerate's newer routing. - aa_stats.py and sdks/kotlin/generate.py do not exist here and stay absent. (cherry picked from commit a518c0d) Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * style(stream): format dispatch.py for the CI ruff check Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Co-authored-by: Marcelo Pires <thesyncim@gmail.com>
Nash0x7E2
pushed a commit
that referenced
this pull request
Oct 1, 2026
…683) * feat(conversation): follow Stream's AI protocol for streamed replies Assistant replies are created with ai_generated: true, which is how Stream's AI components (StreamingMessageView and friends) tell a streamed reply from other messages. User messages are written as the agent too, so the field is set on the assistant's message only. Live updates go out every 100 ms instead of 200 ms, within Stream's guidance for streamed text, so clients animate the reply smoothly. The conversation now sends ai_indicator.update while a reply works (AI_STATE_THINKING, AI_STATE_EXTERNAL_SOURCES while a search runs, AI_STATE_GENERATING once the answer streams) and ai_indicator.clear once the final text is stored. The events are sent from the outbox loop after every pending write succeeds, only when the state changes, and through the SDK's typed SendEvent so Stream flattens ai_state and message_id to the top level where clients read them. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * docs(changelog): note Stream's AI protocol for streamed replies Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Restore CI on accelerate (#676) Brings the CI repair from accelerate onto the Athena runtime line: the dashboard job skips when there is no dashboard, Python CI runs on 3.12.13 with the locked wheels, the Python-versions check syncs only the dev tools, and the lint, typing and Go test fixes that apply here. Adapted to this line: - uv.lock is kept as it is here. The upstream change dropped the lemonslice and liveavatar plugins, which this line still has. - The shared-conversation session test keeps this line's expected author; the upstream expectation matches accelerate's author handling, not this line's. - The simulation config gains llm-flow only; the other aliases upstream belong to accelerate's newer routing. - aa_stats.py and sdks/kotlin/generate.py do not exist here and stay absent. (cherry picked from commit a518c0d) Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * style(stream): format dispatch.py for the CI ruff check Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Co-authored-by: Marcelo Pires <thesyncim@gmail.com>
Nash0x7E2
added a commit
that referenced
this pull request
Oct 1, 2026
…indows, client tools and pacing (#706) Ports #683 and #704 from athena/live-reasoning-ephemeral-runtime onto accelerate: ai_reasoning and ai_tool_call steps, reasoning windows clients append, tools that run on a person's device, paced live updates with backoff, and Stream's AI protocol. Which tool calls become steps follows the agent config's visible_tools. Co-authored-by: Martin Mitrevski <martinmitrevski.oh@gmail.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Why
Athena's iOS app now renders replies with Stream's AI components (GetStream/athena-ai#5), which stream a reply character by character through
StreamingMessageViewand show what the agent is doing withAITypingIndicatorView. Persistent conversations already stream replies the way Stream's AI docs describe: a placeholder message,EphemeralMessageUpdates while it is live, and one storedUpdateMessagePartialwithgenerating: falsewhen it settles. Two parts of the protocol were missing:ai_generatedflag: Stream's AI components useai_generated: trueon a message to recognise a streamed AI reply, and replies didn't carry it.ai_indicator.*events, so clients couldn't show thinking or searching before the first token.The live update interval was also above the 50–100 ms that Stream recommends.
This targets
athena/live-reasoning-ephemeral-runtime, the Athena runtime line it is built on. That line is not inaccelerate, so a PR againstacceleratewould include its 50-odd commits as well.Changes
ai_generated: true. User messages, which are also written as the agent, don't get it.ai_indicator.updateonly when a reply's state changes:AI_STATE_THINKING, thenAI_STATE_EXTERNAL_SOURCESwhile a search runs, thenAI_STATE_GENERATINGonce text streams. It sendsai_indicator.clearafter every pending write, including the final text, has been stored. The events use the SDK's typedSendEventwithCustom, which Stream delivers withai_stateandmessage_idat the top level, where the clients read them.🤖 Generated with Claude Code