Skip to content

feat(conversation): follow Stream's AI protocol for streamed replies - #683

Merged
martinmitrevski merged 4 commits into
athena/live-reasoning-ephemeral-runtimefrom
athena/ai-message-streaming
Sep 30, 2026
Merged

martinmitrevski merged 4 commits into
athena/live-reasoning-ephemeral-runtimefrom
athena/ai-message-streaming

Conversation

@martinmitrevski

Copy link
Copy Markdown

Why

Athena's iOS app now renders replies with Stream's AI components (GetStream/athena-ai#5), which stream a reply character by character through StreamingMessageView and show what the agent is doing with AITypingIndicatorView. Persistent conversations already stream replies the way Stream's AI docs describe: a placeholder message, EphemeralMessageUpdates while it is live, and one stored UpdateMessagePartial with generating: false when it settles. Two parts of the protocol were missing:

  • ai_generated flag: Stream's AI components use ai_generated: true on a message to recognise a streamed AI reply, and replies didn't carry it.
  • Indicator events: there were no ai_indicator.* events, so clients couldn't show thinking or searching before the first token.

The live update interval was also above the 50–100 ms that Stream recommends.

This targets athena/live-reasoning-ephemeral-runtime, the Athena runtime line it is built on. That line is not in accelerate, so a PR against accelerate would include its 50-odd commits as well.

Changes

  • Assistant replies are created with ai_generated: true. User messages, which are also written as the agent, don't get it.
  • Live updates go out every 100 ms instead of 200 ms. The only stored write per reply is still the settled one.
  • The outbox loop sends ai_indicator.update only when a reply's state changes: AI_STATE_THINKING, then AI_STATE_EXTERNAL_SOURCES while a search runs, then AI_STATE_GENERATING once text streams. It sends ai_indicator.clear after every pending write, including the final text, has been stored. The events use the SDK's typed SendEvent with Custom, which Stream delivers with ai_state and message_id at the top level, where the clients read them.
  • A new conversation test checks the flag, the order of indicator states, and that the clear only arrives after the stored final write.
  • Changelog entry under "Athena integration fixes".

🤖 Generated with Claude Code

martinmitrevski and others added 2 commits September 29, 2026 16:24
Assistant replies are created with ai_generated: true, which is how Stream's
AI components (StreamingMessageView and friends) tell a streamed reply from
other messages. User messages are written as the agent too, so the field is
set on the assistant's message only.

Live updates go out every 100 ms instead of 200 ms, within Stream's guidance
for streamed text, so clients animate the reply smoothly.

The conversation now sends ai_indicator.update while a reply works
(AI_STATE_THINKING, AI_STATE_EXTERNAL_SOURCES while a search runs,
AI_STATE_GENERATING once the answer streams) and ai_indicator.clear once the
final text is stored. The events are sent from the outbox loop after every
pending write succeeds, only when the state changes, and through the SDK's
typed SendEvent so Stream flattens ai_state and message_id to the top level
where clients read them.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
thesyncim and others added 2 commits September 29, 2026 17:42
Brings the CI repair from accelerate onto the Athena runtime line: the
dashboard job skips when there is no dashboard, Python CI runs on 3.12.13
with the locked wheels, the Python-versions check syncs only the dev tools,
and the lint, typing and Go test fixes that apply here.

Adapted to this line:
- uv.lock is kept as it is here. The upstream change dropped the
  lemonslice and liveavatar plugins, which this line still has.
- The shared-conversation session test keeps this line's expected author;
  the upstream expectation matches accelerate's author handling, not this
  line's.
- The simulation config gains llm-flow only; the other aliases upstream
  belong to accelerate's newer routing.
- aa_stats.py and sdks/kotlin/generate.py do not exist here and stay absent.

(cherry picked from commit a518c0d)

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@martinmitrevski
martinmitrevski marked this pull request as ready for review September 29, 2026 15:54
@martinmitrevski
martinmitrevski merged commit f768457 into athena/live-reasoning-ephemeral-runtime Sep 30, 2026
15 of 17 checks passed
@martinmitrevski
martinmitrevski deleted the athena/ai-message-streaming branch September 30, 2026 07:12
Nash0x7E2 pushed a commit that referenced this pull request Oct 1, 2026
…683)

* feat(conversation): follow Stream's AI protocol for streamed replies

Assistant replies are created with ai_generated: true, which is how Stream's
AI components (StreamingMessageView and friends) tell a streamed reply from
other messages. User messages are written as the agent too, so the field is
set on the assistant's message only.

Live updates go out every 100 ms instead of 200 ms, within Stream's guidance
for streamed text, so clients animate the reply smoothly.

The conversation now sends ai_indicator.update while a reply works
(AI_STATE_THINKING, AI_STATE_EXTERNAL_SOURCES while a search runs,
AI_STATE_GENERATING once the answer streams) and ai_indicator.clear once the
final text is stored. The events are sent from the outbox loop after every
pending write succeeds, only when the state changes, and through the SDK's
typed SendEvent so Stream flattens ai_state and message_id to the top level
where clients read them.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* docs(changelog): note Stream's AI protocol for streamed replies

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Restore CI on accelerate (#676)

Brings the CI repair from accelerate onto the Athena runtime line: the
dashboard job skips when there is no dashboard, Python CI runs on 3.12.13
with the locked wheels, the Python-versions check syncs only the dev tools,
and the lint, typing and Go test fixes that apply here.

Adapted to this line:
- uv.lock is kept as it is here. The upstream change dropped the
  lemonslice and liveavatar plugins, which this line still has.
- The shared-conversation session test keeps this line's expected author;
  the upstream expectation matches accelerate's author handling, not this
  line's.
- The simulation config gains llm-flow only; the other aliases upstream
  belong to accelerate's newer routing.
- aa_stats.py and sdks/kotlin/generate.py do not exist here and stay absent.

(cherry picked from commit a518c0d)

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* style(stream): format dispatch.py for the CI ruff check

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: Marcelo Pires <thesyncim@gmail.com>
Nash0x7E2 pushed a commit that referenced this pull request Oct 1, 2026
…683)

* feat(conversation): follow Stream's AI protocol for streamed replies

Assistant replies are created with ai_generated: true, which is how Stream's
AI components (StreamingMessageView and friends) tell a streamed reply from
other messages. User messages are written as the agent too, so the field is
set on the assistant's message only.

Live updates go out every 100 ms instead of 200 ms, within Stream's guidance
for streamed text, so clients animate the reply smoothly.

The conversation now sends ai_indicator.update while a reply works
(AI_STATE_THINKING, AI_STATE_EXTERNAL_SOURCES while a search runs,
AI_STATE_GENERATING once the answer streams) and ai_indicator.clear once the
final text is stored. The events are sent from the outbox loop after every
pending write succeeds, only when the state changes, and through the SDK's
typed SendEvent so Stream flattens ai_state and message_id to the top level
where clients read them.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* docs(changelog): note Stream's AI protocol for streamed replies

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Restore CI on accelerate (#676)

Brings the CI repair from accelerate onto the Athena runtime line: the
dashboard job skips when there is no dashboard, Python CI runs on 3.12.13
with the locked wheels, the Python-versions check syncs only the dev tools,
and the lint, typing and Go test fixes that apply here.

Adapted to this line:
- uv.lock is kept as it is here. The upstream change dropped the
  lemonslice and liveavatar plugins, which this line still has.
- The shared-conversation session test keeps this line's expected author;
  the upstream expectation matches accelerate's author handling, not this
  line's.
- The simulation config gains llm-flow only; the other aliases upstream
  belong to accelerate's newer routing.
- aa_stats.py and sdks/kotlin/generate.py do not exist here and stay absent.

(cherry picked from commit a518c0d)

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* style(stream): format dispatch.py for the CI ruff check

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: Marcelo Pires <thesyncim@gmail.com>
Nash0x7E2 added a commit that referenced this pull request Oct 1, 2026
…indows, client tools and pacing (#706)

Ports #683 and #704 from athena/live-reasoning-ephemeral-runtime onto accelerate: ai_reasoning and ai_tool_call steps, reasoning windows clients append, tools that run on a person's device, paced live updates with backoff, and Stream's AI protocol. Which tool calls become steps follows the agent config's visible_tools.

Co-authored-by: Martin Mitrevski <martinmitrevski.oh@gmail.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants