Skip to content

feat(conversation): reply steps as ai_* attachments, live reasoning windows, client tools and pacing (port of #683, #704) - #706

Merged
Nash0x7E2 merged 5 commits into
acceleratefrom
nash/conversation-ai-parts
Oct 1, 2026
Merged

Nash0x7E2 merged 5 commits into
acceleratefrom
nash/conversation-ai-parts

Conversation

@Nash0x7E2

Copy link
Copy Markdown
Member

Why

Athena's clients on main now read a reply's steps as ai_reasoning / ai_tool_call attachments, follow live reasoning as windows, and run device tools (Athena's docs/contracts/AI_MESSAGE_PARTS.md). That contract was built by Martin on the Athena runtime line (athena/live-reasoning-ephemeral-runtime) in #683 and #704, which are not on accelerate. This ports both onto the accelerate stack, on top of #700 (nash/conversation-artifacts), so accelerate's router serves what Athena reads. Thanks to Martin Mitrevski, whose commits are cherry-picked with his authorship.

It supersedes #705, which streamed an older 4 KB reasoning tail behind a show_reasoning config.

Changes

🤖 Generated with Claude Code

https://claude.ai/code/session_0198ttpps8iUcinspwFJg8rF

@Nash0x7E2
Nash0x7E2 force-pushed the nash/conversation-ai-parts branch from c5a2cca to 64d473c Compare October 1, 2026 19:09
@Nash0x7E2
Nash0x7E2 force-pushed the nash/conversation-artifacts branch from 9b4dcf5 to 77b77b6 Compare October 1, 2026 19:45
Base automatically changed from nash/conversation-artifacts to accelerate October 1, 2026 19:51
martinmitrevski and others added 5 commits October 1, 2026 13:51
…683)

* feat(conversation): follow Stream's AI protocol for streamed replies

Assistant replies are created with ai_generated: true, which is how Stream's
AI components (StreamingMessageView and friends) tell a streamed reply from
other messages. User messages are written as the agent too, so the field is
set on the assistant's message only.

Live updates go out every 100 ms instead of 200 ms, within Stream's guidance
for streamed text, so clients animate the reply smoothly.

The conversation now sends ai_indicator.update while a reply works
(AI_STATE_THINKING, AI_STATE_EXTERNAL_SOURCES while a search runs,
AI_STATE_GENERATING once the answer streams) and ai_indicator.clear once the
final text is stored. The events are sent from the outbox loop after every
pending write succeeds, only when the state changes, and through the SDK's
typed SendEvent so Stream flattens ai_state and message_id to the top level
where clients read them.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* docs(changelog): note Stream's AI protocol for streamed replies

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Restore CI on accelerate (#676)

Brings the CI repair from accelerate onto the Athena runtime line: the
dashboard job skips when there is no dashboard, Python CI runs on 3.12.13
with the locked wheels, the Python-versions check syncs only the dev tools,
and the lint, typing and Go test fixes that apply here.

Adapted to this line:
- uv.lock is kept as it is here. The upstream change dropped the
  lemonslice and liveavatar plugins, which this line still has.
- The shared-conversation session test keeps this line's expected author;
  the upstream expectation matches accelerate's author handling, not this
  line's.
- The simulation config gains llm-flow only; the other aliases upstream
  belong to accelerate's newer routing.
- aa_stats.py and sdks/kotlin/generate.py do not exist here and stay absent.

(cherry picked from commit a518c0d)

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* style(stream): format dispatch.py for the CI ruff check

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: Marcelo Pires <thesyncim@gmail.com>
…thinking

Ports the parts of the Athena runtime line's 2d40a96 that #704 builds on. Tool starts and finishes are checkpointed to the local ledger, which restart recovery reads, and reach watchers through ephemeral updates, so Stream Chat is written only when a reply is created and when it settles. A written reply's reasoning (agent.ReasoningDelta) rides on the ephemeral updates and is never stored.
…eltas, client tools and paced live updates (#704)

* feat(conversation): stream live thinking in windows clients append

Every live update used to repeat the latest 4,000 bytes of the model's thinking,
so a long thought cost the same bytes ten times a second and watchers lost
everything before its tail. The ephemeral "reasoning" field is now a window,
{id, offset, text, length, duration_ms}: the thinking from where the last
delivered window ended, in Unicode scalars like answer_start, so a client only
appends and never counts what it has.

- A window holds at most 2,000 bytes, leaving room within Stream's custom data
  limit; a longer backlog goes out over several ticks, and one past the 32 KB
  buffer skips ahead so watchers see a gap.
- Every 3 s a window repeats the last 1,500 bytes, so someone who opens the
  conversation midway or reconnects catches up.
- Thinking alone is sent at most every 200 ms; the answer keeps the 100 ms
  cadence and carries any new thinking with it. Thinking no longer bumps the
  reply's sequence or republishes it to session watchers.
- Each model round starts a new paragraph, and a finished reply's last thoughts
  go out once after its final text is stored.

Thinking is still never written to the outbox, the ledger or the settled message.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* feat(conversation): show a reply's steps as ai_reasoning and ai_tool_call attachments

A persistent reply's steps are now its Stream attachments, in order: each round
of the model's thinking as ai_reasoning and each tool call people may see as
ai_tool_call, followed by the reply's artifacts. The order is the timeline and
the answer stays in the message text. The runtime is the only writer: the steps
ride on the live updates while a reply works and are stored with the settled
reply.

- A reasoning step stores its first sentence as its summary and its first 500
  characters as its preview; while it streams it carries its latest 200. The
  whole of the thinking is still only shown live, through the reasoning
  windows, which now name the step they belong to.
- Stream allows 30 attachments on a message, 5 KB together. Artifacts are kept
  whole; older steps' previews shrink and then go, then summaries, then the
  oldest finished steps.
- Tool steps use the provider's call ID. Only tools already shown to people
  (Athena's own and web search) and client tools become steps, and a server
  tool's arguments and results are never shown.

A caller can now declare client tools, which run on a person's device:
SessionTool gains executor (server or client) and display_title, and
RespondRequest gains client_id, the install a command came from, written on
the person's message. A client tool's call is shown as awaiting_client,
addressed to the command's initiator and install, with its arguments (at most
512 bytes, since every member can read them). The caller answers it over the
events socket once the device reports, and a summary in its result is shown on
the step.

The Go and JavaScript SDKs are regenerated from the spec; the Go client had not
yet picked up the image generation endpoint either.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(conversation): pace live updates and back off when Stream refuses them

A failed live (ephemeral) update was retried on the next 100 ms tick, so a
rate-limited Stream was asked again ten times a second. Now the next live update
waits: Stream's Retry-After, or the end of its rate-limit window, on a 429 (at
most a minute), and otherwise one second doubling to eight, reset by a success.
Stored writes keep their own two-second retry, and nothing held back is lost:
the next update carries the reply as it is by then, and thinking counts as
delivered only once Stream accepts it.

Stream throttles message.updated to 10 a second per channel, and the answer went
out every 100 ms, right at that limit. It now goes out at most every 150 ms,
leaving room for the channel's other updates; thinking alone keeps its 200 ms
pace. The loop ticks every 50 ms so both paces hold closely.

Sending a reply's steps raced with updating them: the copy taken for the send
outside the lock, and the one published to session watchers, shared the live
steps' array. Both, and the stored-write queue, now copy the steps.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
SessionTool gains executor and display_title and RespondRequest gains client_id. Go and JavaScript are regenerated; the rest follow.
@Nash0x7E2
Nash0x7E2 force-pushed the nash/conversation-ai-parts branch from 64d473c to 8d3f8ad Compare October 1, 2026 19:52
@Nash0x7E2
Nash0x7E2 marked this pull request as ready for review October 1, 2026 19:52
@Nash0x7E2
Nash0x7E2 merged commit 5f1ede5 into accelerate Oct 1, 2026
9 of 18 checks passed
@Nash0x7E2
Nash0x7E2 deleted the nash/conversation-ai-parts branch October 1, 2026 19:54
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants