Skip to content

feat(conversation): reply steps as ai_* attachments, live reasoning deltas, client tools and paced live updates - #704

Merged
martinmitrevski merged 3 commits into
athena/live-reasoning-ephemeral-runtimefrom
athena/live-reasoning-deltas
Oct 1, 2026
Merged

martinmitrevski merged 3 commits into
athena/live-reasoning-ephemeral-runtimefrom
athena/live-reasoning-deltas

Conversation

@martinmitrevski

@martinmitrevski martinmitrevski commented Oct 1, 2026 •

Copy link
Copy Markdown

Stacked on athena/live-reasoning-ephemeral-runtime (where #683 landed).

A persistent reply's steps (rounds of reasoning, tool calls) become Stream attachments, its live reasoning streams as deltas instead of a resent 4 KB tail, a caller can declare tools that run on a person's device, and live updates stay within Stream's limits.

What changes

Reasoning windows. The ephemeral reasoning field was the latest 4,000 bytes, resent on every update. It is now {id, offset, text, length} for the reasoning step id: the new thinking since the last delivered window, in Unicode scalars, which clients append.

  • At most 2,000 bytes per window; a longer backlog goes out over several updates.
  • Every 3 s a window repeats the last 1,500 bytes, so a watcher who opened the conversation midway or reconnected catches up.
  • The last thoughts go out once after the reply settles.

Reply steps as attachments. Each round of thinking is an ai_reasoning attachment and each visible tool call an ai_tool_call attachment, in order, followed by the artifacts. They ride on the live updates and are stored with the settled reply.

  • A completed round stores its first sentence (summary) and first 500 characters (preview); the rest of the thinking is only shown live.
  • Tool calls use the provider's call ID. Server tools never show arguments or results.
  • Everything fits Stream's limit of 30 attachments and 5 KB per message: artifacts are kept whole, and older steps' previews and summaries give way first.

Client tools.

  • SessionTool gains executor (server or client) and display_title.
  • RespondRequest gains client_id, the install a command came from, written on the person's message.
  • A client tool's call is shown as awaiting_client, addressed to the command's initiator and install, with its arguments (at most 512 bytes). The caller answers it over the events socket once the device reports, and a summary (or the device's failure reason) is shown on the step.

Pacing and backoff.

  • The answer goes out at most every 150 ms (Stream throttles message.updated to 10/s per channel) and thinking alone at most every 200 ms, on a 50 ms tick.
  • After a refused live update the next one waits: Retry-After or the rate-limit window's reset on a 429 (at most a minute), otherwise 1 s doubling to 8 s. Stored writes keep their own retry, and nothing held back is lost.

Fix. A data race between sending a reply's steps and updating them. The send and the session broadcast shared the live steps' array; they now copy it.

Generated code. The OpenAPI spec changes, so internal/api/generated.go and the Go and JS SDKs are regenerated. The Go SDK also picks up the earlier image endpoint it had missed.

Verification

  • internal/conversation (new tests: windows, keyframes, backlog, gaps, steps, client-tool targeting, attachment limits, pacing, a 429 backoff against the fake Stream server); internal/conversation ×3, internal/session and internal/api with -race. internal/agent and internal/harness pass without -race.
    • TestAgentSuite/TestASecondToolDoesNotStartACompetingReply fails under -race on this branch's base as well.
  • Local Athena stack, DeepSeek V4 Pro:
    • 27 KB of reasoning over 105 s arrived in 588 windows with no gaps: 75.7 KB sent, against about 3.9 MB under the old scheme.
    • Live updates: the answer at 6.2/s, thinking-only at a 196 ms median, no failed updates. The ephemeral endpoint reported a 10,000/min app budget.
    • A device tool call went awaiting_client → completed through the iOS app.
  • Not verified: web search between rounds (no Exa key locally), a runtime restart mid-reply, anything hosted.

Athena client and API: GetStream/athena-ai#14. Stream AI components: stream-chat-swift-ai 0.9.0.

🤖 Generated with Claude Code

martinmitrevski and others added 3 commits October 1, 2026 17:25
Every live update used to repeat the latest 4,000 bytes of the model's thinking,
so a long thought cost the same bytes ten times a second and watchers lost
everything before its tail. The ephemeral "reasoning" field is now a window,
{id, offset, text, length, duration_ms}: the thinking from where the last
delivered window ended, in Unicode scalars like answer_start, so a client only
appends and never counts what it has.

- A window holds at most 2,000 bytes, leaving room within Stream's custom data
  limit; a longer backlog goes out over several ticks, and one past the 32 KB
  buffer skips ahead so watchers see a gap.
- Every 3 s a window repeats the last 1,500 bytes, so someone who opens the
  conversation midway or reconnects catches up.
- Thinking alone is sent at most every 200 ms; the answer keeps the 100 ms
  cadence and carries any new thinking with it. Thinking no longer bumps the
  reply's sequence or republishes it to session watchers.
- Each model round starts a new paragraph, and a finished reply's last thoughts
  go out once after its final text is stored.

Thinking is still never written to the outbox, the ledger or the settled message.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…call attachments

A persistent reply's steps are now its Stream attachments, in order: each round
of the model's thinking as ai_reasoning and each tool call people may see as
ai_tool_call, followed by the reply's artifacts. The order is the timeline and
the answer stays in the message text. The runtime is the only writer: the steps
ride on the live updates while a reply works and are stored with the settled
reply.

- A reasoning step stores its first sentence as its summary and its first 500
  characters as its preview; while it streams it carries its latest 200. The
  whole of the thinking is still only shown live, through the reasoning
  windows, which now name the step they belong to.
- Stream allows 30 attachments on a message, 5 KB together. Artifacts are kept
  whole; older steps' previews shrink and then go, then summaries, then the
  oldest finished steps.
- Tool steps use the provider's call ID. Only tools already shown to people
  (Athena's own and web search) and client tools become steps, and a server
  tool's arguments and results are never shown.

A caller can now declare client tools, which run on a person's device:
SessionTool gains executor (server or client) and display_title, and
RespondRequest gains client_id, the install a command came from, written on
the person's message. A client tool's call is shown as awaiting_client,
addressed to the command's initiator and install, with its arguments (at most
512 bytes, since every member can read them). The caller answers it over the
events socket once the device reports, and a summary in its result is shown on
the step.

The Go and JavaScript SDKs are regenerated from the spec; the Go client had not
yet picked up the image generation endpoint either.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
… them

A failed live (ephemeral) update was retried on the next 100 ms tick, so a
rate-limited Stream was asked again ten times a second. Now the next live update
waits: Stream's Retry-After, or the end of its rate-limit window, on a 429 (at
most a minute), and otherwise one second doubling to eight, reset by a success.
Stored writes keep their own two-second retry, and nothing held back is lost:
the next update carries the reply as it is by then, and thinking counts as
delivered only once Stream accepts it.

Stream throttles message.updated to 10 a second per channel, and the answer went
out every 100 ms, right at that limit. It now goes out at most every 150 ms,
leaving room for the channel's other updates; thinking alone keeps its 200 ms
pace. The loop ticks every 50 ms so both paces hold closely.

Sending a reply's steps raced with updating them: the copy taken for the send
outside the lock, and the one published to session watchers, shared the live
steps' array. Both, and the stored-write queue, now copy the steps.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@martinmitrevski
martinmitrevski merged commit 04d9fcf into athena/live-reasoning-ephemeral-runtime Oct 1, 2026
10 of 12 checks passed
@martinmitrevski
martinmitrevski deleted the athena/live-reasoning-deltas branch October 1, 2026 15:52
Nash0x7E2 added a commit that referenced this pull request Oct 1, 2026
…thinking

Ports the parts of the Athena runtime line's 2d40a96 that #704 builds on. Tool starts and finishes are checkpointed to the local ledger, which restart recovery reads, and reach watchers through ephemeral updates, so Stream Chat is written only when a reply is created and when it settles. A written reply's reasoning (agent.ReasoningDelta) rides on the ephemeral updates and is never stored.
Nash0x7E2 pushed a commit that referenced this pull request Oct 1, 2026
…eltas, client tools and paced live updates (#704)

* feat(conversation): stream live thinking in windows clients append

Every live update used to repeat the latest 4,000 bytes of the model's thinking,
so a long thought cost the same bytes ten times a second and watchers lost
everything before its tail. The ephemeral "reasoning" field is now a window,
{id, offset, text, length, duration_ms}: the thinking from where the last
delivered window ended, in Unicode scalars like answer_start, so a client only
appends and never counts what it has.

- A window holds at most 2,000 bytes, leaving room within Stream's custom data
  limit; a longer backlog goes out over several ticks, and one past the 32 KB
  buffer skips ahead so watchers see a gap.
- Every 3 s a window repeats the last 1,500 bytes, so someone who opens the
  conversation midway or reconnects catches up.
- Thinking alone is sent at most every 200 ms; the answer keeps the 100 ms
  cadence and carries any new thinking with it. Thinking no longer bumps the
  reply's sequence or republishes it to session watchers.
- Each model round starts a new paragraph, and a finished reply's last thoughts
  go out once after its final text is stored.

Thinking is still never written to the outbox, the ledger or the settled message.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* feat(conversation): show a reply's steps as ai_reasoning and ai_tool_call attachments

A persistent reply's steps are now its Stream attachments, in order: each round
of the model's thinking as ai_reasoning and each tool call people may see as
ai_tool_call, followed by the reply's artifacts. The order is the timeline and
the answer stays in the message text. The runtime is the only writer: the steps
ride on the live updates while a reply works and are stored with the settled
reply.

- A reasoning step stores its first sentence as its summary and its first 500
  characters as its preview; while it streams it carries its latest 200. The
  whole of the thinking is still only shown live, through the reasoning
  windows, which now name the step they belong to.
- Stream allows 30 attachments on a message, 5 KB together. Artifacts are kept
  whole; older steps' previews shrink and then go, then summaries, then the
  oldest finished steps.
- Tool steps use the provider's call ID. Only tools already shown to people
  (Athena's own and web search) and client tools become steps, and a server
  tool's arguments and results are never shown.

A caller can now declare client tools, which run on a person's device:
SessionTool gains executor (server or client) and display_title, and
RespondRequest gains client_id, the install a command came from, written on
the person's message. A client tool's call is shown as awaiting_client,
addressed to the command's initiator and install, with its arguments (at most
512 bytes, since every member can read them). The caller answers it over the
events socket once the device reports, and a summary in its result is shown on
the step.

The Go and JavaScript SDKs are regenerated from the spec; the Go client had not
yet picked up the image generation endpoint either.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(conversation): pace live updates and back off when Stream refuses them

A failed live (ephemeral) update was retried on the next 100 ms tick, so a
rate-limited Stream was asked again ten times a second. Now the next live update
waits: Stream's Retry-After, or the end of its rate-limit window, on a 429 (at
most a minute), and otherwise one second doubling to eight, reset by a success.
Stored writes keep their own two-second retry, and nothing held back is lost:
the next update carries the reply as it is by then, and thinking counts as
delivered only once Stream accepts it.

Stream throttles message.updated to 10 a second per channel, and the answer went
out every 100 ms, right at that limit. It now goes out at most every 150 ms,
leaving room for the channel's other updates; thinking alone keeps its 200 ms
pace. The loop ticks every 50 ms so both paces hold closely.

Sending a reply's steps raced with updating them: the copy taken for the send
outside the lock, and the one published to session watchers, shared the live
steps' array. Both, and the stored-write queue, now copy the steps.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Nash0x7E2 added a commit that referenced this pull request Oct 1, 2026
…thinking

Ports the parts of the Athena runtime line's 2d40a96 that #704 builds on. Tool starts and finishes are checkpointed to the local ledger, which restart recovery reads, and reach watchers through ephemeral updates, so Stream Chat is written only when a reply is created and when it settles. A written reply's reasoning (agent.ReasoningDelta) rides on the ephemeral updates and is never stored.
Nash0x7E2 pushed a commit that referenced this pull request Oct 1, 2026
…eltas, client tools and paced live updates (#704)

* feat(conversation): stream live thinking in windows clients append

Every live update used to repeat the latest 4,000 bytes of the model's thinking,
so a long thought cost the same bytes ten times a second and watchers lost
everything before its tail. The ephemeral "reasoning" field is now a window,
{id, offset, text, length, duration_ms}: the thinking from where the last
delivered window ended, in Unicode scalars like answer_start, so a client only
appends and never counts what it has.

- A window holds at most 2,000 bytes, leaving room within Stream's custom data
  limit; a longer backlog goes out over several ticks, and one past the 32 KB
  buffer skips ahead so watchers see a gap.
- Every 3 s a window repeats the last 1,500 bytes, so someone who opens the
  conversation midway or reconnects catches up.
- Thinking alone is sent at most every 200 ms; the answer keeps the 100 ms
  cadence and carries any new thinking with it. Thinking no longer bumps the
  reply's sequence or republishes it to session watchers.
- Each model round starts a new paragraph, and a finished reply's last thoughts
  go out once after its final text is stored.

Thinking is still never written to the outbox, the ledger or the settled message.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* feat(conversation): show a reply's steps as ai_reasoning and ai_tool_call attachments

A persistent reply's steps are now its Stream attachments, in order: each round
of the model's thinking as ai_reasoning and each tool call people may see as
ai_tool_call, followed by the reply's artifacts. The order is the timeline and
the answer stays in the message text. The runtime is the only writer: the steps
ride on the live updates while a reply works and are stored with the settled
reply.

- A reasoning step stores its first sentence as its summary and its first 500
  characters as its preview; while it streams it carries its latest 200. The
  whole of the thinking is still only shown live, through the reasoning
  windows, which now name the step they belong to.
- Stream allows 30 attachments on a message, 5 KB together. Artifacts are kept
  whole; older steps' previews shrink and then go, then summaries, then the
  oldest finished steps.
- Tool steps use the provider's call ID. Only tools already shown to people
  (Athena's own and web search) and client tools become steps, and a server
  tool's arguments and results are never shown.

A caller can now declare client tools, which run on a person's device:
SessionTool gains executor (server or client) and display_title, and
RespondRequest gains client_id, the install a command came from, written on
the person's message. A client tool's call is shown as awaiting_client,
addressed to the command's initiator and install, with its arguments (at most
512 bytes, since every member can read them). The caller answers it over the
events socket once the device reports, and a summary in its result is shown on
the step.

The Go and JavaScript SDKs are regenerated from the spec; the Go client had not
yet picked up the image generation endpoint either.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(conversation): pace live updates and back off when Stream refuses them

A failed live (ephemeral) update was retried on the next 100 ms tick, so a
rate-limited Stream was asked again ten times a second. Now the next live update
waits: Stream's Retry-After, or the end of its rate-limit window, on a 429 (at
most a minute), and otherwise one second doubling to eight, reset by a success.
Stored writes keep their own two-second retry, and nothing held back is lost:
the next update carries the reply as it is by then, and thinking counts as
delivered only once Stream accepts it.

Stream throttles message.updated to 10 a second per channel, and the answer went
out every 100 ms, right at that limit. It now goes out at most every 150 ms,
leaving room for the channel's other updates; thinking alone keeps its 200 ms
pace. The loop ticks every 50 ms so both paces hold closely.

Sending a reply's steps raced with updating them: the copy taken for the send
outside the lock, and the one published to session watchers, shared the live
steps' array. Both, and the stored-write queue, now copy the steps.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Nash0x7E2 added a commit that referenced this pull request Oct 1, 2026
…indows, client tools and pacing (#706)

Ports #683 and #704 from athena/live-reasoning-ephemeral-runtime onto accelerate: ai_reasoning and ai_tool_call steps, reasoning windows clients append, tools that run on a person's device, paced live updates with backoff, and Stream's AI protocol. Which tool calls become steps follows the agent config's visible_tools.

Co-authored-by: Martin Mitrevski <martinmitrevski.oh@gmail.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant