Skip to content

Start replies before the flow controller rules - #682

Merged
darkoatanasovski merged 3 commits into
acceleratefrom
accelerate-improvements
Sep 30, 2026
Merged

darkoatanasovski merged 3 commits into
acceleratefrom
accelerate-improvements

Conversation

@darkoatanasovski

Copy link
Copy Markdown

Why

Every answered turn paid for the flow decision and the reply one after the other: the reply model was only asked once the flow controller had ruled that the words were meant for the agent. The per-turn stage timing from #674 shows that decision at a median of about 1.2 s per turn, and in a Voicebench run with identical models and voice on both sides (Gemini transcribe-live, GPT-5.6 Luna, Inworld TTS-2 Flash, voice Ashley) LiveKit Inference replied about 1.1 s sooner than us. That gap is smaller than the decision itself, so it's the router's turn-taking, not the providers. The benchmark tooling is in #667.

With speculation on, the same matched restaurant run (8 calls per side) gave:

Off On LiveKit Inference
First response P50 4.24 s 3.02 s 3.26 s
Reply delay, all turns P50 5.16 s 3.06 s 3.20 s
Reply delay P95 6.72 s 4.58 s 3.68 s

The decision still takes about 1.24 s and is still measured in full; what changes is that the reply is already written when the ruling arrives, so model to first text drops from about 0.98 s to near zero and the router's roundtrip from 4.34 s to 2.54 s (medians). This is one laptop run of 8 calls each, so it needs a larger run before anyone quotes it.

It is off by default for two reasons. The flow controller asks about the words at every pause, so the router starts many replies that are then dropped: 229 reply requests against 47 for the same 46 answered turns in that run. And false cutoffs went from 2 to 6 (LiveKit: 4), which may be replies landing in a pause the caller had not finished; that needs a look before this is turned on anywhere.

Changes

  • New ROUTER_SPECULATIVE_REPLIES setting (config agent.speculative_replies), passed through the session manager to every agent. The router logs once at startup when it is on.
  • internal/agent/speculate.go: when a settled turn goes to the flow controller and the agent is quiet, the reply is asked for at the same time and its stream is held where a guardrail holds one, so nothing is spoken, recorded or acted on. An answer for the same words, with no clarifying note and an unchanged history, releases it; any other ruling, a supersede, an interruption or a pipeline release drops it.
  • Speculation is skipped for text conversations, when a guardrail is configured, while the agent is speaking or has tools pending, and for another voice on the caller's track.
  • Harness.Forget clears the provider's stored-response state when a held reply is dropped. Respond records its input as sent, so without this the next turn with the same words would send nothing new and continue from a response that never saw them.
  • Turn timing: a released reply counts its model as starting at the ruling, so decision_ms stays honest and the head start shows up in model_to_first_text_ms.
  • Agent tests for a held reply being spoken only after an answer, dropped on ignore, restarted on clarify, and waiting for the ruling when the setting is off; a config test for the setting.

Every answered turn paid for the flow decision and the reply one after
the other, because the reply model was only asked once the ruling came
back. Matched Voicebench runs put that decision at a median of 1.2 s per
turn, more than our whole gap to LiveKit.

With ROUTER_SPECULATIVE_REPLIES=true the agent asks for the reply beside
the ruling and holds its stream where a guardrail already holds one:
nothing is spoken, recorded or acted on until an answer for the same
words releases it. Any other ruling, a clarifying note, a conversation
that moved on or a model switch drops it, and the harness forgets the
provider's stored state so the next turn cannot continue from a reply
nobody used.

Turn timing stays honest: the decision is still measured up to the
ruling, and the head start shows as a shorter wait for the model. Off
by default, since a dropped reply is still paid for.
@darkoatanasovski
darkoatanasovski marked this pull request as ready for review September 30, 2026 09:03
@darkoatanasovski
darkoatanasovski merged commit 9736941 into accelerate Sep 30, 2026
10 of 12 checks passed
@darkoatanasovski
darkoatanasovski deleted the accelerate-improvements branch September 30, 2026 12:29
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant