Start replies before the flow controller rules - #682
Merged
Merged
Conversation
Every answered turn paid for the flow decision and the reply one after the other, because the reply model was only asked once the ruling came back. Matched Voicebench runs put that decision at a median of 1.2 s per turn, more than our whole gap to LiveKit. With ROUTER_SPECULATIVE_REPLIES=true the agent asks for the reply beside the ruling and holds its stream where a guardrail already holds one: nothing is spoken, recorded or acted on until an answer for the same words releases it. Any other ruling, a clarifying note, a conversation that moved on or a model switch drops it, and the harness forgets the provider's stored state so the next turn cannot continue from a reply nobody used. Turn timing stays honest: the decision is still measured up to the ruling, and the head start shows as a shorter wait for the model. Off by default, since a dropped reply is still paid for.
darkoatanasovski
marked this pull request as ready for review
September 30, 2026 09:03
…ovements # Conflicts: # CHANGELOG.md
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Why
Every answered turn paid for the flow decision and the reply one after the other: the reply model was only asked once the flow controller had ruled that the words were meant for the agent. The per-turn stage timing from #674 shows that decision at a median of about 1.2 s per turn, and in a Voicebench run with identical models and voice on both sides (Gemini transcribe-live, GPT-5.6 Luna, Inworld TTS-2 Flash, voice Ashley) LiveKit Inference replied about 1.1 s sooner than us. That gap is smaller than the decision itself, so it's the router's turn-taking, not the providers. The benchmark tooling is in #667.
With speculation on, the same matched restaurant run (8 calls per side) gave:
The decision still takes about 1.24 s and is still measured in full; what changes is that the reply is already written when the ruling arrives, so model to first text drops from about 0.98 s to near zero and the router's roundtrip from 4.34 s to 2.54 s (medians). This is one laptop run of 8 calls each, so it needs a larger run before anyone quotes it.
It is off by default for two reasons. The flow controller asks about the words at every pause, so the router starts many replies that are then dropped: 229 reply requests against 47 for the same 46 answered turns in that run. And false cutoffs went from 2 to 6 (LiveKit: 4), which may be replies landing in a pause the caller had not finished; that needs a look before this is turned on anywhere.
Changes
ROUTER_SPECULATIVE_REPLIESsetting (configagent.speculative_replies), passed through the session manager to every agent. The router logs once at startup when it is on.internal/agent/speculate.go: when a settled turn goes to the flow controller and the agent is quiet, the reply is asked for at the same time and its stream is held where a guardrail holds one, so nothing is spoken, recorded or acted on. An answer for the same words, with no clarifying note and an unchanged history, releases it; any other ruling, a supersede, an interruption or a pipeline release drops it.Harness.Forgetclears the provider's stored-response state when a held reply is dropped.Respondrecords its input as sent, so without this the next turn with the same words would send nothing new and continue from a response that never saw them.decision_msstays honest and the head start shows up inmodel_to_first_text_ms.ignore, restarted onclarify, and waiting for the ruling when the setting is off; a config test for the setting.