Speak sooner and stop answering a caller twice - #688
Merged
Merged
Conversation
…ntence The voice was handed whole sentences, so the first thing a caller heard waited for the model to reach the first full stop. The first chunk of a reply now ends at the first clause once it has 20 characters, and only where the mark follows a letter, so 1,250 and 7:30 stay whole. The rest of the reply still goes by the sentence.
Every revision waited the 350 ms cadence gap before it was asked about, including a final transcript, whose provider has already decided the caller stopped. A final now waits 60 ms (plus any grace owed after an overlap), and a final of words already waited on cuts the pending wait short rather than leaving it. A final that ends in digits still waits the retry gap, because a PIN or a time may still be growing.
…nswered A caller who paused mid-sentence while a floor question was in flight had the first part of their turn queued. They kept talking, the whole sentence was answered, and when the agent finally stopped the queued part was answered again: in the Voicebench restaurant run that second answer went out 17.9 s later, to less than the caller had asked. Answering words that restate or grow the queued ones now drops the queued turn. Words that are not a revision of it leave it queued, since they may be the only copy of what it said.
darkoatanasovski
marked this pull request as ready for review
October 2, 2026 06:23
Agent.simple_response was renamed to responses.create on accelerate; the two Gemini examples that came in with the merge from main still used the old name, which failed mypy for every PR against accelerate.
The types match the spec again, but sdks/js itself has not been moved to the harness API yet, so it no longer compiles against them. That update belongs in its own change; this PR stays router-only.
kanat
pushed a commit
that referenced
this pull request
Oct 5, 2026
* acceleration: speak a reply's first clause without waiting for the sentence The voice was handed whole sentences, so the first thing a caller heard waited for the model to reach the first full stop. The first chunk of a reply now ends at the first clause once it has 20 characters, and only where the mark follows a letter, so 1,250 and 7:30 stay whole. The rest of the reply still goes by the sentence. * acceleration: put finalized words to the controller without the full gap Every revision waited the 350 ms cadence gap before it was asked about, including a final transcript, whose provider has already decided the caller stopped. A final now waits 60 ms (plus any grace owed after an overlap), and a final of words already waited on cuts the pending wait short rather than leaving it. A final that ends in digits still waits the retry gap, because a PIN or a time may still be growing. * acceleration: drop a queued turn once the caller's fuller words are answered A caller who paused mid-sentence while a floor question was in flight had the first part of their turn queued. They kept talking, the whole sentence was answered, and when the agent finally stopped the queued part was answered again: in the Voicebench restaurant run that second answer went out 17.9 s later, to less than the caller had asked. Answering words that restate or grow the queued ones now drops the queued turn. Words that are not a revision of it leave it queued, since they may be the only copy of what it said. * acceleration: note the router speed changes in the changelog * plugins/gemini: call agent.responses.create in the examples Agent.simple_response was renamed to responses.create on accelerate; the two Gemini examples that came in with the merge from main still used the old name, which failed mypy for every PR against accelerate. * sdks/js: regenerate the types from the spec The spec on accelerate moved a session's sandbox, skills, subagent and tasks into harness (7f7bd9b, 41a832e) without regenerating sdks/js/src/generated/api.ts, so the js job's npm run types --check failed for every PR against accelerate. * Revert the regenerated JS types The types match the spec again, but sdks/js itself has not been moved to the harness API yet, so it no longer compiles against them. That update belongs in its own change; this PR stays router-only.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Why
After speculative replies (#682), our first reply matched LiveKit Inference on the median in a Voicebench run where both sides use the same models and voice, but our slow tail was still behind and the router's per-turn stage timing (#674) pointed at three more places where time is spent waiting rather than working:
Measured on the same matched restaurant run as #682 (8 calls, GPT-5.6 Luna, Inworld TTS-2 Flash, voice Ashley,
ROUTER_SPECULATIVE_REPLIES=true, laptop):The P95 improvement over speculation alone is outside the noise (−1.72 s to −0.22 s); the medians move in the right direction but are still inside it at 8 calls. One
coherencetrial in that run had 15 false cutoffs; three more trials each with and without the first-clause change (0/0/0 against 2/0/2) put that down to a single early "respond" ruling mid-sentence rather than to this branch. This needs a larger run before anyone quotes it.Two items from the same investigation are not here. A fast path that skips the flow controller on clear-cut turns would skip exactly the ignore/clarify judgements the quality benchmark is meant to protect. And the flow controller's round trip is still about 1.2 s because none of its fast models are reachable: Cerebras has archived
gemma-4-31b(its other models return 402 on our key), the Baseten Gemma deployment is deactivated, and Jev is registered as anlcmclassifier thatllm-flowcannot pick. Fixing that is an account and config question first.Changes
,;:and the CJK marks) once it has 20 characters and the mark follows a letter, so1,250and7:30stay whole. Later chunks still go by the sentence, and each new reply may open on a clause again.