Skip to content

Speak sooner and stop answering a caller twice - #688

Merged
darkoatanasovski merged 8 commits into
acceleratefrom
accelerate-router-speed
Oct 2, 2026
Merged

darkoatanasovski merged 8 commits into
acceleratefrom
accelerate-router-speed

Conversation

@darkoatanasovski

Copy link
Copy Markdown

Why

After speculative replies (#682), our first reply matched LiveKit Inference on the median in a Voicebench run where both sides use the same models and voice, but our slow tail was still behind and the router's per-turn stage timing (#674) pointed at three more places where time is spent waiting rather than working:

  • The voice was handed whole sentences, so the first thing a caller heard waited for the model to reach the first full stop.
  • Every transcript revision waited the 350 ms cadence gap before it was asked about, including a final one, whose provider has already decided the caller stopped.
  • A turn queued while a floor question was in flight was answered again once the agent stopped, even when the caller had since said the whole thing and it had been answered. In the restaurant run that second answer went out 17.9 s after the first, to less than the caller had asked, and it accounted for the longest "decisions" in the stage timing.

Measured on the same matched restaurant run as #682 (8 calls, GPT-5.6 Luna, Inworld TTS-2 Flash, voice Ashley, ROUTER_SPECULATIVE_REPLIES=true, laptop):

Speculation only This branch LiveKit Inference
First response P50 3.02 s 3.02 s 3.26 s
Reply delay P50 3.06 s 2.72 s 3.20 s
Reply delay P95 4.58 s 3.76 s 3.68 s
Decision stage, median per turn 1.24 s 0.44 s

The P95 improvement over speculation alone is outside the noise (−1.72 s to −0.22 s); the medians move in the right direction but are still inside it at 8 calls. One coherence trial in that run had 15 false cutoffs; three more trials each with and without the first-clause change (0/0/0 against 2/0/2) put that down to a single early "respond" ruling mid-sentence rather than to this branch. This needs a larger run before anyone quotes it.

Two items from the same investigation are not here. A fast path that skips the flow controller on clear-cut turns would skip exactly the ignore/clarify judgements the quality benchmark is meant to protect. And the flow controller's round trip is still about 1.2 s because none of its fast models are reachable: Cerebras has archived gemma-4-31b (its other models return 402 on our key), the Baseten Gemma deployment is deactivated, and Jev is registered as an lcm classifier that llm-flow cannot pick. Fixing that is an account and config question first.

Changes

  • The chunker ends a reply's first chunk at its first clause (, ; : and the CJK marks) once it has 20 characters and the mark follows a letter, so 1,250 and 7:30 stay whole. Later chunks still go by the sentence, and each new reply may open on a clause again.
  • A final transcript is put to the flow controller after 60 ms plus any grace owed after an overlap, and a final of words already waited on cuts the pending wait short. It never lengthens a wait, and a final ending in digits still waits the retry gap.
  • Answering words that restate or grow a queued turn from the same caller drops the queued one. Words that are not a revision of it leave it queued, since with a transcriber that only sends what is new they may be its only copy.
  • Tests for each: the opening clause, lead-ins and numbers that stay whole, a new reply opening on a clause; finals settling early and digit finals waiting; a queued turn dropped by its fuller restatement and kept for different words.

…ntence

The voice was handed whole sentences, so the first thing a caller heard
waited for the model to reach the first full stop. The first chunk of a
reply now ends at the first clause once it has 20 characters, and only
where the mark follows a letter, so 1,250 and 7:30 stay whole. The rest of
the reply still goes by the sentence.
Every revision waited the 350 ms cadence gap before it was asked about,
including a final transcript, whose provider has already decided the
caller stopped. A final now waits 60 ms (plus any grace owed after an
overlap), and a final of words already waited on cuts the pending wait
short rather than leaving it. A final that ends in digits still waits the
retry gap, because a PIN or a time may still be growing.
…nswered

A caller who paused mid-sentence while a floor question was in flight had
the first part of their turn queued. They kept talking, the whole
sentence was answered, and when the agent finally stopped the queued part
was answered again: in the Voicebench restaurant run that second answer
went out 17.9 s later, to less than the caller had asked. Answering words
that restate or grow the queued ones now drops the queued turn. Words
that are not a revision of it leave it queued, since they may be the only
copy of what it said.
@darkoatanasovski
darkoatanasovski marked this pull request as ready for review October 2, 2026 06:23
Agent.simple_response was renamed to responses.create on accelerate; the two Gemini examples that came in with the merge from main still used the old name, which failed mypy for every PR against accelerate.
The spec on accelerate moved a session's sandbox, skills, subagent and tasks into harness (7f7bd9b, 41a832e) without regenerating sdks/js/src/generated/api.ts, so the js job's npm run types --check failed for every PR against accelerate.
The types match the spec again, but sdks/js itself has not been moved to the harness API yet, so it no longer compiles against them. That update belongs in its own change; this PR stays router-only.
@darkoatanasovski
darkoatanasovski merged commit 2009e12 into accelerate Oct 2, 2026
9 of 12 checks passed
@darkoatanasovski
darkoatanasovski deleted the accelerate-router-speed branch October 2, 2026 07:35
kanat pushed a commit that referenced this pull request Oct 5, 2026
* acceleration: speak a reply's first clause without waiting for the sentence

The voice was handed whole sentences, so the first thing a caller heard
waited for the model to reach the first full stop. The first chunk of a
reply now ends at the first clause once it has 20 characters, and only
where the mark follows a letter, so 1,250 and 7:30 stay whole. The rest of
the reply still goes by the sentence.

* acceleration: put finalized words to the controller without the full gap

Every revision waited the 350 ms cadence gap before it was asked about,
including a final transcript, whose provider has already decided the
caller stopped. A final now waits 60 ms (plus any grace owed after an
overlap), and a final of words already waited on cuts the pending wait
short rather than leaving it. A final that ends in digits still waits the
retry gap, because a PIN or a time may still be growing.

* acceleration: drop a queued turn once the caller's fuller words are answered

A caller who paused mid-sentence while a floor question was in flight had
the first part of their turn queued. They kept talking, the whole
sentence was answered, and when the agent finally stopped the queued part
was answered again: in the Voicebench restaurant run that second answer
went out 17.9 s later, to less than the caller had asked. Answering words
that restate or grow the queued ones now drops the queued turn. Words
that are not a revision of it leave it queued, since they may be the only
copy of what it said.

* acceleration: note the router speed changes in the changelog

* plugins/gemini: call agent.responses.create in the examples

Agent.simple_response was renamed to responses.create on accelerate; the two Gemini examples that came in with the merge from main still used the old name, which failed mypy for every PR against accelerate.

* sdks/js: regenerate the types from the spec

The spec on accelerate moved a session's sandbox, skills, subagent and tasks into harness (7f7bd9b, 41a832e) without regenerating sdks/js/src/generated/api.ts, so the js job's npm run types --check failed for every PR against accelerate.

* Revert the regenerated JS types

The types match the spec again, but sdks/js itself has not been moved to the harness API yet, so it no longer compiles against them. That update belongs in its own change; this PR stays router-only.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant