Skip to content

[AI-811] Benchmark average time to reply against LiveKit - #718

Merged
darkoatanasovski merged 4 commits into
acceleratefrom
AI-811
Oct 2, 2026
Merged

darkoatanasovski merged 4 commits into
acceleratefrom
AI-811

Conversation

@darkoatanasovski

Copy link
Copy Markdown

Why

Average time to reply is the second of the three benchmarks in AI-809. Voicebench already measured every reply gap from the recordings, but it had no agreed headline, no mean, no way to tell a real difference from noise in voicebench compare, and no record of where a turn's time went inside our own stack. The router's per-turn stage timing (#674) shows that the waiting is in specific stages, the flow controller's decision above all, so those stages belong next to the number they explain.

Ticket: AI-811

Decisions

The ticket left two choices open; this takes its proposals, and both are easy to change in review.

  • Headline: reply time is P50 and P95 over non-tool turns. The mean is shown beside them, not instead of them, because a few slow turns move it a long way.
  • Tool turns: reported on their own and kept in the all-turns figures. A turn that waited on a tool is slower for a reason the conversation loop does not own.

Changes

  • summary.json gains non_tool_p95_ms, non_tool_mean_ms, tool_p50_ms, tool_samples and v2v_mean_ms, and report.md has a Reply time table with sample counts.
  • voicebench compare shows each reply-time statistic with its 95% bootstrap interval and marks a run whose interval lies wholly above the best one, so a gap inside the noise is not read as a win. Against a baseline it prints the non-tool P50 difference with the interval of that difference and the smallest detectable difference, half the interval's width. The resampling is seeded, so the same summaries print the same intervals.
  • For targets on the router, each call keeps the router's timeline as timeline.json, and metrics.json lists every caller turn's stages: STT settle, cadence, decision, model to first text, text to TTS, TTS to audio and roundtrip. report.md shows their medians pooled over every timed turn. Turns the agent started itself, like the greeting or a tool reply, are left out.
  • For the python target, the averages the agent session reports about itself are read just before the session closes, since they go with it, and kept as agent_metrics.json with their median across calls in report.md. The table is left out when nothing in it was measured, which is the case on the accelerated target because the router runs the pipeline.
  • benchmark/README.md defines reply time, the intervals and the per-stage diagnostics.

The stage tables are diagnostics for our own performance work, not figures to set against LiveKit, which reports nothing comparable. A live restaurant call on the accelerated target filled them as expected, and also showed two turns whose decision took 11.8 s and 13.7 s, which is the kind of thing they are there to surface.

Still open

The run-to-run spread and a stored baseline need repeated runs on the pinned US runner, the same as for AI-810, so the ticket's spread criterion stays open until then. The intervals in compare already show, per comparison, how small a difference the samples can detect.

Reply time's headline is now non-tool turns, P50 and P95, with the mean
beside them and tool turns reported apart: a turn that waited on a tool
is slower for a reason the loop does not own. summary.json and report.md
gain non_tool_p95_ms, non_tool_mean_ms, tool_p50_ms, tool_samples and
v2v_mean_ms and a Reply time table.

voicebench compare shows each reply-time statistic with its 95%
bootstrap interval and marks a run whose interval clears the best one.
Against a baseline it prints the non-tool P50 delta with the interval of
the difference and the smallest difference the samples can detect. The
resampling is seeded, so the same summaries print the same intervals.
For targets on the router, Voicebench now reads the call's timeline after
it ends, keeps it as timeline.json, and records every caller turn's
stages (STT settle, cadence, decision, model to first text, text to TTS,
TTS to audio, roundtrip) in metrics.json. report.md shows their medians,
pooled over every timed turn.

For the python target, the session's own averages are read just before
the session is closed, since they go with it, and kept as
agent_metrics.json, with their median across calls in report.md.

Both are diagnostics for our own performance work rather than figures to
set against LiveKit, which reports neither.
Headline P50 and P95 over non-tool turns, the mean beside them, tool turns apart; the intervals and smallest detectable difference in compare; and the per-stage diagnostics each call keeps.
On the accelerated target the router runs the pipeline, so the Python agent reports its stage averages as null and the table was all dashes.
@darkoatanasovski
darkoatanasovski merged commit d89b5c0 into accelerate Oct 2, 2026
10 of 12 checks passed
@darkoatanasovski
darkoatanasovski deleted the AI-811 branch October 2, 2026 09:14
kanat pushed a commit that referenced this pull request Oct 5, 2026
* AI-811: report reply time on non-tool turns, with intervals in compare

Reply time's headline is now non-tool turns, P50 and P95, with the mean
beside them and tool turns reported apart: a turn that waited on a tool
is slower for a reason the loop does not own. summary.json and report.md
gain non_tool_p95_ms, non_tool_mean_ms, tool_p50_ms, tool_samples and
v2v_mean_ms and a Reply time table.

voicebench compare shows each reply-time statistic with its 95%
bootstrap interval and marks a run whose interval clears the best one.
Against a baseline it prints the non-tool P50 delta with the interval of
the difference and the smallest difference the samples can detect. The
resampling is seeded, so the same summaries print the same intervals.

* AI-811: keep each call's per-stage timing

For targets on the router, Voicebench now reads the call's timeline after
it ends, keeps it as timeline.json, and records every caller turn's
stages (STT settle, cadence, decision, model to first text, text to TTS,
TTS to audio, roundtrip) in metrics.json. report.md shows their medians,
pooled over every timed turn.

For the python target, the session's own averages are read just before
the session is closed, since they go with it, and kept as
agent_metrics.json, with their median across calls in report.md.

Both are diagnostics for our own performance work rather than figures to
set against LiveKit, which reports neither.

* AI-811: define reply time in the Voicebench README

Headline P50 and P95 over non-tool turns, the mean beside them, tool turns apart; the intervals and smallest detectable difference in compare; and the per-stage diagnostics each call keeps.

* AI-811: leave out the agent table when nothing in it was measured

On the accelerated target the router runs the pipeline, so the Python agent reports its stage averages as null and the table was all dashes.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant