Repository navigation
[AI-811] Benchmark average time to reply against LiveKit - #718
Merged
Merged
Conversation
Reply time's headline is now non-tool turns, P50 and P95, with the mean beside them and tool turns reported apart: a turn that waited on a tool is slower for a reason the loop does not own. summary.json and report.md gain non_tool_p95_ms, non_tool_mean_ms, tool_p50_ms, tool_samples and v2v_mean_ms and a Reply time table. voicebench compare shows each reply-time statistic with its 95% bootstrap interval and marks a run whose interval clears the best one. Against a baseline it prints the non-tool P50 delta with the interval of the difference and the smallest difference the samples can detect. The resampling is seeded, so the same summaries print the same intervals.
For targets on the router, Voicebench now reads the call's timeline after it ends, keeps it as timeline.json, and records every caller turn's stages (STT settle, cadence, decision, model to first text, text to TTS, TTS to audio, roundtrip) in metrics.json. report.md shows their medians, pooled over every timed turn. For the python target, the session's own averages are read just before the session is closed, since they go with it, and kept as agent_metrics.json, with their median across calls in report.md. Both are diagnostics for our own performance work rather than figures to set against LiveKit, which reports neither.
Headline P50 and P95 over non-tool turns, the mean beside them, tool turns apart; the intervals and smallest detectable difference in compare; and the per-stage diagnostics each call keeps.
On the accelerated target the router runs the pipeline, so the Python agent reports its stage averages as null and the table was all dashes.
kanat
pushed a commit
that referenced
this pull request
Oct 5, 2026
* AI-811: report reply time on non-tool turns, with intervals in compare Reply time's headline is now non-tool turns, P50 and P95, with the mean beside them and tool turns reported apart: a turn that waited on a tool is slower for a reason the loop does not own. summary.json and report.md gain non_tool_p95_ms, non_tool_mean_ms, tool_p50_ms, tool_samples and v2v_mean_ms and a Reply time table. voicebench compare shows each reply-time statistic with its 95% bootstrap interval and marks a run whose interval clears the best one. Against a baseline it prints the non-tool P50 delta with the interval of the difference and the smallest difference the samples can detect. The resampling is seeded, so the same summaries print the same intervals. * AI-811: keep each call's per-stage timing For targets on the router, Voicebench now reads the call's timeline after it ends, keeps it as timeline.json, and records every caller turn's stages (STT settle, cadence, decision, model to first text, text to TTS, TTS to audio, roundtrip) in metrics.json. report.md shows their medians, pooled over every timed turn. For the python target, the session's own averages are read just before the session is closed, since they go with it, and kept as agent_metrics.json, with their median across calls in report.md. Both are diagnostics for our own performance work rather than figures to set against LiveKit, which reports neither. * AI-811: define reply time in the Voicebench README Headline P50 and P95 over non-tool turns, the mean beside them, tool turns apart; the intervals and smallest detectable difference in compare; and the per-stage diagnostics each call keeps. * AI-811: leave out the agent table when nothing in it was measured On the accelerated target the router runs the pipeline, so the Python agent reports its stage averages as null and the table was all dashes.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Why
Average time to reply is the second of the three benchmarks in AI-809. Voicebench already measured every reply gap from the recordings, but it had no agreed headline, no mean, no way to tell a real difference from noise in
voicebench compare, and no record of where a turn's time went inside our own stack. The router's per-turn stage timing (#674) shows that the waiting is in specific stages, the flow controller's decision above all, so those stages belong next to the number they explain.Ticket: AI-811
Decisions
The ticket left two choices open; this takes its proposals, and both are easy to change in review.
Changes
summary.jsongainsnon_tool_p95_ms,non_tool_mean_ms,tool_p50_ms,tool_samplesandv2v_mean_ms, andreport.mdhas a Reply time table with sample counts.voicebench compareshows each reply-time statistic with its 95% bootstrap interval and marks a run whose interval lies wholly above the best one, so a gap inside the noise is not read as a win. Against a baseline it prints the non-tool P50 difference with the interval of that difference and the smallest detectable difference, half the interval's width. The resampling is seeded, so the same summaries print the same intervals.timeline.json, andmetrics.jsonlists every caller turn's stages: STT settle, cadence, decision, model to first text, text to TTS, TTS to audio and roundtrip.report.mdshows their medians pooled over every timed turn. Turns the agent started itself, like the greeting or a tool reply, are left out.pythontarget, the averages the agent session reports about itself are read just before the session closes, since they go with it, and kept asagent_metrics.jsonwith their median across calls inreport.md. The table is left out when nothing in it was measured, which is the case on theacceleratedtarget because the router runs the pipeline.benchmark/README.mddefines reply time, the intervals and the per-stage diagnostics.The stage tables are diagnostics for our own performance work, not figures to set against LiveKit, which reports nothing comparable. A live restaurant call on the
acceleratedtarget filled them as expected, and also showed two turns whose decision took 11.8 s and 13.7 s, which is the kind of thing they are there to surface.Still open
The run-to-run spread and a stored baseline need repeated runs on the pinned US runner, the same as for AI-810, so the ticket's spread criterion stays open until then. The intervals in
comparealready show, per comparison, how small a difference the samples can detect.