Measure and trim what a Live interview spends - #203
Merged
Merged
Conversation
An interview spends Gemini credit faster than users expect, and the server could not say where: it logged one usage sum, could lose the last turn's usage on shutdown, and treated a depleted prepaid account like any outage, spending its restart budget on it. Usage is now recorded as each frame is decoded and logged per turn, with its session and the input that asked for the generation, and per session with how it ended, billing included; scripts/analyze-gemini-usage.py summarizes those lines offline. A billing failure is retried only on another key. The Live instructions, greeting and tool answers drop repeated explanation, and every generation is billed again on what stays in its context. Measured on gemini-3.1-flash-live-preview with the production opening, the setup and then the greeting, the first generation cost 5,292 prompt tokens against 5,912 on main in every run: 620 fewer (10.5%), 534 from the setup and 86 from the greeting, both billed again on every later generation. Thought tokens stayed at zero under the minimal level the family-keyed thinking setting now picks. Candidate video, when enabled, sends a fifth of the frames at the low media resolution. An optional compression pair adds silent local-state checkpoints for measured experiments and leaves the provider's defaults in place until a comparison says otherwise.
The web command read any agent config error as a missing Gemini key and served the web side alone, so a single host with a mistyped optional entry, such as an inverted compression pair, left every interview waiting for an interviewer that would never join. Only missing keys now mean the web half of a split deployment; an invalid entry stops the server with the error.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
An interview spends Gemini credit faster than users expect, and the server could not say where: it logged one usage sum per interview, could lose the last turn's usage when a socket was replaced or the room shut down, and treated a project out of prepaid credit like any other outage. This records usage as each frame is decoded and logs it per turn with the session and the input that asked for the generation, logs a per-session summary with how it ended (
billingincluded), and addsscripts/analyze-gemini-usage.pyto summarize those lines offline. A 402 or a depleted-credit close is retried only on another key, andcodetrial webrefuses to start on an invalid agent config instead of silently serving without its interviewer.It also trims what every Live generation is billed on. The Live API bills each generation on its whole context, so text that stays in the context (the system instructions, tool declarations, the greeting, stage directions and tool answers) is paid for again on every later turn. The Live instructions drop repeated explanation (live prompt 16, bundle 24); the greeting no longer repeats the exercise's title and brief, which the instructions already carry; test-run reactions, the earlier-steps reminder and the
end_interviewdescription state their rule once. Candidate video, when enabled, sends a fifth of the frames at the low media resolution. An optionalGEMINI_CONTEXT_TRIGGER_TOKENS/GEMINI_CONTEXT_TARGET_TOKENSpair enables silent local-state checkpoints for compression experiments; defaults stay with the provider.Measured token reduction
Production opening (setup, then the greeting with the timer sentence) sent to
gemini-3.1-flash-live-preview, arms alternating, first-generationpromptTokenCountas reported by the API:mainThat is 620 fewer prompt tokens (10.5%) on the opening generation. All of it is text that stays in the context, so every later generation of the interview carries the same saving.
Where it comes from, counted per text with
countTokensongemini-3.1-flash-lite; the parts sum to the live difference exactly (563 + 86 - 29 = 620):mainThe tool declarations grow because the
read_editordescription now asks for only the code the current question needs; without it the credentialed editor probe re-read an editor it had already been shown in every run. The second generation after one candidate turn varies by about 150 tokens between runs (it includes the model's own spoken reply and a varying number of tool calls), so these numbers use the first generation, where the count is deterministic. Thought tokens were zero on every side.Not measured: the optional compression experiment. The credentialed replay in
tests/unit/livekit/cost.rsdid not survive a full trial on this model (a tool answer the model never followed up, then a 45-second silent socket), so this PR makes no compression saving claim.Quality checks
turnCompleteframe, so per-turn sums do not double count).read_editorfirst); removing the unsolicited-hint clause from the editor review (kept deliberately when the watch prompts were last cut, because a review is where an unsolicited hint happens).Verification
./scripts/test.shpasses; the lanes skipped locally are Playwright Chromium, cargo-audit and shellcheck.codetrial check-geminiwith the default model accepts the new setup.cargo mutants --in-difflocally on the functions CI flagged: survivors are covered by new tests, and three wrappers that only send or log on the live socket join.cargo/mutants.tomlwith the reason.Not included: a pre-session readiness check and browser messaging for quota and billing failures, a pre-session cost warning, a session length cap, a text-only mode, and any change to the default compression window.