Skip to content

Measure and trim what a Live interview spends - #203

Merged
jserv merged 2 commits into
mainfrom
token-reduction
Oct 1, 2026
Merged

jserv merged 2 commits into
mainfrom
token-reduction

Conversation

@jserv

@jserv jserv commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

An interview spends Gemini credit faster than users expect, and the server could not say where: it logged one usage sum per interview, could lose the last turn's usage when a socket was replaced or the room shut down, and treated a project out of prepaid credit like any other outage. This records usage as each frame is decoded and logs it per turn with the session and the input that asked for the generation, logs a per-session summary with how it ended (billing included), and adds scripts/analyze-gemini-usage.py to summarize those lines offline. A 402 or a depleted-credit close is retried only on another key, and codetrial web refuses to start on an invalid agent config instead of silently serving without its interviewer.

It also trims what every Live generation is billed on. The Live API bills each generation on its whole context, so text that stays in the context (the system instructions, tool declarations, the greeting, stage directions and tool answers) is paid for again on every later turn. The Live instructions drop repeated explanation (live prompt 16, bundle 24); the greeting no longer repeats the exercise's title and brief, which the instructions already carry; test-run reactions, the earlier-steps reminder and the end_interview description state their rule once. Candidate video, when enabled, sends a fifth of the frames at the low media resolution. An optional GEMINI_CONTEXT_TRIGGER_TOKENS/GEMINI_CONTEXT_TARGET_TOKENS pair enables silent local-state checkpoints for compression experiments; defaults stay with the provider.

Measured token reduction

Production opening (setup, then the greeting with the timer sentence) sent to gemini-3.1-flash-live-preview, arms alternating, first-generation promptTokenCount as reported by the API:

Setup and greeting Prompt tokens, first generation Runs
main 5,912 3 of 3 identical
This PR 5,292 8 of 8 identical

That is 620 fewer prompt tokens (10.5%) on the opening generation. All of it is text that stays in the context, so every later generation of the interview carries the same saving.

Where it comes from, counted per text with countTokens on gemini-3.1-flash-lite; the parts sum to the live difference exactly (563 + 86 - 29 = 620):

Text main This PR Saved Billed
System instructions 4,832 4,269 563 every generation
Greeting 256 170 86 every generation after the first
Tool declarations 407 436 -29 every generation
Each test-run reaction (8 variants) 89 to 245 81 to 237 8 per run, then retained
Earlier-steps reminder 75 46 29 per reminder, then retained

The tool declarations grow because the read_editor description now asks for only the code the current question needs; without it the credentialed editor probe re-read an editor it had already been shown in every run. The second generation after one candidate turn varies by about 150 tokens between runs (it includes the model's own spoken reply and a varying number of tool calls), so these numbers use the first generation, where the count is deterministic. Thought tokens were zero on every side.

Not measured: the optional compression experiment. The credentialed replay in tests/unit/livekit/cost.rs did not survive a full trial on this model (a tool answer the model never followed up, then a 45-second silent socket), so this PR makes no compression saving claim.

Quality checks

  • Greeting, 11 sessions per arm against the real model: the interviewer introduced the scenario and asked for the language in 11 of 11 on both arms. It read the timer aloud once in 11 with the new greeting and never with the old one, which is within noise at this sample; the instructions still forbid volunteering the time.
  • Credentialed checkpoint probes pass on the final code: language reconstruction, a declined and a complete behavioral round, and the usage probe (usage arrives once per turn, on the turnComplete frame, so per-turn sums do not double count).
  • Considered and rejected: slowing silence nudges or editor reviews (a behavior change without quality evidence); answering a requested hint with "editor unchanged" (contract row 16 says a hint returns the editor so it needs no read_editor first); removing the unsolicited-hint clause from the editor review (kept deliberately when the watch prompts were last cut, because a review is where an unsolicited hint happens).

Verification

  • ./scripts/test.sh passes; the lanes skipped locally are Playwright Chromium, cargo-audit and shellcheck.
  • codetrial check-gemini with the default model accepts the new setup.
  • cargo mutants --in-diff locally on the functions CI flagged: survivors are covered by new tests, and three wrappers that only send or log on the live socket join .cargo/mutants.toml with the reason.
  • Each of the two commits builds on its own.

Not included: a pre-session readiness check and browser messaging for quota and billing failures, a pre-session cost warning, a session length cap, a text-only mode, and any change to the default compression window.

cubic-dev-ai[bot]

This comment was marked as resolved.

cubic-dev-ai[bot]

This comment was marked as resolved.

jserv added 2 commits October 2, 2026 04:49
An interview spends Gemini credit faster than users expect, and the
server could not say where: it logged one usage sum, could lose the last
turn's usage on shutdown, and treated a depleted prepaid account like
any outage, spending its restart budget on it. Usage is now recorded as
each frame is decoded and logged per turn, with its session and the
input that asked for the generation, and per session with how it ended,
billing included; scripts/analyze-gemini-usage.py summarizes those lines
offline. A billing failure is retried only on another key.

The Live instructions, greeting and tool answers drop repeated
explanation, and every generation is billed again on what stays in its
context. Measured on gemini-3.1-flash-live-preview with the production
opening, the setup and then the greeting, the first generation cost
5,292 prompt tokens against 5,912 on main in every run: 620 fewer
(10.5%), 534 from the setup and 86 from the greeting, both billed again
on every later generation. Thought tokens stayed at zero under the
minimal level the family-keyed thinking setting now picks. Candidate
video, when enabled, sends a fifth of the frames at the low media
resolution. An optional compression pair adds silent local-state
checkpoints for measured experiments and leaves the provider's defaults
in place until a comparison says otherwise.
The web command read any agent config error as a missing Gemini key and
served the web side alone, so a single host with a mistyped optional
entry, such as an inverted compression pair, left every interview
waiting for an interviewer that would never join. Only missing keys now
mean the web half of a split deployment; an invalid entry stops the
server with the error.
@jserv
jserv merged commit 06eac1f into main Oct 1, 2026
18 checks passed
@jserv
jserv deleted the token-reduction branch October 1, 2026 21:34
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant