Repository navigation
Conversation
JJ27
marked this pull request as ready for review
September 29, 2026 23:28
lilly-luo
reviewed
Oct 1, 2026
Carries PR #850's statusline, pricing, and endpoint-rates client onto the Claude mods stack. Conflicts with the mods branch are resolved by keeping both sides; ENABLE_SMART_ROUTING_SAVINGS stays in v2 since the routing flags moved to ucode.constants. Co-authored-by: Isaac <no-reply@databricks.com>
- Price each response against the main model in effect when it ran, so /model no longer reprices the session's history. - Follow the session's /smart-router toggle via UCODE_SESSION_ENV_FILE, drop the nonexistent "/smart-router savings" hint, and show the "off" row when routing setup fails. - Add a Claude Code mods concern (smart-routing-savings.ts) that sums per-request usage into mod-usage.json, plus a --mod-usage pricer entry point that ug passes to it through UCODE_SAVINGS_PRICER; mods usage has no TTL split, so its cache writes bill at the 1-hour rate ug uses. Co-authored-by: Isaac <no-reply@databricks.com>
Wire smart-routing-savings.ts into the composed mod: register.ts composes it, the status band appends its segments while routing is on, and SMART_ROUTING_UI ships the file. Co-authored-by: Isaac <no-reply@databricks.com>
JJ27
force-pushed
the
josh-joseph/smart-routing-savings-statusline
branch
from
October 6, 2026 05:20
e547da8 to
6f8dc9c
Compare
Co-authored-by: Isaac <no-reply@databricks.com>
The gateway reports non-Claude responses by a deployment id that drifts from the model service name (glm-5.3-flash for system.ai.glm-5-3-flash, glm-5-3-colo-on-sp-v1 for system.ai.glm-5-3), so routed subagent usage never matched a price and voided the estimate. The mod now prices the model the request named, and model_key folds the deployment spellings. Co-authored-by: Isaac <no-reply@databricks.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Adds a Claude Code statusline row for smart routing: it shows whether routing is on/off and an estimate of what routing saved this session, priced from the AI Gateway pay-per-token (PPT) endpoint. Default-on (
ENABLE_SMART_ROUTING_SAVINGS=0opts out).The row, by state:
onuntil there's anything to show;offwhen routing is disabled. A cost increase readsSmart routing cost $X more (Y%).model-orchestratorplugin (shown when present, omitted otherwise)./smart-router savings— @lilly's upcoming skill — will give the granular breakdown.Screenshots
Smart routing on, before any prompt has run:
Estimated savings after a routed subagent:
How it works
saved = Σ tokens × price(baseline) − Σ tokens × price(served)— summed per response and per token class (input, output, cache read, 5-minute and 1-hour cache writes). Computed from the Claude Code transcripts (main +<session>/subagents/agent-*.jsonl), sincecost.total_cost_usdexcludes subagents.onrather than undercount.Price source
GET /api/ai-gateway/v2/endpoint-rates:batchGetwith repeateddatabricks_hosted_model_servicesquery params. The client reads the response'sUSDcost group — per-million-tokentoken_costsfor input, output, and cache (5m/1h writes and reads) — and caches it per workspace (dollar rates depend on the org's DBU→$ conversion; models without it stay unpriced rather than shown in DBUs). Served ids are matched back tosystem.ainames (incl. Bedrock-style ids likeanthropic.claude-haiku-4-5-…). Validated end-to-end against the staging endpoint.Scope & notes
Testing
uv run pytest tests/test_smart_routing_pricing.py tests/test_claude_statusline.py tests/test_claude_smart_routing_v2.py tests/test_databricks.py tests/test_agent_claude.pycovers the cost math per token class, the off/on/estimate states, transcript accounting (split records, incremental reads, rewrites), the wrapped command through realsh/bash/zsh, the stdlib-only import guard, the endpoint request/response mapping, and launch wiring (routed + vanilla-off). Full unit suite is green;ruff/tyclean. End-to-end verified live against the staging PPT endpoint.Also in this PR: tracing CI fix
The Claude and Codex tracing journeys (
test_ug_*_exports_trace_to_configured_table) were failing withSQL statement did not finish within 180 secondsbecausetests/integration/utils/sql.pypicked the first warehouse when none was running and counted warehouse startup against the query's 180s budget. Fix: prefer a serverless warehouse, allow up to 10 minutes inPENDING(warehouse startup) before the 180s execution limit, and report which phase timed out. Can move to its own PR if reviewers prefer.This pull request and its description were written by Isaac.