Skip to content

Add Claude Code statusline estimating smart-routing savings (client side) - #850

Open
JJ27 wants to merge 5 commits into
claude-mods-subagent-routingfrom
josh-joseph/smart-routing-savings-statusline
Open

JJ27 wants to merge 5 commits into
claude-mods-subagent-routingfrom
josh-joseph/smart-routing-savings-statusline

Conversation

@JJ27

@JJ27 JJ27 commented Sep 25, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Adds a Claude Code statusline row for smart routing: it shows whether routing is on/off and an estimate of what routing saved this session, priced from the AI Gateway pay-per-token (PPT) endpoint. Default-on (ENABLE_SMART_ROUTING_SAVINGS=0 opts out).

The row, by state:

💰 Est. saved with smart routing: $0.18 (48%) · smart router plugin v0.4.6 · /smart-router savings for details
Smart routing on · smart router plugin v0.4.6
Smart routing off · smart router plugin v0.4.6
  • The 💰 estimate appears once a subagent has run on a cheaper model; on until there's anything to show; off when routing is disabled. A cost increase reads Smart routing cost $X more (Y%).
  • The version is the installed model-orchestrator plugin (shown when present, omitted otherwise). /smart-router savings — @lilly's upcoming skill — will give the granular breakdown.

Screenshots

Smart routing on, before any prompt has run:

SCR-20260929-ogta

Estimated savings after a routed subagent:

image

How it works

  • Estimate: every token the session used (main agent + subagents) is priced at the baseline main model vs the model that actually served it — saved = Σ tokens × price(baseline) − Σ tokens × price(served) — summed per response and per token class (input, output, cache read, 5-minute and 1-hour cache writes). Computed from the Claude Code transcripts (main + <session>/subagents/agent-*.jsonl), since cost.total_cost_usd excludes subagents.
  • Prices are fetched at launch and cached per workspace; the row reads the cache (a statusline refresh can't wait on the network). When prices are missing or any response can't be priced, the row shows on rather than undercount.
  • Wraps any existing statusline (runs it first, prints the row below), is stdlib-only for fast startup, reads transcripts incrementally, and refreshes on main-agent turns.

Price source

GET /api/ai-gateway/v2/endpoint-rates:batchGet with repeated databricks_hosted_model_services query params. The client reads the response's USD cost group — per-million-token token_costs for input, output, and cache (5m/1h writes and reads) — and caches it per workspace (dollar rates depend on the org's DBU→$ conversion; models without it stay unpriced rather than shown in DBUs). Served ids are matched back to system.ai names (incl. Bedrock-style ids like anthropic.claude-haiku-4-5-…). Validated end-to-end against the staging endpoint.

Scope & notes

  • Claude Code only for now; the pricing/rate-fetch layer is agent-agnostic, so a Codex surface is a straightforward follow-up.
  • Cache rates are priced per class (they're ~87% of a session's cost, and fixed multipliers don't hold across models — e.g. Opus 5.5 cache reads bill at 0.05× input, Opus 4.8's at 0.1×).
  • The client sends no LiteSwap header; once the endpoint is live on a workspace the row populates automatically.
  • Refresh fires on main-agent turns (not during a long subagent); the router offering GLM/Kimi is rejected by the test workspace's Anthropic endpoint (pre-existing); a router fallback to a pricier tier is shown honestly as a cost increase.

Testing

uv run pytest tests/test_smart_routing_pricing.py tests/test_claude_statusline.py tests/test_claude_smart_routing_v2.py tests/test_databricks.py tests/test_agent_claude.py covers the cost math per token class, the off/on/estimate states, transcript accounting (split records, incremental reads, rewrites), the wrapped command through real sh/bash/zsh, the stdlib-only import guard, the endpoint request/response mapping, and launch wiring (routed + vanilla-off). Full unit suite is green; ruff/ty clean. End-to-end verified live against the staging PPT endpoint.

Also in this PR: tracing CI fix

The Claude and Codex tracing journeys (test_ug_*_exports_trace_to_configured_table) were failing with SQL statement did not finish within 180 seconds because tests/integration/utils/sql.py picked the first warehouse when none was running and counted warehouse startup against the query's 180s budget. Fix: prefer a serverless warehouse, allow up to 10 minutes in PENDING (warehouse startup) before the 180s execution limit, and report which phase timed out. Can move to its own PR if reviewers prefer.

This pull request and its description were written by Isaac.

@JJ27
JJ27 marked this pull request as ready for review September 29, 2026 23:28
Copilot AI balanced review requested due to automatic review settings September 29, 2026 23:28

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

Copilot AI balanced review requested due to automatic review settings September 29, 2026 23:33

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

Comment thread src/ucode/smart_routing/claude_statusline.py Outdated
JJ27 and others added 3 commits October 5, 2026 23:53
Carries PR #850's statusline, pricing, and endpoint-rates client onto the
Claude mods stack. Conflicts with the mods branch are resolved by keeping
both sides; ENABLE_SMART_ROUTING_SAVINGS stays in v2 since the routing
flags moved to ucode.constants.

Co-authored-by: Isaac <no-reply@databricks.com>
- Price each response against the main model in effect when it ran, so
  /model no longer reprices the session's history.
- Follow the session's /smart-router toggle via UCODE_SESSION_ENV_FILE,
  drop the nonexistent "/smart-router savings" hint, and show the "off"
  row when routing setup fails.
- Add a Claude Code mods concern (smart-routing-savings.ts) that sums
  per-request usage into mod-usage.json, plus a --mod-usage pricer entry
  point that ug passes to it through UCODE_SAVINGS_PRICER; mods usage
  has no TTL split, so its cache writes bill at the 1-hour rate ug uses.

Co-authored-by: Isaac <no-reply@databricks.com>
Wire smart-routing-savings.ts into the composed mod: register.ts
composes it, the status band appends its segments while routing is on,
and SMART_ROUTING_UI ships the file.

Co-authored-by: Isaac <no-reply@databricks.com>
@JJ27
JJ27 force-pushed the josh-joseph/smart-routing-savings-statusline branch from e547da8 to 6f8dc9c Compare October 6, 2026 05:20
Copilot AI balanced review requested due to automatic review settings October 6, 2026 05:20
@JJ27
JJ27 changed the base branch from main to claude-mods-subagent-routing October 6, 2026 05:20

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

JJ27 and others added 2 commits October 6, 2026 06:07
Co-authored-by: Isaac <no-reply@databricks.com>
The gateway reports non-Claude responses by a deployment id that drifts
from the model service name (glm-5.3-flash for system.ai.glm-5-3-flash,
glm-5-3-colo-on-sp-v1 for system.ai.glm-5-3), so routed subagent usage
never matched a price and voided the estimate. The mod now prices the
model the request named, and model_key folds the deployment spellings.

Co-authored-by: Isaac <no-reply@databricks.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants