Skip to content
2 changes: 2 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -9,6 +9,8 @@
- Roo Code support. Roo Code was discontinued in May 2026 (its repository is archived).

### Fixed
- **Codex long-context calls now use the published higher rates at their prompt-token threshold.** The bundled and live pricing readers retain LiteLLM's context tiers, and Codex applies them per call; below-threshold calls and other providers retain base pricing, and explicit price overrides still win. Audit totals use the same per-call calculation. Pricing and daily-cache versions advance so warm history re-derives with the corrected rates. The regenerated bundle preserves the snapshot refresh's base-rate guards and exact-key carry-forward for previously priced models. (#1076)
- Bundled pricing is refreshed: 5,896 primary entries (298 added, 158 repriced by upstream, 364 no-longer-published ids moved to the fallback table, which grows from 212 to 543). `grok-4.6` base input moves to $1.25/M; `claude-opus-4` is pinned at Anthropic list price since LiteLLM dropped it.
- **OMP: side calls outside the chat transcript now count.** OMP logs `find`, `judge`, TTSR checks and cache warming as `model_usage` entries, not assistant messages, and the parser read only messages, so that spend was missing. On one real week that was 605 `find` decisions on TypeSafe Jev and a $0.048 cache warm. Each side call now counts at its reported cost, with its purpose in the tools column as `omp:find`, `omp:cache-warm` and so on, and a zero reported cost is priced from tokens. A side call never duplicates an assistant turn. OMP sessions re-parse once and settled days re-derive (daily cache v41).
- **Text selection works again in the terminal dashboard.** Mouse tracking, added in 0.9.20 so the wheel scrolls the view, makes the terminal send mouse events to the application, and in most terminals that stops click and drag from selecting text, so copying a number out of the dashboard needed Shift held down. Tracking is now off by default and `m` turns it on and off while the dashboard runs.
- **Cursor Enterprise and team seats connect in Plan & Usage instead of showing an unrecognized quota response.** Their `usage-summary` carries no individual meter, only a team on-demand meter, which now fills the bar under the Enterprise plan label. On other plans it is listed beside the monthly window. In the macOS app a rate limit, outage or unrecognized response no longer adds sign-in guidance, which is only shown when the session is missing or rejected. Reported in #1546, thanks @paulodearaujo.
Expand Down
9 changes: 9 additions & 0 deletions app/renderer/lib/types.ts
Original file line number Diff line number Diff line change
Expand Up @@ -969,6 +969,15 @@ export type ModelCosts = {
cacheReadCostPerToken: number
webSearchCostPerRequest: number
fastMultiplier: number
/** The vendor's long-context tier, applied once prompt tokens reach the
* threshold. Optional: absent on models without a published tier. */
longContextTier?: {
thresholdTokens: number
inputCostPerToken: number
outputCostPerToken: number
cacheWriteCostPerToken?: number
cacheReadCostPerToken?: number
}
}

/** One (provider, model) audit bucket (src/audit-report.ts AuditRow): raw
Expand Down
65 changes: 60 additions & 5 deletions scripts/bundle-litellm.mjs
Original file line number Diff line number Diff line change
Expand Up @@ -54,6 +54,11 @@ const MANUAL_ENTRIES = {
// exact gpt-5.6 tuple (Sol-tier: $5/$30 per million, 1.25x cache-write).
'gpt-5.6-codex': [5e-6, 3e-5, 6.25e-6, 5e-7],
'gpt-5.6-codex-max': [5e-6, 3e-5, 6.25e-6, 5e-7],
// LiteLLM dropped `claude-opus-4` upstream (a refresh moves dropped ids to
// the fallback tier), but the Cursor-style alias `claude-4-opus` resolves
// against PRIMARY rows - without this pin the bare id falls to the
// snowflake gateway row and under-prices by 3x. Anthropic list price.
'claude-opus-4': [15e-6, 75e-6, 18.75e-6, 1.5e-6],
}

const snapshot = {}
Expand All @@ -64,11 +69,43 @@ if (!res.ok) throw new Error(`HTTP ${res.status}`)
const data = await res.json()
const entries = Object.entries(data).filter(([k]) => k !== 'sample_spec')

// The plain context-length tiers only: `input_cost_per_token_above_272k_tokens`
// and siblings. Service-tier variants (`_above_272k_priority_tokens`,
// `_above_272k_flex_tokens`) and the 1-hour cache-write combination are NOT
// context thresholds and are deliberately not matched. The threshold comes
// from the key suffix (272k -> 272000) because LiteLLM carries no numeric
// threshold field (#1076). Mirrored in src/models.ts parseLiteLLMEntry.
const TIER_KEY_RE = /^(input_cost_per_token|output_cost_per_token|cache_read_input_token_cost|cache_creation_input_token_cost)_above_(\d+)k_tokens$/

function tierOf(entry) {
// Rates are read ONLY from the largest threshold a model carries, so a
// hypothetical entry with two tiers can never mix a smaller tier's rates
// under the bigger threshold. Values must be finite and non-negative, the
// same validation src/models.ts applies on the live path.
const byThreshold = new Map()
for (const [key, value] of Object.entries(entry)) {
const m = TIER_KEY_RE.exec(key)
if (!m || typeof value !== 'number' || !Number.isFinite(value) || value < 0) continue
const tokens = Number(m[2]) * 1000
const rates = byThreshold.get(tokens) ?? {}
if (m[1] === 'input_cost_per_token') rates.input = value
else if (m[1] === 'output_cost_per_token') rates.output = value
else if (m[1] === 'cache_read_input_token_cost') rates.cacheRead = value
else rates.cacheWrite = value
byThreshold.set(tokens, rates)
}
if (byThreshold.size === 0) return null
const threshold = Math.max(...byThreshold.keys())
const rates = byThreshold.get(threshold)
if (rates.input == null || rates.output == null) return null
return { threshold, input: rates.input, output: rates.output, cacheWrite: rates.cacheWrite ?? null, cacheRead: rates.cacheRead ?? null }
}

function toVal(entry) {
const inp = entry.input_cost_per_token
const out = entry.output_cost_per_token
if (inp == null || out == null) return null
return [inp, out, entry.cache_creation_input_token_cost ?? null, entry.cache_read_input_token_cost ?? null, entry.provider_specific_entry?.fast ?? null]
return [inp, out, entry.cache_creation_input_token_cost ?? null, entry.cache_read_input_token_cost ?? null, entry.provider_specific_entry?.fast ?? null, tierOf(entry)]
}

// Pass 1: direct entries (no prefix) get priority
Expand All @@ -80,8 +117,10 @@ for (const [name, entry] of entries) {
// A tuple's completeness: how many optional rate slots (cache-write,
// cache-read) carry a published value. A richer upstream row may cite it to
// FILL a sparser entry's missing slots - never as a license to re-price it
// (see the fillsOnly guard in Pass 2).
const completeness = (val) => (val[2] != null ? 1 : 0) + (val[3] != null ? 1 : 0) + (val[5] != null ? 1 : 0)
// (see the fillsOnly guard in Pass 2). The tier slot (5) is deliberately not
// counted: tier presence must never decide which row wins, or a tier-bearing
// row would outrank the base-richer row main would have picked.
const completeness = (val) => (val[2] != null ? 1 : 0) + (val[3] != null ? 1 : 0)

// Pass 2: prefixed entries - store full key + stripped (slot-fill-only)
for (const [name, entry] of entries) {
Expand All @@ -100,13 +139,18 @@ for (const [name, entry] of entries) {
// changes across a refresh; only missing slots fill. The completeness-wins
// version re-priced 43 input/output and 34 cache rates by swapping in a
// different upstream row (grok-3 3/15 -> 1.25/2.5, mistral-large-latest
// 8/24 -> 0.5/1.5).
// 8/24 -> 0.5/1.5). Slot 5 (the tier object) stays out of the guard: it is
// built fresh per row, so a reference compare is always false and would
// veto fills main performs (it silently dropped the azure cache-read fill
// for gpt-5.4-pro-class rows); and since the replacement only fires when
// the candidate fills a missing BASE slot, the winning row's tier travels
// with its own base rates - splicing the old row's tier onto the new row's
// base would mix two different upstream rows.
const fillsOnly = (cand, prev) =>
cand[0] === prev[0]
&& cand[1] === prev[1]
&& (prev[2] == null || cand[2] === prev[2])
&& (prev[3] == null || cand[3] === prev[3])
&& (prev[5] == null || cand[5] === prev[5])
if (!existing || (completeness(val) > completeness(existing) && fillsOnly(val, existing))) snapshot[stripped] = val
}

Expand Down Expand Up @@ -236,6 +280,17 @@ const coveredByKey = (key) =>
for (const [k, v] of [...Object.entries(previousSnapshot), ...Object.entries(previousFallback)]) {
if (coveredByKey(k)) continue
if (fallback[k] !== undefined) continue
// Same hygiene the gap-fill passes enforce: a carried row must not be
// @pin or date-suffixed (a query can never arrive in those forms - the
// runtime only ever peels them off, never adds them; a vendor-prefixed
// key CAN arrive verbatim, so it stays carriable) and must not be free on
// both ends (an unpriced model falls to expected-free handling, not a $0
// fallback row). Primary files legitimately hold such rows, so the guard
// lives here: when upstream drops one, carrying it verbatim would
// re-import exactly what tests/pricing-fallback-data.test.ts keeps out
// of this file.
if (/@/.test(k) || /-\d{8}$/.test(k)) continue
if (!validRates(v[0], v[1])) continue
fallback[k] = v
carried += 1
}
Expand Down
45 changes: 35 additions & 10 deletions src/audit-report.ts
Original file line number Diff line number Diff line change
@@ -1,5 +1,5 @@
import { behavioralCallWeight } from './behavioral-weight.js'
import { billableOutputTokens, cacheWriteCostPerToken, fallbackRawModelDisplayName, getModelCosts, getShortModelName, sanitizeModelForDisplay, type ModelCosts } from './models.js'
import { billableOutputTokens, cacheWriteCostPerToken, fallbackRawModelDisplayName, getModelCosts, getShortModelName, sanitizeModelForDisplay, tieredCostsFor, type ModelCosts } from './models.js'
import { getProvider } from './providers/index.js'
import { formatCost, formatTokens } from './format.js'
import { renderTable, type TableColumn } from './text-table.js'
Expand Down Expand Up @@ -57,6 +57,12 @@ export async function aggregateAudit(projects: ProjectSummary[]): Promise<AuditR
attributedCostUSD: number
cacheReadDisplayed: number
raw: AuditRow['raw']
/** Per-call recomputed components, each priced with THAT call's tier swap —
* a (provider, model) bucket's summed tokens must never cross the
* per-call 272k threshold on their own (review: a $4,979 phantom gap). */
recomputed: { input: number, output: number, cacheWrite: number, cacheRead: number, webSearch: number }
/** Resolved rates for the bucket's model, fetched once per bucket. */
rates: ModelCosts | null
}
const buckets = new Map<string, Bucket>()

Expand All @@ -75,6 +81,8 @@ export async function aggregateAudit(projects: ProjectSummary[]): Promise<AuditR
calls: 0,
attributedCostUSD: 0,
cacheReadDisplayed: 0,
recomputed: { input: 0, output: 0, cacheWrite: 0, cacheRead: 0, webSearch: 0 },
rates: null,
raw: {
inputTokens: 0,
outputTokens: 0,
Expand All @@ -85,6 +93,7 @@ export async function aggregateAudit(projects: ProjectSummary[]): Promise<AuditR
webSearchRequests: 0,
},
}
bucket.rates = getModelCosts(bucket.model)
buckets.set(key, bucket)
}
const u = call.usage
Expand All @@ -97,8 +106,25 @@ export async function aggregateAudit(projects: ProjectSummary[]): Promise<AuditR
bucket.raw.webSearchRequests += u.webSearchRequests
// Per-call max (then summed) mirrors how the reports collapse the two
// cache-read vocabularies, so the audit's displayed total matches.
bucket.cacheReadDisplayed += Math.max(u.cacheReadInputTokens, u.cachedInputTokens)
const cacheReadForCall = Math.max(u.cacheReadInputTokens, u.cachedInputTokens)
bucket.cacheReadDisplayed += cacheReadForCall
bucket.attributedCostUSD += call.costUSD
// Recompute per call through the same tier swap calculateCost applies
// (prompt tokens = input + cached input of THIS call), so a
// long-context request shows the rates that priced it while a bucket
// of small calls never crosses the threshold on the sum. Fast-mode
// and the 1-hour cache-write rate remain visible gaps on purpose.
if (bucket.rates) {
const promptTokens = u.inputTokens + cacheReadForCall
const tiered = tieredCostsFor(bucket.model, bucket.rates, promptTokens, bucket.provider)
const outputForCall = billableOutputTokens(bucket.provider, u.outputTokens, u.reasoningTokens)
bucket.recomputed.input += u.inputTokens * tiered.inputCostPerToken
bucket.recomputed.output += outputForCall * tiered.outputCostPerToken
bucket.recomputed.cacheWrite += u.cacheCreationInputTokens * cacheWriteCostPerToken(bucket.model, tiered)
bucket.recomputed.cacheRead += cacheReadForCall * tiered.cacheReadCostPerToken
// Web search never participates in a tier; keep it on the base row.
bucket.recomputed.webSearch += u.webSearchRequests * bucket.rates.webSearchCostPerRequest
}
// Supplementary accounting calls keep their tokens and cost above but are not
// distinct requests, so they add no call weight (see behavioral-weight.ts).
bucket.calls += behavioralCallWeight(call)
Expand Down Expand Up @@ -134,16 +160,15 @@ export async function aggregateAudit(projects: ProjectSummary[]): Promise<AuditR
cacheWriteTokens: bucket.raw.cacheCreationInputTokens,
cacheReadTokens: bucket.cacheReadDisplayed,
}
const rates = getModelCosts(bucket.model)
const cost = {
input: rates ? displayed.inputTokens * rates.inputCostPerToken : 0,
output: rates ? displayed.outputTokens * rates.outputCostPerToken : 0,
cacheWrite: rates ? displayed.cacheWriteTokens * cacheWriteCostPerToken(bucket.model, rates) : 0,
cacheRead: rates ? displayed.cacheReadTokens * rates.cacheReadCostPerToken : 0,
webSearch: rates ? bucket.raw.webSearchRequests * rates.webSearchCostPerRequest : 0,
recomputedTotalUSD: 0,
input: bucket.recomputed.input,
output: bucket.recomputed.output,
cacheWrite: bucket.recomputed.cacheWrite,
cacheRead: bucket.recomputed.cacheRead,
webSearch: bucket.recomputed.webSearch,
recomputedTotalUSD: bucket.recomputed.input + bucket.recomputed.output + bucket.recomputed.cacheWrite + bucket.recomputed.cacheRead + bucket.recomputed.webSearch,
}
cost.recomputedTotalUSD = cost.input + cost.output + cost.cacheWrite + cost.cacheRead + cost.webSearch
const rates = bucket.rates
rows.push({
provider: bucket.provider,
providerDisplayName: meta.displayName,
Expand Down
4 changes: 3 additions & 1 deletion src/daily-cache.ts
Original file line number Diff line number Diff line change
Expand Up @@ -226,7 +226,9 @@ import type { DateRange, ProjectSummary } from './types.js'
// v41: OMP side calls (`model_usage` entries: find, judge, cache warming) now
// count. Days finalized at v40 miss them; the bump re-derives surviving days.
// Call counts and cost only rise.
export const DAILY_CACHE_VERSION = 41
// v42: #1076 long-context tiers. A Codex call past a model's published
// long-context threshold prices at the tier rate, so settled days re-derive.
export const DAILY_CACHE_VERSION = 42
const MIN_SUPPORTED_VERSION = 28

/// Providers whose per-day CALL COUNT means something different at
Expand Down
2 changes: 1 addition & 1 deletion src/data/litellm-snapshot.json

Large diffs are not rendered by default.

Loading
Loading