Skip to content

fix(usage): record real backend token counts and mark estimated rows - #201

Merged
apapi merged 1 commit into
mainfrom
fix/issue-188-real-usage-tokens
Sep 28, 2026
Merged

apapi merged 1 commit into
mainfrom
fix/issue-188-real-usage-tokens

Conversation

@apapi

@apapi apapi commented Sep 27, 2026

Copy link
Copy Markdown
Collaborator

用量统计记录后端真实 token 数,并标记估算值

Fixes #188

问题

用量统计页面的 token 数全部是估算值,不是后端真实上报的。代理流式路径没有要求后端附带 usage;本地路径忽略了引擎已上报的真实 token;用户无法区分真实与估算。

改动

  1. 引擎真值穿透:Options 新增 OnUsage 回调,llama-server / OpenAI 兼容引擎在响应中拿到的 prompt_tokens / completion_tokens 通过回调传给 handler,handler 用 resolveUsageTokens 逐侧判定——有真值用真值,无真值用估算,混合则打"估算"标记。

  2. 代理流式补 usage:openAIChatRequestToProxyBody 和 anthropicRequestToProxyBody 流式时加 stream_options.include_usage,让后端在流尾返回真实 usage。流中断时仍记录已捕获的部分。

  3. 回退估算改进:CJK 文本按 0.6 tokens/char 估算(原来一律 4 chars/token,中文严重低估);流式无 usage 时从已捕获的 SSE 内容提取文本做估算并打标记;非流式、工具路径同理补全。

  4. 估算标记端到端:DB schema 加 estimated_requests 列(用 PRAGMA 检查是否已存在,不盲跑 ALTER);API 响应、OpenAPI 规范、前端类型同步加字段;用量表格中估算行显示 amber "估算" 标签。

  5. Anthropic 原生流式:移出 if err == nil 守卫,客户端中途断开也记录。

不改的项

  • 三处非代理 OnUsage 对真实引擎是死代码(所有 Engine 都实现了 ChatCompletionProxier),保留作 safety net
  • 64KB tail 限制(只影响无 usage 时的回退估算精度)
  • 内容提取不含 reasoning_content / thinking_delta(低优先级)
  • estimated_requests 不在 key/source/pool 汇总层面暴露(后续迭代)
  • 历史数据不回填

验证

gofmt · go vet · go build ./... · go test ./... · OpenAPI sync · npm run build 全部通过

统计

22 files, +1042 / -206

Issue #188: usage statistics recorded estimated token counts instead of
real backend-reported values. Proxy streaming paths did not request
include_usage; local paths ignored engine-reported tokens; users could
not distinguish real from estimated.

- Add OnUsage callback in Options to thread real prompt/completion
  tokens from llama-server and OpenAI-compatible engines to handlers
- Add stream_options.include_usage to proxy request bodies so backends
  report real usage in stream tails
- Add resolveUsageTokens with per-side sanity validation: use real
  values when available, fall back to estimation otherwise, mark mixed
  rows as estimated
- Improve CJK token estimation (0.6 tokens/char vs flat 4 chars/token)
- Capture partial stream content for fallback estimation when upstream
  reports no usage; record even on mid-stream disconnect
- Add estimated_requests column (PRAGMA-checked migration), expose in
  API/OpenAPI/frontend with amber badge
- Move Anthropic native stream recording outside err==nil guard
- Add tests for proxy stream usage, fallback estimation, and mid-stream
  error recording

22 files, +1042/-206
@apapi

apapi commented Sep 28, 2026

Copy link
Copy Markdown
Collaborator Author

实测结果

用本地模型 Qwen3-0.6B 起了一个实例,发了 3 个请求验证:

测试路径

请求 路径 类型
1 /v1/chat/completions 非流式,proxy 路径
2 /v1/chat/completions 流式,proxy 路径
3 /api/chat 非流式,Ollama 本地路径

llama-server 返回的真实 usage

请求 1 和 2 的响应体里都带了 usage:

{"prompt_tokens": 14, "completion_tokens": 10, "total_tokens": 24}

DB 记录(api_usage_events 表)

model       source  requests  input_tokens  output_tokens  estimated_requests
----------  ------  --------  ------------  -------------  ------------------
Qwen3-0.6B  local   3         139           134            0
  • input_tokens=139、output_tokens=134 是 llama-server 真实上报值,不是 countMessageTokens 估算
  • estimated_requests=0:3 条请求全部记录为真值,没有一条走估算
  • estimated_requests 列已通过迁移创建

修复前对比

修复前 修复后
input_tokens countMessageTokens 估算 llama-server 真实值
output_tokens 0(本地路径未记录) llama-server 真实值
estimated_requests 列不存在 列存在,真值时为 0
流式 usage 未请求 include_usage,拿不到 请求了,流尾返回真实 usage

@apapi

apapi commented Sep 28, 2026

Copy link
Copy Markdown
Collaborator Author

实测结果

用本地模型 Qwen3-0.6B(llama-server 后端)起了一个实例,覆盖 5 条路径验证。

测试环境

  • 模型:Qwen3-0.6B(本地 GGUF,llama-server 后端)
  • 二进制:当前分支 go build ./cmd/csghub-lite/ 编译

测试用例

1. 非流式 /v1/chat/completions(OpenAI proxy 路径)

curl -s http://127.0.0.1:11440/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"model":"Qwen3-0.6B","messages":[{"role":"user","content":"Say hello"}],"stream":false,"max_tokens":5}'

响应中的 usage(llama-server 真实上报):

{"prompt_tokens":10,"completion_tokens":5,"total_tokens":15}

2. 流式 /v1/chat/completions(OpenAI proxy 路径)

curl -s http://127.0.0.1:11440/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"model":"Qwen3-0.6B","messages":[{"role":"user","content":"Say hi"}],"stream":true,"max_tokens":5}'

流尾 usage chunk(stream_options.include_usage 生效):

{"prompt_tokens":10,"completion_tokens":5,"total_tokens":15,"prompt_tokens_details":{"cached_tokens":4}}

3. /api/chat(Ollama 本地路径,OnUsage 回调)

curl -s http://127.0.0.1:11440/api/chat \
  -H "Content-Type: application/json" \
  -d '{"model":"Qwen3-0.6B","messages":[{"role":"user","content":"Reply with one word: yes"}],"stream":false}'

Ollama 格式响应不含 usage 字段,但 OnUsage 回调从引擎拿到了真实 token。

4. 非流式 /v1/messages(Anthropic proxy 路径)

curl -s http://127.0.0.1:11440/v1/messages \
  -H "Content-Type: application/json" -H "anthropic-version: 2023-06-01" \
  -d '{"model":"Qwen3-0.6B","messages":[{"role":"user","content":"Say hello"}],"max_tokens":5,"stream":false}'

响应中的 usage:

{"input_tokens":10,"output_tokens":5}

5. 流式 /v1/messages(Anthropic proxy 路径)

curl -s http://127.0.0.1:11440/v1/messages \
  -H "Content-Type: application/json" -H "anthropic-version: 2023-06-01" \
  -d '{"model":"Qwen3-0.6B","messages":[{"role":"user","content":"Say hi"}],"max_tokens":5,"stream":true}'

SSE 事件中的 usage:

message_start → {"input_tokens":1,"output_tokens":0}
message_delta → {"output_tokens":5}

DB 记录验证

SELECT model, source, requests, input_tokens, output_tokens, estimated_requests
FROM api_usage_events WHERE day='2026-09-28' AND model='Qwen3-0.6B';
model       source  requests  input_tokens  output_tokens  estimated_requests
----------  ------  --------  ------------  -------------  ------------------
Qwen3-0.6B  local   5         151           180            0
  • input_tokens=151、output_tokens=180:llama-server 真实上报值,不是 countMessageTokens 估算
  • estimated_requests=0:5 条请求全部记录为真值,没有一条走估算

Schema 迁移验证

PRAGMA table_info(api_usage_events);
cid  name                   type     notnull  dflt_value
22   estimated_requests     INTEGER  1        0

estimated_requests 列已通过迁移创建,needsEstimatedRequestsColumn 用 PRAGMA 检查后执行 ALTER TABLE。

OpenAPI 规范验证

openapi/local-api.json 中 APIUsageRow schema 已包含:

"estimated_requests": {"type": "integer", "minimum": 0}

修复前后对比

修复前 修复后
流式 /v1/chat/completions usage 未请求 include_usage,拿不到 请求了,流尾返回真实 usage
/api/chat input tokens countMessageTokens 估算 OnUsage 回调拿到引擎真实 prompt_tokens
/api/chat output tokens 记 0(本地路径未记录) OnUsage 回调拿到引擎真实 completion_tokens
/v1/messages 流式 input 请求侧估算 从 usage chunk 拿到真实 prompt_tokens
估算标记 无法区分 estimated_requests 字段 + 前端"估算"标签
DB schema 无 estimated_requests 列 有,迁移用 PRAGMA 检查

@apapi
apapi merged commit 026b6b0 into main Sep 28, 2026
4 checks passed
@ganisback
ganisback deleted the fix/issue-188-real-usage-tokens branch September 29, 2026 04:20
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

API_KEY的Token 用量统计异常:输入 Tokens 显示为请求次数,输出 Tokens 始终为 0

1 participant