Skip to content

feat(api): expose the lcm router at POST /v1/classify - #680

Merged
Nash0x7E2 merged 3 commits into
acceleratefrom
nash/lcm-classify
Sep 28, 2026
Merged

Nash0x7E2 merged 3 commits into
acceleratefrom
nash/lcm-classify

Conversation

@Nash0x7E2

Copy link
Copy Markdown
Member

Why

The lcm router (TypeSafe's Jev behind classify-fast) could only be reached from inside the router, by guardrails, the prompt-injection screen and sessions. /v1/{modality}/stream serves stt, tts, llm and sts only, so the hosted router answers /v1/lcm/stream with "this deployment does not stream this modality", and a product agent that classifies signal with Jev had to call TypeSafe directly, outside routing, failover and billing.

A classifier is one request and one set of answers, so this exposes it the way search is exposed: a plain POST, not a socket. A caller that retries also has to be able to tell a rate limit from a bad question, which every upstream failure being a 400 made impossible, and the product eval reads input tokens from the response, so the usage comes back with it.

Changes

  • POST /v1/classify: a state (text or a JSON object), questions keyed by the caller's ids (noul, choice, score), an optional target (defaults to classify-fast) and tags. Each answer carries only the fields its type uses, alongside provider, the model version that answered and usage. Server-side only, not client-accessible.
  • An unknown target is a 404, a rate-limited provider a 429, and an overloaded or unreachable one a 503. lcm.ErrRateLimited and lcm.ErrUnavailable carry that without the API layer knowing a vendor's status codes; TypeSafe's StatusError unwraps to them.
  • Clients regenerated for Python, JS, Go, Rust, Ruby, PHP, .NET and Dart. Swift and Kotlin only take client-accessible operations and are unchanged. PHP, .NET and Rust also pick up fields they had not been regenerated for (the latency timeline from Expose the voice-agent latency critical path #674 and command_id), and the Rust Responses::create now sets command_id: None.

Tested locally against Jev: a three-question request returns answers with model: jev-1.13.0 and usage in about 0.3 s, an unknown target returns 404, and the call is recorded in /v1/lcm/stats.

🤖 Generated with Claude Code

Add lcm.ErrRateLimited and lcm.ErrUnavailable so a caller can back off
without knowing a vendor's status codes. TypeSafe's StatusError unwraps
to them for 429 and 503/529, and an API that cannot be reached wraps
ErrUnavailable.
Put typed questions (noul, choice, score) about a piece of text to a
routed classifier and get each answer back with its distribution and
the request's usage. Routed like search: classify-fast by default,
failover, one stat row per request, server-side only.

An unknown target is a 404, a rate-limited provider a 429 and an
overloaded or unreachable one a 503, so a caller knows when to retry.
Clients are not regenerated yet.
Regenerate the Python, JS, Go, Rust, Ruby, PHP, .NET and Dart clients
from the spec. Swift and Kotlin take only client-accessible operations,
so they are unchanged.

PHP, .NET and Rust also pick up the latency timeline and command_id
fields they had not been regenerated for, and the Rust Responses::create
now sets command_id to None.
@Nash0x7E2 Nash0x7E2 self-assigned this Sep 28, 2026
@Nash0x7E2
Nash0x7E2 marked this pull request as ready for review September 28, 2026 20:49
@Nash0x7E2
Nash0x7E2 merged commit 70877f4 into accelerate Sep 28, 2026
14 of 18 checks passed
@Nash0x7E2
Nash0x7E2 deleted the nash/lcm-classify branch September 28, 2026 20:49
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant