Skip to content

[CLM 2/2] Reach the encoder, answer requests, and check against CLM - #29

Open
xiaoyu-xyz wants to merge 14 commits into
ThinkFlowLab:mainfrom
xiaoyu-xyz:clm-2-engine
Open

xiaoyu-xyz wants to merge 14 commits into
ThinkFlowLab:mainfrom
xiaoyu-xyz:clm-2-engine

Conversation

@xiaoyu-xyz

@xiaoyu-xyz xiaoyu-xyz commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

Second of two, stacked on #28. Merge #28 first — this branch is stacked on it, so until then the diff below shows both PRs and after #28 merges it is only this one's. Reviewing commit by commit is cleanest in the meantime; the last nine commits are this PR's.

Core code: 681 lines (embedding 142, serve 539), plus the 15 lines it adds to scoring for the legend, counting lines that are neither blank nor start with //. serve was 350 by that count when this PR was opened, and the literal handling below is 189 of the lines since. clm-run (92 lines) and the recipe scripts are model tooling, like laya-run and export_weights.py, so they are not counted.

The embeddings client

The engine does not compute embeddings, so this is the client to the process that does. An Encoder trait with two implementations:

  • HttpEncoder reproduces the reference client — a POST of {model, input, encoding_format, truncate_prompt_tokens}, a base64 f32 payload per input in index order, L2-normalised on receipt, batched. That is what vllm serve --runner pooling exposes. The truncation limit defaults to the reference's 2048, which is the deployment's --max-model-len; a text past it is truncated there but rejected by a server asked to embed it whole.
  • HashingEncoder derives a vector from a SHA-256 of the text, exactly as head_oracle.py does, so the whole decision path runs with no GPU and no weights and can still be compared against the Python oracle.

The request path

Two text functions decide what reaches the encoder and both are reproduced exactly, because a separator in the wrong place shifts every probability without failing anything:

  • the state head sees the context and the question joined by a blank line, so a question belongs in instructions and not repeated in the state;
  • the action head sees each candidate's own text with nothing prefixed for choice, but "<key>: <text>" for noul, falling back to "Yes. This is true: <question>".

to_text renders a state, description or criteria as the prose the heads were trained on: objects become key: value, top-level fields separated by a blank line and nested ones indented, arrays one - item line each, empty containers taking the key: value branch.

Numbers are parsed and spelled the way the reference does: serde_json's default float parsing is not correctly rounded, so a literal holding more digits than a double does could land one ULP away and render as a different number — 7.8190461323667115 came out 7.819046132366712. The float_roundtrip feature makes the parse bit-identical to json.loads, and is what keeps the following true: numbers are spelled the way str(float) spells them — 1e-05, 1e-07, and a .0 kept on an integral float. serde_json disagrees on both ends, and since the heads are trained on these strings a differently spelled number is a different input. Rust's shortest digits also differ from CPython's on the last digit when a value sits exactly between two equally short decimals, so the digits are re-rendered at the shortest round-tripping precision; fuzzed against repr over 41k doubles, 0 differences.

serde_json gains preserve_order, which is a correctness fix rather than a preference: the default BTreeMap sorts keys, and both to_text and the reference preserve the caller's field order. Without it a state whose fields are not in alphabetical order reaches the encoder re-sorted, which is a different text and a different decision. The oracle caught exactly that.

clm-run takes a converted checkpoint and an embeddings endpoint and answers one request per line, in the shape the other engines use, so the same binary drives a stub encoder on a laptop and a real one on a GPU box.

Checks

Text construction, byte for byte. text_oracle.py writes the reference's own output for to_text, state_text and candidates — 9 and 72 cases, plus three whole request lines that a Value cannot express — by importing clm.schema, so tests/clm/text.rs compares against the implementation rather than a transcription of it.

Against a real encoder. transformers_encoder.py serves Qwen3-8B behind the exact endpoint the deployment uses, so vLLM is not needed to verify the client, and compare_with_reference.py compares omni-clm with CLM's own heads and schema over the same vectors:

PASS choice_two   choice  max|dp|=1.72e-05 choice=billing/billing
PASS choice_five  choice  max|dp|=1.97e-02 choice=a/a
PASS score_three  score   max|dp|=1.08e-02 score=0.385166/0.403241
PASS noul_stmt    noul    0.631038 vs 0.646805
PASS noul_numbers noul    0.323871 vs 0.321280

That comparison is what established the exp(logit_scale) cap in #28 — with the uncapped 100.82 every probability was about 0.8 % off, and no CPU-side oracle can see it, because both sides of those share the constant.

On the tolerance. CLM scores with exp(logit_scale) = 100, so a difference of d in a cosine becomes 100 d in the logit. Two encoder forward passes that disagree by ~1e-4 in the cosine — which is what vLLM's bf16 pooling and Transformers' forward produce for Qwen3-8B — therefore disagree by ~1e-2 in the probabilities. That is a property of how this model scores, not a defect on either side.

The engine's own arithmetic is checked separately and much more tightly: serve both sides one frozen set of vectors and the implementations agree to 3e-06. That is the number that says the port is faithful; this script is a smoke test over a real encoder.

--temperature is checked the same way against a stub endpoint, where both sides get identical vectors and any difference is the two implementations': at 2.5 the comparison is 4/4 PASS with max|dp| around 1e-6, and with the argument not passed on to clm-run it is 4/4 FAIL at around 0.18.

15 tests with the oracles, 11 passed / 4 ignored without them. The CI job's commands pass locally: cargo fmt --all --check, cargo clippy --workspace --locked --all-targets -- -D warnings, cargo test --workspace --locked, cargo build --workspace --release --locked.

@hsliuustc0106 hsliuustc0106 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed commit fd841e1967e0a58b0a1af7b8425e83579f1c576c.

  • [P2] Preserve empty-container candidate text. src/models/clm/src/serve.rs:158 falls back to the key whenever to_text is empty; upstream falls back only for null or the empty string. Criteria {"a":{},"b":[]} therefore become ["a","b"] here instead of ["", ""], changing embeddings/probabilities. The noul fallback has the same problem. Reproduced with a CPU Rust harness and compared with upstream schema.py.
  • [P2] Include the score legend in serialized answers. src/models/clm/src/serve.rs:355 omits the legend that upstream answer_from_probs emits for every score question. Criterion descriptions are lost, so consumers cannot recover the score-level mapping from the response. The current answer type discards these descriptions before serialization.
  • [P2] Pass --temperature to the native comparison process. recipe/clm/native/compare_with_reference.py:83 omits this argument when launching clm-run, while line 98 uses args.temperature for the reference. Any nondefault temperature compares different configurations and can falsely report a port regression.
    Tests: 2 passed, 4 ignored without external checkpoint/text oracles. Inherits #28 loader finding.

@xiaoyu-xyz

Copy link
Copy Markdown
Contributor Author

Fixed in d276bf8, all three. (Rebased on the new #28, so the SHAs moved.)

  • candidates now falls back to the key only for null and the empty string, in both the choice and noul branches. Both shapes are covered by text_oracle.py now (30 → 42 cases), and reverting either branch turns the byte comparison red.
  • A score answer carries legend, built the way answer_from_probs builds it — the keys zipped with the candidate texts.
  • compare_with_reference.py passes --temperature to clm-run. At 2.5 the comparison is 4/4 PASS; with the argument still missing it is 4/4 FAIL with max|dp| around 0.18.

@Levius-Fubuki Levius-Fubuki left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed d276bf8a038abb42682aa24bceb1d0ca338ec183, including its incremental changes on #28 and the inherited config/scoring code.

The original empty-container, score-legend, and temperature-forwarding repairs are present. Fresh Linux/Rust 1.98.1 formatting, strict all-target Clippy, workspace tests and release build passed (17 tests in the default workspace run). All eight CLM tests subsequently passed with real head weights and external oracles: 16 FP32 tensor hashes, five head decisions, six state renderings and 42 state/question pairs. The text oracle imports official clm.schema pinned to bb42c6c5bf914fd449bed2f6ca65be80602cb1f7.

A separate end-to-end CPU comparison used the real head weights, official PyTorch heads/schema, one deterministic HTTP embedding stub, and --temperature 2.5: all four cases passed on this head at tolerance 1e-4. The previous comparison script failed all four against the same setup because it omitted the native temperature argument (maximum probability difference about 0.177).

The published head checkpoint at e939398d4556fcd9400c76fa8c5a513202f42b0a was SHA-256 verified. The head oracle uses NumPy and is not itself independent PyTorch numerical certification. The targeted protocol/text findings below were reproduced separately and are not covered by the passing 42-case oracle. This PR also inherits the #28 findings, and the stack needs its existing main conflicts resolved. No GPU or real Qwen/vLLM inference was run.

Changes requested for the two reference-compatibility issues below.

let body = serde_json::json!({
"model": self.model,
"input": texts,
"encoding_format": "base64",

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] Preserve the reference client's prompt-truncation setting

The upstream Embedder sends truncate_prompt_tokens=2048 by default, matching its documented vLLM --max-model-len 2048 deployment, but this request body omits that setting. Long state/candidate text therefore uses different server behavior: an over-limit request that the reference truncates can be rejected by the native path. I captured both real clients against a local endpoint enforcing that wire contract: upstream included the field and succeeded; HttpEncoder omitted it and received 400 for the same long input. This was a protocol stub, not a real vLLM run. Add a configurable truncation limit with the reference default and serialize it in the request. Reference: https://github.com/Contrastive-LM/CLM/blob/bb42c6c5bf914fd449bed2f6ca65be80602cb1f7/src/clm/embedder.py#L42 .

Comment thread src/models/clm/src/serve.rs Outdated
Value::String(s) => s.clone(),
Value::Bool(true) => "true".to_string(),
Value::Bool(false) => "false".to_string(),
Value::Number(n) => n.to_string(),

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] Match the reference's number formatting before embedding

JSON-number to_string() is not equivalent to Python str(float), which upstream uses in to_text. For the same JSON value 1e-5, this function produces 0.00001 while the reference produces 1e-05; for 1e-7, it produces 1e-7 versus 1e-07. These values are valid in structured state or criteria, so the two engines embed different strings and can make different decisions even with identical model weights. The current oracle only exercises 0.5 and misses exponent-format cases. Match Python's numeric rendering for supported JSON values and add small/large exponent cases to the upstream-derived oracle.

@xiaoyu-xyz

Copy link
Copy Markdown
Contributor Author

Fixed in df2da33; the stack is rebased on main (58b8cbe), so the conflicts are gone, and #28's two findings are fixed in e1a1211.

  • Truncation. HttpEncoder sends truncate_prompt_tokens, defaulting to the reference's 2048, and clm-run exposes it as --max-tokens N|none. Captured off the wire: default → 2048, --max-tokens 64 → 64, none → the field absent.
  • Number spelling. to_text reproduces str(float): 1e-05, 1e-07, -1e-06, and a .0 kept on an integral float. Rust's shortest digits agree with CPython's on the count but not always on the last one — for a value exactly between two equally short decimals CPython rounds to even and Rust away from zero — so the digits are re-rendered at the shortest round-tripping precision, which goes through the same correct rounding as %.*e. Fuzzed against repr over 41k doubles covering every decimal exponent: 0 differences.

The text oracle gained an exponent rendering (now 7 renderings, 49 pairs), and the byte comparison fails without the fix. Strict Clippy, workspace tests and the release build pass locally.

@xiaoyu-xyz

Copy link
Copy Markdown
Contributor Author

Ran the real-encoder comparison on the current head, 968b61b — which supersedes the df2da33 named above; that one adds numeric criteria to the text oracle (7 renderings, 56 pairs).

Qwen3-8B bf16 served behind /v1/embeddings on an RTX 4090 (sm_89, CUDA 13.0), with the CLM reference heads reading the same endpoint:

run result
--temperature 1.0 4/4 PASS — choice_two 1.72e-05, choice_five 1.97e-02, score_three 0.385166/0.403241, noul 0.631038 vs 0.646805
--temperature 2.5 4/4 PASS — max|dp| 7.57e-04 to 9.04e-03
--temperature 2.5, script from before the fix 4/4 FAIL — score=0.385166/0.761482

The first row reproduces the numbers already in the description to the last digit but one: score_three moves from 0.385165 to 0.385166, which is the ~1e-7 the redundant second normalisation was contributing.

The third row is the --temperature finding on a real encoder rather than a stub: with the argument not forwarded, clm-run runs at 1.0 while the reference runs at 2.5, and the comparison reports a port regression that is not there.

@Levius-Fubuki

Copy link
Copy Markdown
Collaborator

Rechecked 968b61b: truncation forwarding and the original exponent-format cases are fixed. Numeric text still differs after JSON parsing: 7.8190461323667115 becomes 7.819046132366712 in both state and candidate embedding inputs, confirmed through the actual clm-run HTTP requests. Please cover the full JSON-input-to-embedding path and align numeric parsing with the reference; formatting already-parsed floats alone misses this case.

@xiaoyu-xyz

Copy link
Copy Markdown
Contributor Author

Fixed in 44a55b3, and you were right that formatting was the wrong place to look — the parse was wrong.

serde_json's default float parsing is not correctly rounded. 7.8190461323667115 parsed one ULP high (…b8a5 where json.loads gives …b8a4) and rendered 7.819046132366712; 9007199254740993.0 became 9007199254740994.0 where the reference gives 9007199254740992.0. So the heads were shown a different number, in the state and in the criteria alike.

serde_json's float_roundtrip feature makes the parse bit-identical to json.loads and pulls in no new dependency — the lockfile is unchanged.

Coverage now runs the full path rather than the formatter:

  • json_numbers_parse_to_the_same_doubles_as_the_reference parses the literals with from_str and asserts the bit patterns CPython produces, then the rendered text.
  • The text oracle gained a state carrying both literals, so the byte-for-byte comparison catches it through to_text and through candidates — 8 renderings and 64 state/question pairs now.

Removing the feature turns all three red, including left: "seventeen: 7.819046132366712" against right: "seventeen: 7.8190461323667115".

12 tests pass with the oracles, strict Clippy and the workspace build are clean on the head.

@xiaoyu-xyz

Copy link
Copy Markdown
Contributor Author

Follow-up on the layer you named: the earlier fix was checked at to_text/candidates, which is one step too late. b9907af covers the request line itself.

Captured from a real clm-run request holding both literals, against a server that records the embeddings payload:

state_text    "seventeen: 7.8190461323667115\n\ninexact: 9007199254740992.0\n\nPick"
candidates    ["7.8190461323667115", "plain"]

Both match clm.schema.state_text and clm.schema.candidates on the same request exactly.

That is now pinned without a server: a_request_line_reaches_the_encoder_with_the_reference_text takes a JSON line as clm-run reads it, runs Request::parse and prepare, and asserts the strings that reach the encoder. Without float_roundtrip it fails at left: "seventeen: 7.819046132366712…" against right: "seventeen: 7.8190461323667115…".

13 tests pass with the oracles on the head.

@Levius-Fubuki

Copy link
Copy Markdown
Collaborator

Rechecked b9907af: all 3,573 previous float mismatches are fixed. Large integers still differ: 18446744073709551616 reaches the encoder as "1.8446744073709552e+19" in both state and criteria. Please preserve exact integer text and add raw-request coverage.

@hsliuustc0106 hsliuustc0106 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Independent local review — CLM 2/2

Verdict: changes requested — same convention finding as #28; the engineering is verified.

[Medium] Test bodies in src/models/clm/tests/{checkpoint,text}.rs — crate-local tests/ under src/, per CONTRIBUTING (at your base and on current main) they belong in repo-root tests/clm/. Same fix as #28.

Verified: clm-run is fully synchronous (no async runtime), so the blocking reqwest encoder client is the correct shape; serde_json carries both preserve_order and float_roundtrip as claimed; Engine::decide keeps state/candidate text construction, the noul "<key>: <text>" fallback, and answer assembly model-owned. Locally at head b9907af3: 9 passed / 4 ignored (oracle/checkpoint-gated, disclosed), clippy clean. The byte-for-byte text-oracle discipline (importing the reference clm.schema rather than transcribing it), 3e-06 frozen-vector agreement, and the temperature stub-endpoint check are author-reported with recorded evidence.

Architecture note for the future worker: when CLM grows an HTTP worker, adopt the prepare → execute → finish split and omni-runtime admission directly (per docs/architecture.md and #78/#80) — the library's clean separation makes that easy. #28 merges first.

Local reviewer report per the repo review skill; reflects head b9907af3 only.

@hsliuustc0106 hsliuustc0106 mentioned this pull request Oct 5, 2026
2 of 4 tasks
@xiaoyu-xyz
xiaoyu-xyz force-pushed the clm-2-engine branch 2 times, most recently from 6a40e0e to a44e269 Compare October 5, 2026 06:40
@xiaoyu-xyz

Copy link
Copy Markdown
Contributor Author

The float cases are fixed, but the integer one is real and its natural fix is blocked — I would rather show you the block than pick for you.

arbitrary_precision is what preserves an integer literal's digits. I implemented it, and 18446744073709551616 then reaches the encoder as itself, covered by the oracle, a bit-level test and the raw request-line test.

But a Cargo feature is not scoped to the crate that asks for it. serde_json is one package at one version across this workspace and features unify, so enabling it here also changes Value::Number inside src/models/cua_s1/native — whose json.rs builds numbers with Number::from_f64 and serializes with serde_json::to_string. Toggling that single feature flips three of its tests:

cargo test -p omni-cua-s1-native
  with arbitrary_precision:  3 passed, 3 failed
    json::tests::dumps_matches_python
    json::tests::rejects_what_the_contract_rejects
    contract::tests::maps_a_request
  without it:                6 passed, 0 failed

So I reverted it. The branch is green again: cargo test --workspace --locked → 34 passed, 0 failed.

Two ways out, and the choice is not only mine:

  1. CLM carries the number text itself. Deserialize state and criteria as RawValue and render numbers from the source literal with a small JSON walker instead of serde_json::Value. Scoped to this crate, but it means a second traversal that has to be kept in step with to_text.
  2. omni-cua-s1-native stops depending on Number's Display. If its dumper formatted numbers itself instead of round-tripping through serde_json::to_string, it would be immune to this feature and CLM could set it.

I lean to (2): it removes a coupling rather than adding a second JSON implementation. But it edits a crate this PR does not otherwise touch, so I did not do it unasked.

The limitation as it stands: an integer outside i64::MIN ..= u64::MAX is parsed as a double and rounded. Everything inside that range is exact, including the 9007199254740993.0 case from the previous round.

@xiaoyu-xyz

Copy link
Copy Markdown
Contributor Author

Fixed in 1a15d44, and it needed the text rather than a feature.

Why not arbitrary_precision. It fixes this exactly, and I had it working — but a Cargo feature is not scoped to the crate that asks for it. serde_json is one package at one version across this workspace and features unify, so it also changed Value::Number inside src/models/cua_s1/native, whose json.rs serialises numbers with serde_json::to_string:

cargo test -p omni-cua-s1-native
  with arbitrary_precision:  3 passed, 3 failed
    json::tests::dumps_matches_python
    json::tests::rejects_what_the_contract_rejects
    contract::tests::maps_a_request
  without it:                6 passed, 0 failed

What it does instead. The fields that get rendered — state, and each question's instructions and criteria — are deserialized a second time as RawValue, so each keeps the text it arrived as, and each is rendered from its own text with its own literals. raw_value adds RawValue without changing Number, and cargo test --workspace --locked is 36 passed, 0 failed.

Two things had to be right, and both have tests:

  • Rendering a field twice must not shift the literals. The state is rendered once per request now, not once per question.
  • Nothing may depend on the order the request's keys appear in. Each field is lexed from its own text, so {"questions": …, "state": …} works. the_literals_line_up_with_the_values pins the lexer against a value walk, including a document whose strings contain digits and escapes.

to_text still works on a Value and falls back to serde_json's integer types; to_text_json and Request::parse_line are the paths that keep the digits, and clm-run now uses the latter. The oracle gained a rendering with 2**64, 2**128 and a negative of each, and the case loop now drives a real request line instead of calling candidates on a Value.

Removing the text turns three red at left: "…big: 1.8446744073709552e+19…" against right: "…big: 18446744073709551616…".

One thing to flag: this costs 169 core lines in serve.rs, so #29 is now 661 against the 500-line budget rather than 492. Carrying the text is the correct fix and it is the smallest one that is order-independent and repeat-safe, but I would rather say the number than leave you to count it. If you want the budget held, the number handling can come out as a [CLM 3/3] — it is self-contained and #29 goes back to 492.

xiaoyu-xyz added 9 commits October 5, 2026 15:10
CLM is the second model tracked in ThinkFlowLab#9, and it decides differently in a way LAYA does not
cover: the engine does not compute embeddings. A frozen Qwen3-8B encoder runs as its own
process, and the engine owns everything after it. This is the half that can be checked
without a GPU -- the checkpoint, the head geometry, and the decision arithmetic.

A CLM checkpoint is a torch.save dict, so recipe/clm/native/export_weights.py converts it
first: tensors to safetensors with the head name as a prefix, and cfg, hidden_size,
projection_dim and logit_scale into the metadata. The published checkpoint is 75 MB and
holds the two heads, not the 8B encoder; the export is 16 tensors and 18.9 M parameters.

- config reads the head geometry from the metadata and checks it.
- weights lists the expected inventory from that geometry, checks shapes before reading,
  and loads FP32.
- scoring projects a state and each candidate, L2-normalises both, scores them by cosine
  under exp(logit_scale), softmaxes across the question's candidates, and assembles
  choice, noul and score. The three types differ only after the distribution exists, so
  they share one path. confidence is the top probability minus the mean of the rest,
  clamped, and 1.0 for a single candidate, which is the definition src/clm/schema.py uses.

exp(logit_scale) is capped at 100, as the reference caps it in heads.py. The published
logit_scale is 4.6132, whose exponential is 100.82, so the cap binds: without it every
probability is about 0.8 % off.

Checks. head_oracle.py is an independent NumPy implementation of the same arithmetic, and
tests/checkpoint.rs checks 16 tensor conversion hashes against the export oracle and five
decisions against head_oracle.py. Default cargo test needs no checkpoint. fmt, clippy -D
warnings.
`head_tensors` lists `norms.N.*` only when `cfg.layernorm` is set, so a
`layernorm: false` checkpoint is complete without them — but `load_head`
read them unconditionally. A valid checkpoint passed the inventory check
and then failed to load.

The new test writes both shapes plus one that declares LayerNorm and
omits the tensors, and checks the loader accepts the first two and
rejects the third.

Found in review.
…labels

- `Config::from_metadata` applied the top-level `hidden_size`/`projection_dim`
  fallback *after* deserializing `cfg`, but both are required fields, so a
  `cfg` that omits them failed to parse and the fallback never ran. The
  geometry is now filled in before deserialization, and a `cfg` that omits it
  loads while one that omits it with no top-level copy is still an error.
- `Answer::label` used `max_by`, which keeps the *last* of several equal
  values. `schema.label_of` is `max(p, key=p.__getitem__)`, which keeps the
  first, so a tie in a score distribution named a different level.

Found in review.
CONTRIBUTING puts every test body in the repository-level `tests/` tree and
asks for explicit `[[test]]` registration, which is what `src/models/laya`
does. `src/models/clm/tests/` was a crate-local directory under `src/`, which
that rule prohibits; the tests ran either way, so this is placement rather than
breakage.

Found in review.
The engine does not compute embeddings: a frozen Qwen3-8B encoder runs as its own process
behind an OpenAI-compatible /v1/embeddings endpoint, and this is the client to it.

An Encoder trait with two implementations. HttpEncoder reproduces the reference client --
a POST of {model, input, encoding_format}, a base64 f32 payload per input in index order,
L2-normalised on receipt, batched -- which is what `vllm serve --runner pooling` exposes.
HashingEncoder derives a vector from a SHA-256 of the text, exactly as
recipe/clm/native/head_oracle.py does, so the decision path runs with no GPU and no
weights and can still be compared against the Python oracle.

fmt, clippy -D warnings, builds on the head loader.
The crate could load heads and score candidates. This adds the client to the encoder and
the path from a request to typed answers, plus the harness that compares the result with
CLM's own implementation.

Core code: 489 lines (embedding 178, serve 311). clm-run and the recipe scripts are model
tooling, like laya-run and export_weights.py, so they are not counted.

embedding carries an Encoder trait with two implementations. HttpEncoder reproduces the
reference client -- a POST of {model, input, encoding_format}, a base64 f32 payload per
input in index order, L2-normalised on receipt, batched -- which is what
`vllm serve --runner pooling` exposes. HashingEncoder derives a vector from a SHA-256 of
the text, exactly as head_oracle.py does, so the decision path runs with no GPU and no
weights and can still be compared against the Python oracle.

serve is the request path. Two text functions decide what reaches the encoder and both are
reproduced exactly, because a separator in the wrong place shifts every probability
without failing:

- the state head sees the context and the question joined by a blank line, so a question
  belongs in `instructions` and not repeated in the state;
- the action head sees each candidate's own text with nothing prefixed for `choice`, but
  "<key>: <text>" for `noul`, falling back to "Yes. This is true: <question>".

to_text renders a state, description or criteria as the prose the heads were trained on:
objects become `key: value`, top-level fields separated by a blank line and nested ones
indented, arrays one `- item` line each, empty containers taking the `key: value` branch.

serde_json gains `preserve_order`, which is a correctness fix rather than a preference:
the default BTreeMap sorts keys, and both to_text and the reference preserve the caller's
field order. Without it a state whose fields are not in alphabetical order reaches the
encoder re-sorted, which is a different text and a different decision.

clm-run takes a converted checkpoint and an embeddings endpoint and answers one request per
line, in the shape the other engines use, so the same binary drives a stub encoder on a
laptop and a real one on a GPU box.

Checks. text_oracle.py writes the reference's own output for to_text, state_text and
candidates -- 6 and 30 cases -- by importing clm.schema, so tests/text.rs compares against
the implementation rather than a transcription of it, byte for byte.

transformers_encoder.py serves Qwen3-8B behind the endpoint the deployment uses -- vLLM is
not needed to verify the client -- and compare_with_reference.py compares omni-clm with
CLM's own heads and schema over the same vectors. That comparison is what showed the
exp(logit_scale) cap in the first PR: with the uncapped 100.82 every probability was about
0.8 % off, and no CPU-side oracle can see it, because both sides of those share the
constant. All four cases agree within 2e-3 for the two-candidate question.

8 tests, fmt, clippy -D warnings.
…imposes

All four cases now pass against a real encoder, and the script says why the number is what
it is rather than leaving the next reader to rediscover it.

CLM scores with exp(logit_scale) = 100, so a difference of d in a cosine similarity becomes
100 d in the logit. Two encoder forward passes that disagree by about 1e-4 in the cosine --
which is what vLLM's bf16 pooling and Transformers' forward produce for Qwen3-8B -- therefore
disagree by about 1e-2 in the probabilities. That is a property of how this model scores, not
a defect on either side, and the old 2e-3 default was asking for agreement the pipeline
cannot give.

The engine's own arithmetic is much tighter and is checked separately: hold the vectors
fixed -- serve both sides one frozen set -- and the implementations agree to 3e-06.

    PASS choice_two   max|dp|=1.72e-05  choice=billing/billing
    PASS choice_five  max|dp|=1.97e-02  choice=a/a
    PASS score_three  score=0.385165/0.403241
    PASS noul_stmt    0.631038 vs 0.646805
…ature

Three findings, all measured against the reference:

- `candidates` fell back to the option key whenever a rendering came out
  empty. `schema.candidates` falls back only for null and the empty string,
  so `{"a": {}, "b": []}` is `["", ""]`, not `["a", "b"]`. The same test
  applies to the `noul` branch. `text_oracle.py` now covers both shapes.
- A `score` answer dropped the level legend that `answer_from_probs`
  returns and that `client.ScoreAnswer` declares, so a consumer could not
  read a score back as a level. The legend is the keys zipped with the
  candidate texts, which is what the reference builds it from.
- `compare_with_reference.py` launched `clm-run` at its default
  temperature while running the reference at `--temperature`, so any
  non-default compared two different configurations.

Found in review.
- `HttpEncoder` omitted `truncate_prompt_tokens`, which `embedder.py` sends as its
  `max_tokens` — 2048 by default, matching the deployment's `--max-model-len`. A
  text past the limit is truncated by the reference but rejected by a server asked
  to embed it whole. The limit is now a field with the reference default, it is
  serialized, and `clm-run` exposes it as `--max-tokens N|none`.
- `to_text` rendered numbers with `serde_json`, which disagrees with Python's
  `str(float)`: `1e-5` came out `0.00001` where the reference writes `1e-05`, and an
  integral float lost its `.0`. The heads are trained on these strings, so the two
  engines embedded different text for the same state. `python_float` reproduces
  CPython's rendering; it was fuzzed against `repr` over 41k doubles covering every
  decimal exponent.

Found in review.
xiaoyu-xyz added 4 commits October 5, 2026 15:10
`serde_json`'s default float parsing is not correctly rounded, so a literal with
more digits than a double holds can land on the neighbouring double:
`7.8190461323667115` parsed one ULP high and rendered `7.819046132366712`, and
`9007199254740993.0` became `9007199254740994.0` where `json.loads` gives
`9007199254740992.0`. Getting the formatting of an already-parsed float right
did not help — the parse was wrong, so the heads were shown a different number.

`serde_json`'s `float_roundtrip` feature makes the parse bit-identical to
`json.loads`, and pulls in no new dependency. The text oracle gained a state with
17 significant digits, which the byte comparison catches through both `to_text`
and `candidates`, and a test pins the parsed bit patterns directly.

Found in review.
The parsing fix was checked at `to_text` and `candidates`. The layer the bug
actually lived at is one step earlier: a request line goes through
`Request::parse` and `prepare` before either. The new test takes a JSON line
as `clm-run` reads it and asserts the exact state text and candidate texts that
reach the encoder -- the two strings also confirmed against the reference over a
real HTTP request, by capturing what `clm-run` sends.

Without `float_roundtrip` it fails at `left: "seventeen: 7.819046132366712..."`
against `right: "seventeen: 7.8190461323667115..."`.
The same rule `src/models/clm/tests/` broke in ThinkFlowLab#28 applies here, and to the
inline test in `embedding.rs`: CONTRIBUTING keeps test bodies in the
repository-level `tests/` tree, and for a private item only the `#[cfg(test)]`
and the module path may stay in `src/`. `HttpEncoder::body` is private, so that
is the `#[path]` form, with the body in `tests/clm/embedding.rs`.

Found in review.
`serde_json` cannot hand back what an integer literal said: one that does not fit
an `i64` or a `u64` is parsed straight to a double, so `18446744073709551616`
reached the encoder as `1.8446744073709552e+19`. `json.loads` gives Python an
arbitrary-precision `int` instead and `str` prints every digit.

`arbitrary_precision` would fix it but is a Cargo feature, so it would reach
every crate in the workspace — it breaks `omni-cua-s1-native`'s
`json::tests::dumps_matches_python`, which serialises `Value::Number`. This
carries the text instead: the fields that get rendered are deserialized a second
time as `RawValue`, and each is rendered from its own text with its own literals.
`raw_value` adds `RawValue` without changing `Number`, and the workspace stays
green.

Two things had to be got right, and both are tested:

- Rendering the same field twice must not shift which literal belongs to which
  number. The state is rendered once now rather than once per question.
- Nothing may depend on the order the request's keys appear in. Each field is
  lexed from its own text, so `{"questions": …, "state": …}` is fine, and
  `the_literals_line_up_with_the_values` pins the lexer against a value walk.

`to_text` keeps working on a `Value` and falls back to `serde_json`'s integer
types; `to_text_json` and `Request::parse_line` are the paths that keep the
digits, and `clm-run` uses the latter. The oracle gained a rendering with 2**64,
2**128 and a negative of each, and the case loop now goes through a real request
line rather than calling `candidates` on a `Value`.

Removing the text turns three red at `left: "…big: 1.8446744073709552e+19…"`
against `right: "…big: 18446744073709551616…"`.

Found in review.
@Levius-Fubuki

Copy link
Copy Markdown
Collaborator

Rechecked 1a15d44: the original bigint case is fixed, but noul now mismatches the reference. Criteria {"true":1,"false":2} produce false: 1 / true: 2 instead of false: 2 / true: 1, because literals are consumed in source order. Default candidates also still round bigint instructions. Please bind literals to the selected values and reuse the raw-rendered instructions, with raw-request regression tests.

The previous commit kept an integer literal's digits by lexing each rendered field
and pairing the literals with the numbers a depth-first walk meets. That pairing is
positional, and it is only right when the walk reads the values in the order they
were written. A `noul` does not: the reference reads its two descriptions in
`NOUL_KEYS` order, `false` before `true`, so

    "criteria": {"true": 1, "false": 2}

reached the action head as `false: 1` / `true: 2`, where `clm.schema` produces
`false: 2` / `true: 1`. A wrong candidate text is not a wrong spelling, it is a
different answer.

The descriptions of a `choice` or a `noul` are now found by the key they were
written under — `HashMap<String, &RawValue>` over the criteria text — and each is
rendered from its own literal, so the order the keys are read in cannot matter. A
`score`'s levels are a list, which has one order, and keep the cursor.

The default `noul` candidates are the question's own statement, and they were built
from the parsed `instructions`: a statement that is a number lost its digits there
even though `state_text` had them,

    "instructions": 18446744073709551616

giving `false: No. This is false: 1.8446744073709552e+19`. The statement rendered
from its text is now passed down and reused, so both heads see one text.

The oracle gained three whole request lines, since a `Value` cannot hold an integer
above `u64::MAX` and `QuestionRequest` holds `instructions` as a `String`, and
`raw_requests_match_the_reference_byte_for_byte` runs them through
`Request::parse_line`. Against the previous commit it fails at
`left: ["false: 1", "true: 2"]` where the reference says `["false: 2", "true: 1"]`.

Found in review.
@xiaoyu-xyz

Copy link
Copy Markdown
Contributor Author

Both findings were real, and the first one is worse than a spelling difference. Fixed in
a2019f9, on a branch rebased onto main (7f39ac4).

The literal now follows its key. You are right that the pairing was positional. The
lexer collects literals in source order and the render walk consumed them in the order it
met the numbers, which is only the same order for a choice (source order) and a score
(a list). A noul reads its two descriptions in NOUL_KEYS order, false before true,
so {"true": 1, "false": 2} walked false first and took 1. The descriptions of a
choice or a noul are now looked up by the key they were written under —
HashMap<String, &RawValue> over the criteria text — and each value is rendered from its
own literals, so the order the keys are read in cannot matter. A score's levels are a
list, which has one order, and keep the cursor.

The default candidates reuse the rendered statement. They were built from the parsed
instructions, so a statement that is a number lost its digits there even though
state_text kept them: "instructions": 18446744073709551616 gave
false: No. This is false: 1.8446744073709552e+19. The string rendered from the
statement's own text is now passed down and used for both heads.

Tests. A Value cannot hold an integer above u64::MAX, and QuestionRequest holds
instructions as a String, so the cases that need the literal cannot be built from
Python values. text_oracle.py gained three whole request lines and
raw_requests_match_the_reference_byte_for_byte runs them through Request::parse_line
and compares byte-for-byte against clm.schema. Against 66bffe9 it fails at

left:  ["false: 1", "true: 2"]
right: ["false: 2", "true: 1"]

which is your case. Two unit tests cover the same two things without the oracle.

This was not cosmetic. compare_with_reference.py gained the noul_numbers case and
I ran the same case list, the same Qwen3-8B encoder and the same checkpoint against two
binaries that differ only in this commit (RTX 4090, CUDA 13.0, driver 595.71.05):

############ PRE-FIX BINARY ############
PASS choice_two   choice  max|dp|=1.72e-05 choice=billing/billing
PASS choice_five  choice  max|dp|=1.97e-02 choice=a/a
PASS score_three  score   max|dp|=1.08e-02 score=0.385166/0.403241
PASS noul_stmt    noul    0.631038 vs 0.646805
FAIL noul_numbers noul    0.429956 vs 0.321280

1 FAILED (tolerance 0.025; the engine alone agrees to 3e-06 when the vectors are fixed)

############ FIXED BINARY ############
... the same four lines ...
PASS noul_numbers noul    0.323871 vs 0.321280

all cases agree

noul was off by 0.109 where the tolerance is 0.025, so the wrong pairing was a different
answer rather than a differently spelled one, and every other case is identical between
the two binaries.

Checks on the rebased head: cargo fmt --all --check and
cargo clippy --workspace --locked --all-targets -- -D warnings clean,
cargo test --workspace --locked 95 passed / 0 failed, and the oracle-backed tests
(CLM_EXPORT and CLM_TEXT_ORACLE) 5 passed / 0 failed including the new one.

@hsliuustc0106

Copy link
Copy Markdown
Contributor

DO WE HAVE AN INTEGRATION example with system1-agents with a video demo?

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants