Skip to content

[CLM 1/2] Checkpoint and typed decision scoring - #26

Closed
xiaoyu-xyz wants to merge 1 commit into
ThinkFlowLab:mainfrom
xiaoyu-xyz:feat/clm-engine
Closed

xiaoyu-xyz wants to merge 1 commit into
ThinkFlowLab:mainfrom
xiaoyu-xyz:feat/clm-engine

Conversation

@xiaoyu-xyz

@xiaoyu-xyz xiaoyu-xyz commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

CLM is the second model tracked in #9, and it decides differently in a way LAYA does not cover: the engine does not compute embeddings. A frozen Qwen3-8B encoder runs as its own process, and the engine owns everything after it. This is the half that can be checked without a GPU — the checkpoint, the head geometry, and the decision arithmetic.

Core code: 483 lines (config 75, weights 146, scoring 256, lib 6), within the 500-line budget. Tests and the export/oracle scripts are separate.

This is the first of three, stacked:

core
this PR checkpoint, geometry, scoring 483
next the embeddings client 178
then the request path, clm-run, and the comparison against CLM 392

The checkpoint is converted first

A CLM checkpoint is a torch.save dict, so it is a pickle and no non-Python reader can open it. recipe/clm/native/export_weights.py writes the tensors to safetensors with the head name as a prefix, and keeps cfg, hidden_size, projection_dim and logit_scale in the metadata. The published CLM_v0.1-8B.pt is 75 MB and holds the two heads, not the 8B encoder; the export is 16 tensors and 18.9 M parameters.

The decision

softmax(exp(logit_scale) * cos(state_head(s), action_head(c)) / temperature) over a question's candidates. Both heads are inp → [LayerNorm →] hidden → out with GELU, and both projections are L2-normalised before the dot product.

The three types differ only after the distribution exists, so they share one path: choice is the argmax key, noul is the true entry of a two-candidate distribution, and score is the expected level index. confidence is the top probability minus the mean of the rest, clamped, and 1.0 for a single candidate — the definition src/clm/schema.py uses.

The scale cap

exp(logit_scale) is capped at 100, as heads.py caps it. The published logit_scale is 4.6132, whose exponential is 100.82, so the cap binds — without it every probability is about 0.8 % off, and more where candidates are close.

I had this wrong until a real encoder disagreed with me. The CPU-side oracle could not catch it: head_oracle.py was written from the same reading of the format and made the same mistake, so both sides agreed and the test was green. It took compare_with_reference.py (in the third PR) comparing against CLM's own heads over real Qwen3-8B embeddings.

Checks

With the checkpoint exported:

python recipe/clm/native/export_weights.py CLM_v0.1-8B.pt /tmp/clm-export
python recipe/clm/native/head_oracle.py /tmp/clm-export /tmp/clm-export/head-oracle.json
CLM_EXPORT=/tmp/clm-export cargo test -p omni-clm -- --include-ignored

5 tests: 16 tensor conversion hashes against the export oracle, the inventory against the head configuration, and five decisions against head_oracle.py — an independent NumPy implementation of the same arithmetic — on hash-derived embeddings. Default cargo test needs no checkpoint (2 passed, 3 ignored). fmt and clippy -D warnings clean.

CLM is the second model tracked in ThinkFlowLab#9. It decides differently in a way LAYA does
not cover: the engine does not compute embeddings. A frozen Qwen3-8B encoder runs
as its own process behind an OpenAI-compatible /v1/embeddings endpoint, and the
engine owns everything after it -- two projection heads, the cosine score, and the
typed answer. That split is the reason to implement it second.

A CLM checkpoint is a torch.save dict, so recipe/clm/native/export_weights.py
converts it first: tensors to safetensors with the head name as a prefix, and
cfg, hidden_size, projection_dim and logit_scale into the metadata. The published
checkpoint is 75 MB and holds the two heads, not the 8B encoder; the export is 16
tensors and 18.9 M parameters.

omni-clm reads that and computes decisions:

- config reads the head geometry from the metadata and checks it.
- weights lists the expected inventory from that geometry, checks shapes before
  reading, and loads FP32.
- scoring projects a state and each candidate, L2-normalises both, scores them by
  cosine under exp(logit_scale), softmaxes across the question's candidates, and
  assembles choice, noul and score. The three types differ only after the
  distribution exists, so they share one path. confidence is the top probability
  minus the mean of the rest, clamped, and 1.0 for a single candidate, which is
  the definition src/clm/schema.py uses.

No GPU is needed for any of it. This crate does not call an embeddings endpoint
and does not serve HTTP; those belong with the runtime that owns the request path.

Checks, with the checkpoint exported:

  python recipe/clm/native/export_weights.py CLM_v0.1-8B.pt /tmp/clm-export
  python recipe/clm/native/head_oracle.py /tmp/clm-export /tmp/clm-export/head-oracle.json
  CLM_EXPORT=/tmp/clm-export cargo test -p omni-clm -- --ignored

- every tensor's FP32 conversion hash against the export oracle (16 tensors);
- the inventory against the head configuration;
- five decisions against head_oracle.py, an independent NumPy implementation of
  the same arithmetic, on hash-derived embeddings. Both sides load the same
  exported weights, so the check is that two implementations of the same maths
  agree; the embeddings carry no meaning as model output. The erf coefficients
  are the same digits on both sides on purpose.

Default cargo test needs no checkpoint. workspace: fmt, clippy -D warnings,
15 tests, release build.
@xiaoyu-xyz xiaoyu-xyz changed the title [CLM] Add the second model engine: heads and typed decision scoring [CLM 1/2] Checkpoint and typed decision scoring Sep 28, 2026
@xiaoyu-xyz xiaoyu-xyz closed this Sep 28, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant