Repository navigation
[CLM 1/2] Checkpoint and typed decision scoring - #26
Closed
xiaoyu-xyz wants to merge 1 commit into
Closed
xiaoyu-xyz wants to merge 1 commit into
xiaoyu-xyz wants to merge 1 commit into
Conversation
CLM is the second model tracked in ThinkFlowLab#9. It decides differently in a way LAYA does not cover: the engine does not compute embeddings. A frozen Qwen3-8B encoder runs as its own process behind an OpenAI-compatible /v1/embeddings endpoint, and the engine owns everything after it -- two projection heads, the cosine score, and the typed answer. That split is the reason to implement it second. A CLM checkpoint is a torch.save dict, so recipe/clm/native/export_weights.py converts it first: tensors to safetensors with the head name as a prefix, and cfg, hidden_size, projection_dim and logit_scale into the metadata. The published checkpoint is 75 MB and holds the two heads, not the 8B encoder; the export is 16 tensors and 18.9 M parameters. omni-clm reads that and computes decisions: - config reads the head geometry from the metadata and checks it. - weights lists the expected inventory from that geometry, checks shapes before reading, and loads FP32. - scoring projects a state and each candidate, L2-normalises both, scores them by cosine under exp(logit_scale), softmaxes across the question's candidates, and assembles choice, noul and score. The three types differ only after the distribution exists, so they share one path. confidence is the top probability minus the mean of the rest, clamped, and 1.0 for a single candidate, which is the definition src/clm/schema.py uses. No GPU is needed for any of it. This crate does not call an embeddings endpoint and does not serve HTTP; those belong with the runtime that owns the request path. Checks, with the checkpoint exported: python recipe/clm/native/export_weights.py CLM_v0.1-8B.pt /tmp/clm-export python recipe/clm/native/head_oracle.py /tmp/clm-export /tmp/clm-export/head-oracle.json CLM_EXPORT=/tmp/clm-export cargo test -p omni-clm -- --ignored - every tensor's FP32 conversion hash against the export oracle (16 tensors); - the inventory against the head configuration; - five decisions against head_oracle.py, an independent NumPy implementation of the same arithmetic, on hash-derived embeddings. Both sides load the same exported weights, so the check is that two implementations of the same maths agree; the embeddings carry no meaning as model output. The erf coefficients are the same digits on both sides on purpose. Default cargo test needs no checkpoint. workspace: fmt, clippy -D warnings, 15 tests, release build.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
CLM is the second model tracked in #9, and it decides differently in a way LAYA does not cover: the engine does not compute embeddings. A frozen Qwen3-8B encoder runs as its own process, and the engine owns everything after it. This is the half that can be checked without a GPU — the checkpoint, the head geometry, and the decision arithmetic.
Core code: 483 lines (
config75,weights146,scoring256,lib6), within the 500-line budget. Tests and the export/oracle scripts are separate.This is the first of three, stacked:
clm-run, and the comparison against CLMThe checkpoint is converted first
A CLM checkpoint is a
torch.savedict, so it is a pickle and no non-Python reader can open it.recipe/clm/native/export_weights.pywrites the tensors to safetensors with the head name as a prefix, and keepscfg,hidden_size,projection_dimandlogit_scalein the metadata. The publishedCLM_v0.1-8B.ptis 75 MB and holds the two heads, not the 8B encoder; the export is 16 tensors and 18.9 M parameters.The decision
softmax(exp(logit_scale) * cos(state_head(s), action_head(c)) / temperature)over a question's candidates. Both heads areinp → [LayerNorm →] hidden → outwith GELU, and both projections are L2-normalised before the dot product.The three types differ only after the distribution exists, so they share one path:
choiceis the argmax key,noulis thetrueentry of a two-candidate distribution, andscoreis the expected level index.confidenceis the top probability minus the mean of the rest, clamped, and1.0for a single candidate — the definitionsrc/clm/schema.pyuses.The scale cap
exp(logit_scale)is capped at 100, asheads.pycaps it. The publishedlogit_scaleis 4.6132, whose exponential is 100.82, so the cap binds — without it every probability is about 0.8 % off, and more where candidates are close.I had this wrong until a real encoder disagreed with me. The CPU-side oracle could not catch it:
head_oracle.pywas written from the same reading of the format and made the same mistake, so both sides agreed and the test was green. It tookcompare_with_reference.py(in the third PR) comparing against CLM's own heads over real Qwen3-8B embeddings.Checks
With the checkpoint exported:
5 tests: 16 tensor conversion hashes against the export oracle, the inventory against the head configuration, and five decisions against
head_oracle.py— an independent NumPy implementation of the same arithmetic — on hash-derived embeddings. Defaultcargo testneeds no checkpoint (2 passed, 3 ignored).fmtandclippy -D warningsclean.