[CLM 2/2] Reach the encoder, answer requests, and check against CLM - #29
xiaoyu-xyz wants to merge 14 commits into
Conversation
91022a7 to
fd841e1
Compare
hsliuustc0106
left a comment
There was a problem hiding this comment.
Reviewed commit fd841e1967e0a58b0a1af7b8425e83579f1c576c.
- [P2] Preserve empty-container candidate text. src/models/clm/src/serve.rs:158 falls back to the key whenever to_text is empty; upstream falls back only for null or the empty string. Criteria {"a":{},"b":[]} therefore become ["a","b"] here instead of ["", ""], changing embeddings/probabilities. The noul fallback has the same problem. Reproduced with a CPU Rust harness and compared with upstream schema.py.
- [P2] Include the score legend in serialized answers. src/models/clm/src/serve.rs:355 omits the legend that upstream answer_from_probs emits for every score question. Criterion descriptions are lost, so consumers cannot recover the score-level mapping from the response. The current answer type discards these descriptions before serialization.
- [P2] Pass --temperature to the native comparison process. recipe/clm/native/compare_with_reference.py:83 omits this argument when launching clm-run, while line 98 uses args.temperature for the reference. Any nondefault temperature compares different configurations and can falsely report a port regression.
Tests: 2 passed, 4 ignored without external checkpoint/text oracles. Inherits #28 loader finding.
fd841e1 to
d276bf8
Compare
|
Fixed in d276bf8, all three. (Rebased on the new #28, so the SHAs moved.)
|
Levius-Fubuki
left a comment
There was a problem hiding this comment.
Reviewed d276bf8a038abb42682aa24bceb1d0ca338ec183, including its incremental changes on #28 and the inherited config/scoring code.
The original empty-container, score-legend, and temperature-forwarding repairs are present. Fresh Linux/Rust 1.98.1 formatting, strict all-target Clippy, workspace tests and release build passed (17 tests in the default workspace run). All eight CLM tests subsequently passed with real head weights and external oracles: 16 FP32 tensor hashes, five head decisions, six state renderings and 42 state/question pairs. The text oracle imports official clm.schema pinned to bb42c6c5bf914fd449bed2f6ca65be80602cb1f7.
A separate end-to-end CPU comparison used the real head weights, official PyTorch heads/schema, one deterministic HTTP embedding stub, and --temperature 2.5: all four cases passed on this head at tolerance 1e-4. The previous comparison script failed all four against the same setup because it omitted the native temperature argument (maximum probability difference about 0.177).
The published head checkpoint at e939398d4556fcd9400c76fa8c5a513202f42b0a was SHA-256 verified. The head oracle uses NumPy and is not itself independent PyTorch numerical certification. The targeted protocol/text findings below were reproduced separately and are not covered by the passing 42-case oracle. This PR also inherits the #28 findings, and the stack needs its existing main conflicts resolved. No GPU or real Qwen/vLLM inference was run.
Changes requested for the two reference-compatibility issues below.
| let body = serde_json::json!({ | ||
| "model": self.model, | ||
| "input": texts, | ||
| "encoding_format": "base64", |
There was a problem hiding this comment.
[P2] Preserve the reference client's prompt-truncation setting
The upstream Embedder sends truncate_prompt_tokens=2048 by default, matching its documented vLLM --max-model-len 2048 deployment, but this request body omits that setting. Long state/candidate text therefore uses different server behavior: an over-limit request that the reference truncates can be rejected by the native path. I captured both real clients against a local endpoint enforcing that wire contract: upstream included the field and succeeded; HttpEncoder omitted it and received 400 for the same long input. This was a protocol stub, not a real vLLM run. Add a configurable truncation limit with the reference default and serialize it in the request. Reference: https://github.com/Contrastive-LM/CLM/blob/bb42c6c5bf914fd449bed2f6ca65be80602cb1f7/src/clm/embedder.py#L42 .
| Value::String(s) => s.clone(), | ||
| Value::Bool(true) => "true".to_string(), | ||
| Value::Bool(false) => "false".to_string(), | ||
| Value::Number(n) => n.to_string(), |
There was a problem hiding this comment.
[P2] Match the reference's number formatting before embedding
JSON-number to_string() is not equivalent to Python str(float), which upstream uses in to_text. For the same JSON value 1e-5, this function produces 0.00001 while the reference produces 1e-05; for 1e-7, it produces 1e-7 versus 1e-07. These values are valid in structured state or criteria, so the two engines embed different strings and can make different decisions even with identical model weights. The current oracle only exercises 0.5 and misses exponent-format cases. Match Python's numeric rendering for supported JSON values and add small/large exponent cases to the upstream-derived oracle.
d276bf8 to
df2da33
Compare
|
Fixed in df2da33; the stack is rebased on
The text oracle gained an exponent rendering (now 7 renderings, 49 pairs), and the byte comparison fails without the fix. Strict Clippy, workspace tests and the release build pass locally. |
df2da33 to
968b61b
Compare
|
Ran the real-encoder comparison on the current head, Qwen3-8B bf16 served behind
The first row reproduces the numbers already in the description to the last digit but one: The third row is the |
|
Rechecked |
|
Fixed in 44a55b3, and you were right that formatting was the wrong place to look — the parse was wrong.
Coverage now runs the full path rather than the formatter:
Removing the feature turns all three red, including 12 tests pass with the oracles, strict Clippy and the workspace build are clean on the head. |
|
Follow-up on the layer you named: the earlier fix was checked at Captured from a real Both match That is now pinned without a server: 13 tests pass with the oracles on the head. |
|
Rechecked |
hsliuustc0106
left a comment
There was a problem hiding this comment.
Independent local review — CLM 2/2
Verdict: changes requested — same convention finding as #28; the engineering is verified.
[Medium] Test bodies in src/models/clm/tests/{checkpoint,text}.rs — crate-local tests/ under src/, per CONTRIBUTING (at your base and on current main) they belong in repo-root tests/clm/. Same fix as #28.
Verified: clm-run is fully synchronous (no async runtime), so the blocking reqwest encoder client is the correct shape; serde_json carries both preserve_order and float_roundtrip as claimed; Engine::decide keeps state/candidate text construction, the noul "<key>: <text>" fallback, and answer assembly model-owned. Locally at head b9907af3: 9 passed / 4 ignored (oracle/checkpoint-gated, disclosed), clippy clean. The byte-for-byte text-oracle discipline (importing the reference clm.schema rather than transcribing it), 3e-06 frozen-vector agreement, and the temperature stub-endpoint check are author-reported with recorded evidence.
Architecture note for the future worker: when CLM grows an HTTP worker, adopt the prepare → execute → finish split and omni-runtime admission directly (per docs/architecture.md and #78/#80) — the library's clean separation makes that easy. #28 merges first.
Local reviewer report per the repo review skill; reflects head b9907af3 only.
6a40e0e to
a44e269
Compare
|
The float cases are fixed, but the integer one is real and its natural fix is blocked — I would rather show you the block than pick for you.
But a Cargo feature is not scoped to the crate that asks for it. So I reverted it. The branch is green again: Two ways out, and the choice is not only mine:
I lean to (2): it removes a coupling rather than adding a second JSON implementation. But it edits a crate this PR does not otherwise touch, so I did not do it unasked. The limitation as it stands: an integer outside |
|
Fixed in Why not What it does instead. The fields that get rendered — Two things had to be right, and both have tests:
Removing the text turns three red at One thing to flag: this costs 169 core lines in |
CLM is the second model tracked in ThinkFlowLab#9, and it decides differently in a way LAYA does not cover: the engine does not compute embeddings. A frozen Qwen3-8B encoder runs as its own process, and the engine owns everything after it. This is the half that can be checked without a GPU -- the checkpoint, the head geometry, and the decision arithmetic. A CLM checkpoint is a torch.save dict, so recipe/clm/native/export_weights.py converts it first: tensors to safetensors with the head name as a prefix, and cfg, hidden_size, projection_dim and logit_scale into the metadata. The published checkpoint is 75 MB and holds the two heads, not the 8B encoder; the export is 16 tensors and 18.9 M parameters. - config reads the head geometry from the metadata and checks it. - weights lists the expected inventory from that geometry, checks shapes before reading, and loads FP32. - scoring projects a state and each candidate, L2-normalises both, scores them by cosine under exp(logit_scale), softmaxes across the question's candidates, and assembles choice, noul and score. The three types differ only after the distribution exists, so they share one path. confidence is the top probability minus the mean of the rest, clamped, and 1.0 for a single candidate, which is the definition src/clm/schema.py uses. exp(logit_scale) is capped at 100, as the reference caps it in heads.py. The published logit_scale is 4.6132, whose exponential is 100.82, so the cap binds: without it every probability is about 0.8 % off. Checks. head_oracle.py is an independent NumPy implementation of the same arithmetic, and tests/checkpoint.rs checks 16 tensor conversion hashes against the export oracle and five decisions against head_oracle.py. Default cargo test needs no checkpoint. fmt, clippy -D warnings.
`head_tensors` lists `norms.N.*` only when `cfg.layernorm` is set, so a `layernorm: false` checkpoint is complete without them — but `load_head` read them unconditionally. A valid checkpoint passed the inventory check and then failed to load. The new test writes both shapes plus one that declares LayerNorm and omits the tensors, and checks the loader accepts the first two and rejects the third. Found in review.
…labels - `Config::from_metadata` applied the top-level `hidden_size`/`projection_dim` fallback *after* deserializing `cfg`, but both are required fields, so a `cfg` that omits them failed to parse and the fallback never ran. The geometry is now filled in before deserialization, and a `cfg` that omits it loads while one that omits it with no top-level copy is still an error. - `Answer::label` used `max_by`, which keeps the *last* of several equal values. `schema.label_of` is `max(p, key=p.__getitem__)`, which keeps the first, so a tie in a score distribution named a different level. Found in review.
CONTRIBUTING puts every test body in the repository-level `tests/` tree and asks for explicit `[[test]]` registration, which is what `src/models/laya` does. `src/models/clm/tests/` was a crate-local directory under `src/`, which that rule prohibits; the tests ran either way, so this is placement rather than breakage. Found in review.
The engine does not compute embeddings: a frozen Qwen3-8B encoder runs as its own process
behind an OpenAI-compatible /v1/embeddings endpoint, and this is the client to it.
An Encoder trait with two implementations. HttpEncoder reproduces the reference client --
a POST of {model, input, encoding_format}, a base64 f32 payload per input in index order,
L2-normalised on receipt, batched -- which is what `vllm serve --runner pooling` exposes.
HashingEncoder derives a vector from a SHA-256 of the text, exactly as
recipe/clm/native/head_oracle.py does, so the decision path runs with no GPU and no
weights and can still be compared against the Python oracle.
fmt, clippy -D warnings, builds on the head loader.
The crate could load heads and score candidates. This adds the client to the encoder and
the path from a request to typed answers, plus the harness that compares the result with
CLM's own implementation.
Core code: 489 lines (embedding 178, serve 311). clm-run and the recipe scripts are model
tooling, like laya-run and export_weights.py, so they are not counted.
embedding carries an Encoder trait with two implementations. HttpEncoder reproduces the
reference client -- a POST of {model, input, encoding_format}, a base64 f32 payload per
input in index order, L2-normalised on receipt, batched -- which is what
`vllm serve --runner pooling` exposes. HashingEncoder derives a vector from a SHA-256 of
the text, exactly as head_oracle.py does, so the decision path runs with no GPU and no
weights and can still be compared against the Python oracle.
serve is the request path. Two text functions decide what reaches the encoder and both are
reproduced exactly, because a separator in the wrong place shifts every probability
without failing:
- the state head sees the context and the question joined by a blank line, so a question
belongs in `instructions` and not repeated in the state;
- the action head sees each candidate's own text with nothing prefixed for `choice`, but
"<key>: <text>" for `noul`, falling back to "Yes. This is true: <question>".
to_text renders a state, description or criteria as the prose the heads were trained on:
objects become `key: value`, top-level fields separated by a blank line and nested ones
indented, arrays one `- item` line each, empty containers taking the `key: value` branch.
serde_json gains `preserve_order`, which is a correctness fix rather than a preference:
the default BTreeMap sorts keys, and both to_text and the reference preserve the caller's
field order. Without it a state whose fields are not in alphabetical order reaches the
encoder re-sorted, which is a different text and a different decision.
clm-run takes a converted checkpoint and an embeddings endpoint and answers one request per
line, in the shape the other engines use, so the same binary drives a stub encoder on a
laptop and a real one on a GPU box.
Checks. text_oracle.py writes the reference's own output for to_text, state_text and
candidates -- 6 and 30 cases -- by importing clm.schema, so tests/text.rs compares against
the implementation rather than a transcription of it, byte for byte.
transformers_encoder.py serves Qwen3-8B behind the endpoint the deployment uses -- vLLM is
not needed to verify the client -- and compare_with_reference.py compares omni-clm with
CLM's own heads and schema over the same vectors. That comparison is what showed the
exp(logit_scale) cap in the first PR: with the uncapped 100.82 every probability was about
0.8 % off, and no CPU-side oracle can see it, because both sides of those share the
constant. All four cases agree within 2e-3 for the two-candidate question.
8 tests, fmt, clippy -D warnings.
…imposes
All four cases now pass against a real encoder, and the script says why the number is what
it is rather than leaving the next reader to rediscover it.
CLM scores with exp(logit_scale) = 100, so a difference of d in a cosine similarity becomes
100 d in the logit. Two encoder forward passes that disagree by about 1e-4 in the cosine --
which is what vLLM's bf16 pooling and Transformers' forward produce for Qwen3-8B -- therefore
disagree by about 1e-2 in the probabilities. That is a property of how this model scores, not
a defect on either side, and the old 2e-3 default was asking for agreement the pipeline
cannot give.
The engine's own arithmetic is much tighter and is checked separately: hold the vectors
fixed -- serve both sides one frozen set -- and the implementations agree to 3e-06.
PASS choice_two max|dp|=1.72e-05 choice=billing/billing
PASS choice_five max|dp|=1.97e-02 choice=a/a
PASS score_three score=0.385165/0.403241
PASS noul_stmt 0.631038 vs 0.646805
…ature
Three findings, all measured against the reference:
- `candidates` fell back to the option key whenever a rendering came out
empty. `schema.candidates` falls back only for null and the empty string,
so `{"a": {}, "b": []}` is `["", ""]`, not `["a", "b"]`. The same test
applies to the `noul` branch. `text_oracle.py` now covers both shapes.
- A `score` answer dropped the level legend that `answer_from_probs`
returns and that `client.ScoreAnswer` declares, so a consumer could not
read a score back as a level. The legend is the keys zipped with the
candidate texts, which is what the reference builds it from.
- `compare_with_reference.py` launched `clm-run` at its default
temperature while running the reference at `--temperature`, so any
non-default compared two different configurations.
Found in review.
- `HttpEncoder` omitted `truncate_prompt_tokens`, which `embedder.py` sends as its `max_tokens` — 2048 by default, matching the deployment's `--max-model-len`. A text past the limit is truncated by the reference but rejected by a server asked to embed it whole. The limit is now a field with the reference default, it is serialized, and `clm-run` exposes it as `--max-tokens N|none`. - `to_text` rendered numbers with `serde_json`, which disagrees with Python's `str(float)`: `1e-5` came out `0.00001` where the reference writes `1e-05`, and an integral float lost its `.0`. The heads are trained on these strings, so the two engines embedded different text for the same state. `python_float` reproduces CPython's rendering; it was fuzzed against `repr` over 41k doubles covering every decimal exponent. Found in review.
`serde_json`'s default float parsing is not correctly rounded, so a literal with more digits than a double holds can land on the neighbouring double: `7.8190461323667115` parsed one ULP high and rendered `7.819046132366712`, and `9007199254740993.0` became `9007199254740994.0` where `json.loads` gives `9007199254740992.0`. Getting the formatting of an already-parsed float right did not help — the parse was wrong, so the heads were shown a different number. `serde_json`'s `float_roundtrip` feature makes the parse bit-identical to `json.loads`, and pulls in no new dependency. The text oracle gained a state with 17 significant digits, which the byte comparison catches through both `to_text` and `candidates`, and a test pins the parsed bit patterns directly. Found in review.
The parsing fix was checked at `to_text` and `candidates`. The layer the bug actually lived at is one step earlier: a request line goes through `Request::parse` and `prepare` before either. The new test takes a JSON line as `clm-run` reads it and asserts the exact state text and candidate texts that reach the encoder -- the two strings also confirmed against the reference over a real HTTP request, by capturing what `clm-run` sends. Without `float_roundtrip` it fails at `left: "seventeen: 7.819046132366712..."` against `right: "seventeen: 7.8190461323667115..."`.
The same rule `src/models/clm/tests/` broke in ThinkFlowLab#28 applies here, and to the inline test in `embedding.rs`: CONTRIBUTING keeps test bodies in the repository-level `tests/` tree, and for a private item only the `#[cfg(test)]` and the module path may stay in `src/`. `HttpEncoder::body` is private, so that is the `#[path]` form, with the body in `tests/clm/embedding.rs`. Found in review.
`serde_json` cannot hand back what an integer literal said: one that does not fit
an `i64` or a `u64` is parsed straight to a double, so `18446744073709551616`
reached the encoder as `1.8446744073709552e+19`. `json.loads` gives Python an
arbitrary-precision `int` instead and `str` prints every digit.
`arbitrary_precision` would fix it but is a Cargo feature, so it would reach
every crate in the workspace — it breaks `omni-cua-s1-native`'s
`json::tests::dumps_matches_python`, which serialises `Value::Number`. This
carries the text instead: the fields that get rendered are deserialized a second
time as `RawValue`, and each is rendered from its own text with its own literals.
`raw_value` adds `RawValue` without changing `Number`, and the workspace stays
green.
Two things had to be got right, and both are tested:
- Rendering the same field twice must not shift which literal belongs to which
number. The state is rendered once now rather than once per question.
- Nothing may depend on the order the request's keys appear in. Each field is
lexed from its own text, so `{"questions": …, "state": …}` is fine, and
`the_literals_line_up_with_the_values` pins the lexer against a value walk.
`to_text` keeps working on a `Value` and falls back to `serde_json`'s integer
types; `to_text_json` and `Request::parse_line` are the paths that keep the
digits, and `clm-run` uses the latter. The oracle gained a rendering with 2**64,
2**128 and a negative of each, and the case loop now goes through a real request
line rather than calling `candidates` on a `Value`.
Removing the text turns three red at `left: "…big: 1.8446744073709552e+19…"`
against `right: "…big: 18446744073709551616…"`.
Found in review.
|
Rechecked |
The previous commit kept an integer literal's digits by lexing each rendered field
and pairing the literals with the numbers a depth-first walk meets. That pairing is
positional, and it is only right when the walk reads the values in the order they
were written. A `noul` does not: the reference reads its two descriptions in
`NOUL_KEYS` order, `false` before `true`, so
"criteria": {"true": 1, "false": 2}
reached the action head as `false: 1` / `true: 2`, where `clm.schema` produces
`false: 2` / `true: 1`. A wrong candidate text is not a wrong spelling, it is a
different answer.
The descriptions of a `choice` or a `noul` are now found by the key they were
written under — `HashMap<String, &RawValue>` over the criteria text — and each is
rendered from its own literal, so the order the keys are read in cannot matter. A
`score`'s levels are a list, which has one order, and keep the cursor.
The default `noul` candidates are the question's own statement, and they were built
from the parsed `instructions`: a statement that is a number lost its digits there
even though `state_text` had them,
"instructions": 18446744073709551616
giving `false: No. This is false: 1.8446744073709552e+19`. The statement rendered
from its text is now passed down and reused, so both heads see one text.
The oracle gained three whole request lines, since a `Value` cannot hold an integer
above `u64::MAX` and `QuestionRequest` holds `instructions` as a `String`, and
`raw_requests_match_the_reference_byte_for_byte` runs them through
`Request::parse_line`. Against the previous commit it fails at
`left: ["false: 1", "true: 2"]` where the reference says `["false: 2", "true: 1"]`.
Found in review.
1a15d44 to
a2019f9
Compare
|
Both findings were real, and the first one is worse than a spelling difference. Fixed in The literal now follows its key. You are right that the pairing was positional. The The default candidates reuse the rendered statement. They were built from the parsed Tests. A which is your case. Two unit tests cover the same two things without the oracle. This was not cosmetic.
Checks on the rebased head: |
|
DO WE HAVE AN INTEGRATION example with system1-agents with a video demo? |
Second of two, stacked on #28. Merge #28 first — this branch is stacked on it, so until then the diff below shows both PRs and after #28 merges it is only this one's. Reviewing commit by commit is cleanest in the meantime; the last nine commits are this PR's.
Core code: 681 lines (
embedding142,serve539), plus the 15 lines it adds toscoringfor thelegend, counting lines that are neither blank nor start with//.servewas 350 by that count when this PR was opened, and the literal handling below is 189 of the lines since.clm-run(92 lines) and the recipe scripts are model tooling, likelaya-runandexport_weights.py, so they are not counted.The embeddings client
The engine does not compute embeddings, so this is the client to the process that does. An
Encodertrait with two implementations:HttpEncoderreproduces the reference client — a POST of{model, input, encoding_format, truncate_prompt_tokens}, a base64f32payload per input in index order, L2-normalised on receipt, batched. That is whatvllm serve --runner poolingexposes. The truncation limit defaults to the reference's 2048, which is the deployment's--max-model-len; a text past it is truncated there but rejected by a server asked to embed it whole.HashingEncoderderives a vector from a SHA-256 of the text, exactly ashead_oracle.pydoes, so the whole decision path runs with no GPU and no weights and can still be compared against the Python oracle.The request path
Two text functions decide what reaches the encoder and both are reproduced exactly, because a separator in the wrong place shifts every probability without failing anything:
instructionsand not repeated in the state;choice, but"<key>: <text>"fornoul, falling back to"Yes. This is true: <question>".to_textrenders a state, description or criteria as the prose the heads were trained on: objects becomekey: value, top-level fields separated by a blank line and nested ones indented, arrays one- itemline each, empty containers taking thekey: valuebranch.Numbers are parsed and spelled the way the reference does:
serde_json's default float parsing is not correctly rounded, so a literal holding more digits than a double does could land one ULP away and render as a different number —7.8190461323667115came out7.819046132366712. Thefloat_roundtripfeature makes the parse bit-identical tojson.loads, and is what keeps the following true: numbers are spelled the waystr(float)spells them —1e-05,1e-07, and a.0kept on an integral float.serde_jsondisagrees on both ends, and since the heads are trained on these strings a differently spelled number is a different input. Rust's shortest digits also differ from CPython's on the last digit when a value sits exactly between two equally short decimals, so the digits are re-rendered at the shortest round-tripping precision; fuzzed againstreprover 41k doubles, 0 differences.serde_jsongainspreserve_order, which is a correctness fix rather than a preference: the defaultBTreeMapsorts keys, and bothto_textand the reference preserve the caller's field order. Without it a state whose fields are not in alphabetical order reaches the encoder re-sorted, which is a different text and a different decision. The oracle caught exactly that.clm-runtakes a converted checkpoint and an embeddings endpoint and answers one request per line, in the shape the other engines use, so the same binary drives a stub encoder on a laptop and a real one on a GPU box.Checks
Text construction, byte for byte.
text_oracle.pywrites the reference's own output forto_text,state_textandcandidates— 9 and 72 cases, plus three whole request lines that aValuecannot express — by importingclm.schema, sotests/clm/text.rscompares against the implementation rather than a transcription of it.Against a real encoder.
transformers_encoder.pyserves Qwen3-8B behind the exact endpoint the deployment uses, so vLLM is not needed to verify the client, andcompare_with_reference.pycomparesomni-clmwith CLM's own heads and schema over the same vectors:That comparison is what established the
exp(logit_scale)cap in #28 — with the uncapped 100.82 every probability was about 0.8 % off, and no CPU-side oracle can see it, because both sides of those share the constant.On the tolerance. CLM scores with
exp(logit_scale) = 100, so a difference ofdin a cosine becomes100 din the logit. Two encoder forward passes that disagree by ~1e-4 in the cosine — which is what vLLM's bf16 pooling and Transformers' forward produce for Qwen3-8B — therefore disagree by ~1e-2 in the probabilities. That is a property of how this model scores, not a defect on either side.The engine's own arithmetic is checked separately and much more tightly: serve both sides one frozen set of vectors and the implementations agree to 3e-06. That is the number that says the port is faithful; this script is a smoke test over a real encoder.
--temperatureis checked the same way against a stub endpoint, where both sides get identical vectors and any difference is the two implementations': at 2.5 the comparison is 4/4 PASS withmax|dp|around 1e-6, and with the argument not passed on toclm-runit is 4/4 FAIL at around 0.18.15 tests with the oracles, 11 passed / 4 ignored without them. The CI job's commands pass locally:
cargo fmt --all --check,cargo clippy --workspace --locked --all-targets -- -D warnings,cargo test --workspace --locked,cargo build --workspace --release --locked.