Repository navigation
Conversation
This was referenced Oct 6, 2026
Open-Jev-9B uses the same method, prompt format and head design as Open-Jev-27B-v1.1. Select the pinned revisions, request alias and backbone dimensions by the export's model_id so one worker serves either checkpoint, and key the export script's pins and head width the same way. Add the 9B recipe steps and its reference comparison. 9B runs one forward pass per candidate: candidate packing is validated on 27B only.
Contributor
Author
|
Rebased onto main after #102. 9B keeps one forward pass per candidate, as you suggested there; 27B packs as before. The 9B answers are unchanged: 253 of 253 identical to the previous head apart from timing. |
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Purpose
First preparation PR for #85, following the plan. Open-Jev-9B uses the same method, prompt format and head design as Open-Jev-27B-v1.1, so the existing native worker can serve it once the pins are selected by the export's
model_id. The prefix-reuse work can then be developed and checked on a model that fits a 48 GB card, while 27B keeps working as before.contract.rs: a table of the two pinned checkpoints (base and checkpoint revisions, request alias, backbone dimensions, and whether candidates are packed). The export manifest selects the entry, and the request model check, the response and/healthmodel andmetadata.base_revisioncome from it. A 9B worker acceptsQwen/Qwen3.5-9B,open-jev,jev-latestandopen-jev-9b; each worker rejects the other checkpoint's names with 422.executor.rs: the head loader checks the selected checkpoint's backbone dimensions. Candidate packing from perf(open-jev): pack candidate projections within requests #102 stays on for 27B; 9B runs one forward pass per candidate, as discussed in perf(open-jev): pack candidate projections within requests #102, until packing is validated for it.export_merged.py: the pins and the head width (5120 or 4096) followmodel.json'smodel_id. The temperature still comes from the package'stemperature.json, and the existing output, max-length and non-finite head checks are unchanged.recipe/open_jev/validation-9b.md, "27B" added to the H200 validation's title so it isn't read as covering 9B, and the packing notes in the recipe, the model README anddocs/architecture.md, which now name 27B.The shared Qwen executor, the CUDA backend, the runtime and the request and response format are unchanged. For 27B, the accepted names, pins, response model, metadata and candidate packing are the same as on
main. The error messages change slightly: the single export-manifest message becomes three (format, unsupportedmodel_id, revision mismatch), and the backbone message now namesQwen/Qwen3.8-27B.Test Plan
System1-Omni Version / Commit:
f731183onmain9986574, rebased after #102. The GPU results below ran at1e789ebon7f39ac4; after the rebase, the 9B worker's 253 responses are identical to that build's apart from timing (see the end of the results).CPU, from the repository root:
cargo fmt --all --checkcargo clippy --workspace --locked --all-targets -- -D warningscargo test -p omni-jev --test frontend --lockedcargo test --workspace --lockedcargo build --workspace --release --lockedmkdocs build --strictwithdocs/requirements.txtNew CPU tests cover:
model_id, wrong format);config.jsonatc202236);GPU, one RTX 6000 Ada (sm_89), Open-Jev-9B only:
Test Result
CPU. All six commands pass on macOS (Rust 1.97.0). On Linux (Ubuntu 22.04, Rust 1.97.0) the workspace tests and release build pass; that toolchain has no rustfmt or clippy, so those two ran on macOS only.
Export. This PR's script produced the same bytes, in all 10 files, as the export used for the plan's measurements. The CPU export peaked at about 20 GB RSS and wrote 15.9 GB. An unknown base (also with a null revision), a 9B
model_idwith 27B's revision, and--max-length 0each stop with aValueErrorbefore writing anything.Worker startup errors. A 27B manifest over the 9B weights stops with
expected the Qwen/Qwen3.8-27B backbone dimensions, and a 9B manifest with 27B's checkpoint revision withexpected a pinned Open-Jev-9B export, both before CUDA loading.Tokenization. The ignored test passes against the 9B export: the existing fixture's token IDs also match 9B's tokenizer. The test now swaps the fixture's
open-jev-27b-v1.1model name for the loaded checkpoint's alias.Worker. Ready after the real warmup, with
/healthreturning{"status":"ready","model":"Qwen/Qwen3.5-9B"}.Qwen/Qwen3.5-9B,open-jev-9b,open-jevandjev-latestreturn 200;open-jev-27b-v1.1andQwen/Qwen3.8-27Breturn 422. The recipe's example request gives the same answers through the frontend as directly. Device memory: 15,874 MiB after warmup, 16,672 MiB after all requests.M1 to M5 (253 requests). Answers and usage are identical to the build measured for the plan, so the plan's numbers hold for this code. Against the reference (
3308a15, uncached, 9B's head and temperature), with the same input token counts everywhere:noul(74)The two changed decisions are close calls on both sides; details are in
validation-9b.md.Readout hidden state, native against reference (2,157 candidates). Relative L2 difference: median 0.0096, p99 0.023, largest 0.063. Head output (the scalar before temperature, from each side's hidden state through the same FP64 head): median 0.028, p99 0.21, largest 1.49. The two sides differ in LoRA handling (merged BF16 weights vs unmerged PEFT), Gated DeltaNet path and batching.
For #85, this places the plan's reuse gates. 1e-2 relative L2 is about the median of the native vs reference spread (903 of 2,157 candidates are above it), and 0.035 on the head output is near its median too (905 above). Both gates stay as declared; they compare reuse with the full forward on one build.
After the rebase onto
9986574. On the same GPU, the 9B worker atf731183, with its own library, answered all 253 requests byte-identically to the1e789ebbuild apart from timing, and the six qwen3_5 kernel tests pass.Not verified: 27B. It needs more than 48 GB. The 27B entry carries
main's values unchanged, packing included, and the CPU tests cover its selection, names and backbone check. A 27B startup on the H200 would confirm it.Demo / evidence
recipe/open_jev/validation-9b.mdhas the aggregate results, workload and controls. JevBench's license allows publishing aggregates only, so its requests and responses are not included; the M1 to M3 manifests, raw responses, hidden-state dumps and scripts are kept locally. No demo: the PR adds a checkpoint without new behavior to show.Self-review