Skip to content

open_jev: support the Open-Jev-9B checkpoint - #92

Open
twu3202 wants to merge 1 commit into
ThinkFlowLab:mainfrom
twu3202:open-jev-9b
Open

twu3202 wants to merge 1 commit into
ThinkFlowLab:mainfrom
twu3202:open-jev-9b

Conversation

@twu3202

@twu3202 twu3202 commented Oct 5, 2026 •

Copy link
Copy Markdown
Contributor

Purpose

First preparation PR for #85, following the plan. Open-Jev-9B uses the same method, prompt format and head design as Open-Jev-27B-v1.1, so the existing native worker can serve it once the pins are selected by the export's model_id. The prefix-reuse work can then be developed and checked on a model that fits a 48 GB card, while 27B keeps working as before.

  • contract.rs: a table of the two pinned checkpoints (base and checkpoint revisions, request alias, backbone dimensions, and whether candidates are packed). The export manifest selects the entry, and the request model check, the response and /health model and metadata.base_revision come from it. A 9B worker accepts Qwen/Qwen3.5-9B, open-jev, jev-latest and open-jev-9b; each worker rejects the other checkpoint's names with 422.
  • executor.rs: the head loader checks the selected checkpoint's backbone dimensions. Candidate packing from perf(open-jev): pack candidate projections within requests #102 stays on for 27B; 9B runs one forward pass per candidate, as discussed in perf(open-jev): pack candidate projections within requests #102, until packing is validated for it.
  • export_merged.py: the pins and the head width (5120 or 4096) follow model.json's model_id. The temperature still comes from the package's temperature.json, and the existing output, max-length and non-finite head checks are unchanged.
  • Docs: the native recipe (9B download, export, serving, model names and the tokenizer test command), the model README, the root README and supported-models table, the recipe index, a new recipe/open_jev/validation-9b.md, "27B" added to the H200 validation's title so it isn't read as covering 9B, and the packing notes in the recipe, the model README and docs/architecture.md, which now name 27B.

The shared Qwen executor, the CUDA backend, the runtime and the request and response format are unchanged. For 27B, the accepted names, pins, response model, metadata and candidate packing are the same as on main. The error messages change slightly: the single export-manifest message becomes three (format, unsupported model_id, revision mismatch), and the backbone message now names Qwen/Qwen3.8-27B.

Test Plan

System1-Omni Version / Commit: f731183 on main 9986574, rebased after #102. The GPU results below ran at 1e789eb on 7f39ac4; after the rebase, the 9B worker's 253 responses are identical to that build's apart from timing (see the end of the results).

CPU, from the repository root:

  • cargo fmt --all --check
  • cargo clippy --workspace --locked --all-targets -- -D warnings
  • cargo test -p omni-jev --test frontend --locked
  • cargo test --workspace --locked
  • cargo build --workspace --release --locked
  • mkdocs build --strict with docs/requirements.txt

New CPU tests cover:

  • export selection for both checkpoints, and its rejections (wrong base or checkpoint revision, unknown model_id, wrong format);
  • accepted and rejected request model names for each checkpoint;
  • the response model and base revision of a 9B processor;
  • the head loader's backbone check against the existing Qwen3.8-27B config fixture and a new Qwen3.5-9B one (config.json at c202236);
  • which checkpoints pack candidates.

GPU, one RTX 6000 Ada (sm_89), Open-Jev-9B only:

  • export with this PR's script, compared with the export used for the plan's measurements;
  • export and worker startup error paths;
  • the ignored tokenization test against the 9B export;
  • worker health, request model names, and the recipe's example request directly and through the frontend;
  • the plan's M1 to M5 requests: responses against the build measured for the plan, and against the reference;
  • the readout hidden state, native against reference.

Test Result

CPU. All six commands pass on macOS (Rust 1.97.0). On Linux (Ubuntu 22.04, Rust 1.97.0) the workspace tests and release build pass; that toolchain has no rustfmt or clippy, so those two ran on macOS only.

Export. This PR's script produced the same bytes, in all 10 files, as the export used for the plan's measurements. The CPU export peaked at about 20 GB RSS and wrote 15.9 GB. An unknown base (also with a null revision), a 9B model_id with 27B's revision, and --max-length 0 each stop with a ValueError before writing anything.

Worker startup errors. A 27B manifest over the 9B weights stops with expected the Qwen/Qwen3.8-27B backbone dimensions, and a 9B manifest with 27B's checkpoint revision with expected a pinned Open-Jev-9B export, both before CUDA loading.

Tokenization. The ignored test passes against the 9B export: the existing fixture's token IDs also match 9B's tokenizer. The test now swaps the fixture's open-jev-27b-v1.1 model name for the loaded checkpoint's alias.

Worker. Ready after the real warmup, with /health returning {"status":"ready","model":"Qwen/Qwen3.5-9B"}. Qwen/Qwen3.5-9B, open-jev-9b, open-jev and jev-latest return 200; open-jev-27b-v1.1 and Qwen/Qwen3.8-27B return 422. The recipe's example request gives the same answers through the frontend as directly. Device memory: 15,874 MiB after warmup, 16,672 MiB after all requests.

M1 to M5 (253 requests). Answers and usage are identical to the build measured for the plan, so the plan's numbers hold for this code. Against the reference (3308a15, uncached, 9B's head and temperature), with the same input token counts everywhere:

Requests Questions Same decision Largest probability difference Above 0.01
M1 to M3 (22) 87 87 0.041 5
M4, JevBench noul (74) 74 74 0.068 4
M5, other JevBench tasks (157) 157 155 0.104 25

The two changed decisions are close calls on both sides; details are in validation-9b.md.

Readout hidden state, native against reference (2,157 candidates). Relative L2 difference: median 0.0096, p99 0.023, largest 0.063. Head output (the scalar before temperature, from each side's hidden state through the same FP64 head): median 0.028, p99 0.21, largest 1.49. The two sides differ in LoRA handling (merged BF16 weights vs unmerged PEFT), Gated DeltaNet path and batching.

For #85, this places the plan's reuse gates. 1e-2 relative L2 is about the median of the native vs reference spread (903 of 2,157 candidates are above it), and 0.035 on the head output is near its median too (905 above). Both gates stay as declared; they compare reuse with the full forward on one build.

After the rebase onto 9986574. On the same GPU, the 9B worker at f731183, with its own library, answered all 253 requests byte-identically to the 1e789eb build apart from timing, and the six qwen3_5 kernel tests pass.

Not verified: 27B. It needs more than 48 GB. The 27B entry carries main's values unchanged, packing included, and the CPU tests cover its selection, names and backbone check. A 27B startup on the H200 would confirm it.

Demo / evidence

recipe/open_jev/validation-9b.md has the aggregate results, workload and controls. JevBench's license allows publishing aggregates only, so its requests and responses are not included; the M1 to M3 manifests, raw responses, hidden-state dumps and scripts are kept locally. No demo: the PR adds a checkpoint without new behavior to show.

Self-review

  • I have reviewed the full diff and addressed the issues I found.
  • I have checked that the change follows the project's architecture and stays focused on the stated purpose.
  • I have run the checks appropriate to this change and reported commands, results, and anything I could not verify above.
  • I have checked that the PR description, documentation, and any accuracy or performance claims match the implementation and available evidence.

Open-Jev-9B uses the same method, prompt format and head design as
Open-Jev-27B-v1.1. Select the pinned revisions, request alias and
backbone dimensions by the export's model_id so one worker serves
either checkpoint, and key the export script's pins and head width the
same way. Add the 9B recipe steps and its reference comparison. 9B
runs one forward pass per candidate: candidate packing is validated on
27B only.
@twu3202

twu3202 commented Oct 7, 2026

Copy link
Copy Markdown
Contributor Author

Rebased onto main after #102. 9B keeps one forward pass per candidate, as you suggested there; 27B packs as before. The 9B answers are unchanged: 253 of 253 identical to the previous head apart from timing.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant