Skip to content

open_jev: add native Rust/CUDA text worker - #55

Merged
hsliuustc0106 merged 12 commits into
mainfrom
codex-open-jev-native
Oct 4, 2026
Merged

hsliuustc0106 merged 12 commits into
mainfrom
codex-open-jev-native

Conversation

@hsliuustc0106

@hsliuustc0106 hsliuustc0106 commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

Purpose

Closes #54

Add native Open-Jev-27B-v1.1 text inference to the existing Rust frontend, sharing the Qwen prefill implementation with Cua-S1. On one H200, the matched single-candidate JevBench comparison reduces mean warm HTTP latency from 362.209 ms with raw HF Transformers to 48.503 ms with native Rust/CUDA: 7.47× faster, or 86.61% lower latency.

A separate native CUDA Graph experiment measures 48.086→47.112 ms (2.03% lower) on warm mixed-length requests with a bounded 64-entry cache. These are separate comparisons; the HF result uses native eager execution.

The implementation:

  • Compiles and tokenizes independent candidate prompts, merges LoRA weights, applies the trained scalar decision head and saved temperature, and returns typed choice, score and noul answers.
  • Shares model execution between Open-Jev and Cua-S1 while preserving Cua-S1's public imports; adds a CPU checkpoint export recipe and real inference before worker readiness.
  • Fuses attention gating, retains residual RMSNorm values in registers, and packs BF16 SiLU loads/stores. RMSNorm reduction order and BF16 rounding are preserved; unsupported SiLU layouts retain the scalar path.
  • Retains up to 64 exact-length CUDA Graphs for opt-in replay with CUA_S1_GRAPH=1; graph entries are cleared before scratch storage is replaced.
  • Adds contract, tokenizer and CUDA coverage under the repository-level tests tree, along with setup instructions, model-support news and measured validation results.

Compatibility: CUDA ABI 4 requires rebuilding the shared CUDA library and both native workers together. Graph mode remains opt-in. See the native recipe and validation details.

Test Plan

System1-Omni Version / Commit: PR head c25ebc198306dd42f7a3c5d22325df0d9a0b27b7; target-main snapshot and merge base 10a05540a28c9af8c360ceaff0d67c36f17ba3bb. Latest main is merged; the README resolution preserves its badges, navigation and section layout alongside the Open-Jev support and news.

Validate prompts, typed answers and token IDs against pinned Open-Jev fixtures; run workspace checks and reserved-GPU kernel and worker validation. Open-Jev tests and fixtures reside under tests/open_jev/; shared Qwen configuration and CUDA tests reside under tests/qwen3_5/. Cargo registers integration tests explicitly, and private unit tests load their files from the same top-level tree. Related frontend CPU mock-worker coverage is tracked in issue #46 and PR #58.

For performance comparisons, freeze sources and controls before execution, exclude one feasibility pass, and collect exactly two measured passes per configuration. Reuse one server per configuration. The mixed workload contains 74 real JevBench noul cases, one candidate per request, 80–3399 input tokens, BF16, maximum length 16384 and concurrency 1. Downloads, export, loading, process-to-readiness, first-inference validation and warmup are outside measured execution.

Warm HTTP timing includes the localhost request through the Rust frontend, tokenization, model execution and UTF8 response decoding; it excludes client body serialization and response JSON parsing. P50 is the median; P95 uses JevBench's sorted[int(0.95*N)-1] rank. Two-pass ranges describe observed variability, not confidence intervals. Separately profiled kernel totals are not HTTP latency.

Test Result

1. Step-by-step implementation and optimization comparison

Each measured row uses its own paired baseline. The sequence documents the changes and their evidence; it is not a cumulative latency waterfall. The RMSNorm experiment used GPU 5, while the later experiments used GPU 2. October 1–2 HTTP workers may have retained CUPTI instrumentation from an inactive Nsight session; October 3 HTTP measurements were unprofiled. Therefore, the gains below must not be added or multiplied across experiments.

Step Change and comparison Mean warm HTTP before → after (ms) Observed effect Experiment
1 Native model path: shared Qwen execution, merged LoRA, scalar head and fused attention gating No isolated measurement for these changes individually Their combined contribution is included in the complete HF/backend comparison below Native implementation
2 Uncached → register-cached residual RMSNorm, with fixed-width specialization 51.424 → 50.097 2.58% lower 2026-10-01, H200 GPU 5
3 Scalar → eight-element packed BF16 SiLU, with cached RMSNorm fixed 49.527 → 48.297 2.48% lower 2026-10-02, H200 GPU 2
4 Eager → existing eight-entry graph replay, repeated 107-token requests 19.829 → 18.882 4.77% lower; mixed-length latency instead rises from 48.086 to 95.289 ms (98.16% higher) 2026-10-03, H200 GPU 2
5 Eight-entry → 64-entry graph cache, warm mixed-length requests 95.289 → 47.112 50.56% lower than graph-eight; 2.03% lower than the matched eager baseline 2026-10-03, H200 GPU 2

For steps 2 and 3, the two measured pass means establish the observed spread:

Isolated native A/B Baseline pass 1 / pass 2 (ms) Optimized pass 1 / pass 2 (ms) Output validation
Residual RMSNorm 51.440 / 51.409 50.144 / 50.049 All 74 probabilities and decisions unchanged; maximum delta 0.0
Packed SiLU 49.563 / 49.490 48.286 / 48.307 All 74 probabilities and decisions unchanged; maximum delta 0.0

Separate two-trace kernel measurements support the intended mechanism: the 107-token residual-RMSNorm family falls 1.944→0.464 ms (76.16%), and the 3399-token SiLU family falls 21.217→6.380 ms (69.93%). Original Fast's long-request SiLU total is 5.466 ms, so packing closes 94.19% of that measured kernel-family gap. These are kernel-family sums on selected requests, not incremental HTTP savings. Nsight Systems supplied the traces because Nsight Compute hardware counters were unavailable.

2. Complete backend comparison against raw HF Transformers

2026-10-03, exact H200 GPU 2 (UUID GPU-cbf66259-f4ab-0ede-1811-82037dde5924), NUMA 0, CPUs 0–15. All three backends were newly measured through the same frozen Rust frontend, using the same model, temperature and 74 requests. Each has 148 measured responses across two passes; no Nsight launcher or trace collection was used.

Configuration Overall mean (ms) Mean, pass 1 / pass 2 (ms) P50, pass 1 / pass 2 (ms) P95, pass 1 / pass 2 (ms) Correct / 74
Raw HF Transformers: unmerged LoRA, PyTorch fallback 362.209 362.238 / 362.180 324.822 / 321.321 651.999 / 652.911 64
Native Rust/CUDA: merged LoRA, cached RMSNorm, packed SiLU; eager 48.503 48.471 / 48.535 25.797 / 24.895 215.696 / 218.882 64
Original OpenJev-Fast: custom kernels and graph stack 50.936 51.097 / 50.775 24.651 / 24.622 248.121 / 251.193 63

Native versus raw HF: 7.47× speedup and 86.61% lower mean latency. This measures the complete backend change; it does not isolate LoRA merging or individual kernels. Native's observed mean is 4.78% below original Fast, while Fast has a lower median. The two-pass budget and single-candidate subset do not establish a general backend winner. The author's 17.3 ms B300 result uses different hardware and workload.

The HF baseline uses the original Open-Jev server and DecisionModel, its unmerged PEFT adapter, trained head and saved temperature. Full attention uses stock SDPA. Runtime checks confirm 160 unmerged LoRA modules and all 48 linear-attention layers on torch_chunk_gated_delta_rule, stock convolution and Qwen3_5RMSNormGated. Optional FLA/causal-conv1d availability checks are disabled before model imports in the same prepared environment. HF uses no custom Fast model, torch.compile, CUDA Graph replay or prefix cache. Original Fast retains its original kernel and CUDA Graph path, reporting 30 retained graphs.

Native and HF agree on all 74 thresholded decisions and both score 64/74; maximum probability difference is 0.020423. Fast scores 63/74 with one different decision, hard-opus-a-temporal_numeric-09; native/Fast maximum probability difference is 0.034353. Each backend's outputs remain exactly stable across its feasibility and measured passes. These results do not establish full numerical parity or statistical accuracy superiority. Multi-candidate prefix sharing and broader JevBench coverage remain unvalidated.

Frozen inputs: native worker/frontend 202c0e1 with the accepted packed-SiLU library; PR source at this measurement ad1cb81; Open-Jev 3308a15, Fast c52b8bb, JevBench f8ce713; base 1d4bf0f, adapter 28cf730, temperature 2.5343690298472983. Environment: Torch 2.13.0+cu130, Transformers 5.10.2, PEFT 0.19.1.

A collector cleanup assertion treated native's intentional SIGTERM as a failure after all HF/native results were saved. Only the remaining Fast configuration continued under a second reservation of the same exact GPU and affinity; no additional measured passes were introduced.

3. Latest native CUDA Graph comparison

The mixed workload has 57 distinct input lengths, exceeding the eight-entry cache capacity. Increasing the bounded cache to 64 allows all these lengths to remain cached after warmup. All variants use identical frozen sources, build options, dependencies, CUDA library, frontend and inputs; graph capacities differ only in the model's cache limit. Eager uses the eight-entry worker with graph mode disabled.

Workload Configuration Mean (ms) Mean, pass 1 / pass 2 (ms)
Repeated 107-token request Eager 19.829 19.835 / 19.823
Repeated 107-token request Graph, eight entries 18.882 18.902 / 18.863
Repeated 107-token request Graph, 64 entries 19.081 18.968 / 19.194
74 mixed-length cases Eager 48.086 48.062 / 48.110
74 mixed-length cases Graph, eight entries 95.289 95.425 / 95.153
74 mixed-length cases Graph, 64 entries 47.112 47.050 / 47.174

Graph-64 reduces warm mean latency by 3.77% for the short workload and 2.03% for the mixed workload, each against matched eager execution. Increasing capacity resolves the mixed-length regression; it does not improve the short mean over graph-eight in this run. Both graph-64 measured pass means are below eager, and the mixed aggregate narrowly exceeds the prespecified 2% improvement gate.

Controls: same exact GPU 2 and affinity as above, one reused server per configuration, first long inference validated after real readiness to preallocate maximum workload scratch storage, then short-request validation. Each workload receives one excluded feasibility pass and exactly two measured passes: 32 repeated short requests or all 74 mixed cases per pass. HTTP measurements are unprofiled. All 954 feasibility/measured responses and six first-inference validations preserve probabilities and decisions exactly (maximum delta 0.0).

After the mixed passes, graph-64's observed device memory is 122 MB above graph-eight and 142 MB above eager; these are scheduler samples, not peak measurements. The excluded graph-64 mixed feasibility mean is 110.255 ms, documenting capture cost rather than a controlled cold-latency comparison. New lengths, more than 64 active lengths or scratch growth can require recapture, so CUA_S1_GRAPH=1 remains opt-in.

A separate preceding Nsight experiment confirms one graph launch, zero recaptures and zero individual runtime kernel-launch calls per warm short request, with all 834 GPU kernels retained. Node-level graph tracing shows larger gaps despite lower unprofiled HTTP latency, so those traces establish replay behavior rather than a speedup estimate. The shared Cua-S1 graph cache also increases to 64; full Cua-S1 checkpoint graph inference has not been revalidated.

4. Automated checks and evidence

The following checks passed after test relocation and the graph-cache update:

cargo fmt --all --check
cargo clippy --workspace --locked --offline --all-targets -- -D warnings
cargo test --workspace --locked --offline
cargo build --workspace --release --locked --offline
  • Strict MkDocs build, documentation link/claim checks and git diff --check passed. Head CI reports success for Rust, benchmarks and the documentation build; documentation deployment is skipped. Format, workspace Clippy/tests/release build, strict MkDocs, all 25 local README links and rendered navigation anchors also passed locally after this merge.
  • Test relocation preserved all 38 registered test names, including ignored GPU/checkpoint cases; 29 CPU tests and the pinned-checkpoint tokenizer check passed. The relocated six-test GPU suite compiled in release mode without device execution.
  • Six CUDA ABI 4 kernel tests and live worker/frontend smoke checks passed previously on reserved H200 GPU 2 at 61b83b3. Kernel sources and test contents are unchanged since those checks; graph-cache GPU measurements are reported separately above. SM89 compilation passed, but SM89 device execution and full Cua-S1 checkpoint inference were not checked.
  • The integrated PR model source matches the measured 64-entry candidate exactly. Raw commands, plans, frozen source copies, responses, token IDs, traces, memory samples and checksums are preserved outside the PR. Task-owned GPU processes exited and their reservations were released.

Local experiment archives under profile/:

Evidence Archive directory
Residual RMSNorm A/B jev-single-candidate-rmsnorm-register-cache-20261001/
Packed SiLU A/B and kernel traces jev-single-candidate-silu-pack8-20261002/
Matched HF/native/original-Fast comparison jev-hf-transformers-comparison-20261003/
Graph replay proof and traces jev-cuda-graph-20261003-074613/
Matched eager/graph-eight/graph-64 comparison jev-cuda-graph-cache64-20261003-075238/

Published controls, revisions and reproduction details are in the validation page. The local archives retain complete raw evidence; they are not included in this PR.

Self-review

Contributor checklist from CONTRIBUTING.md, pending completion before requesting review:

  • I have reviewed the full diff and addressed the issues I found.
  • I have checked that the change follows the project's architecture and stays focused on the stated purpose.
  • I have run the checks appropriate to this change and reported commands, results, and anything I could not verify above.
  • I have checked that the PR description, documentation, and any accuracy or performance claims match the implementation and available evidence.

Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>
Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>
Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>
Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>
Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>
Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>
Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>
@hsliuustc0106
hsliuustc0106 marked this pull request as ready for review October 3, 2026 06:09
Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>
Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>
Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>
Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>
Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>
@hsliuustc0106
hsliuustc0106 merged commit 873655b into main Oct 4, 2026
4 checks passed
xiaoyu-xyz pushed a commit to xiaoyu-xyz/system1-omni that referenced this pull request Oct 4, 2026
- Rebased on main; the recipe index conflict was this PR and ThinkFlowLab#55 both adding an
  entry, so both are kept.
- The comparator's docstring said it compares body bytes. It does that first, and
  falls back to the parsed answers when only the envelope differs; the fallback
  also makes key order and whitespace irrelevant, which is wider than "usage
  only" and is now stated rather than implied.
- stub_embedder.py counted requests in a class attribute nothing read.
- apple-silicon.md still called compare_with_backend.py the Laya recipe's; it is
  shared now and lives in recipe/.
xiaoyu-xyz pushed a commit to xiaoyu-xyz/system1-omni that referenced this pull request Oct 4, 2026
- Rebased on main. The recipe index conflict was this PR and ThinkFlowLab#55 both adding an
  entry, so both are kept.
- `CLM_CKPT_DIR` is not a variable `clm-serve` reads. It takes `CLM_CKPT` as the
  file path, and without it looks in `~/.cache/clm` and downloads -- so the
  command in this recipe would have quietly fetched a second copy of the
  checkpoint instead of using the one it names.
- The comparator's docstring said it compares body bytes. It does that first and
  falls back to the parsed answers when only the envelope differs; that fallback
  also makes key order and whitespace irrelevant, which is wider than "usage
  only" and is now stated rather than implied.
- `stub_embedder.py` counted requests in a class attribute nothing read.
- `apple-silicon.md` still called `compare_with_backend.py` the Laya recipe's;
  it is shared now and lives in `recipe/`.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[New Model]: Native Rust/CUDA support for Open-Jev-27B-v1.1

1 participant