open_jev: add native Rust/CUDA text worker - #55
Merged
Merged
Conversation
Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>
5 of 7 tasks
Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>
1 of 8 tasks
1 task done
4 tasks
Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>
Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>
Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>
Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>
Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>
4 tasks done
hsliuustc0106
marked this pull request as ready for review
October 3, 2026 06:09
Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>
Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>
Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>
1 of 10 tasks
Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>
4 tasks
Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>
xiaoyu-xyz
pushed a commit
to xiaoyu-xyz/system1-omni
that referenced
this pull request
Oct 4, 2026
- Rebased on main; the recipe index conflict was this PR and ThinkFlowLab#55 both adding an entry, so both are kept. - The comparator's docstring said it compares body bytes. It does that first, and falls back to the parsed answers when only the envelope differs; the fallback also makes key order and whitespace irrelevant, which is wider than "usage only" and is now stated rather than implied. - stub_embedder.py counted requests in a class attribute nothing read. - apple-silicon.md still called compare_with_backend.py the Laya recipe's; it is shared now and lives in recipe/.
xiaoyu-xyz
pushed a commit
to xiaoyu-xyz/system1-omni
that referenced
this pull request
Oct 4, 2026
- Rebased on main. The recipe index conflict was this PR and ThinkFlowLab#55 both adding an entry, so both are kept. - `CLM_CKPT_DIR` is not a variable `clm-serve` reads. It takes `CLM_CKPT` as the file path, and without it looks in `~/.cache/clm` and downloads -- so the command in this recipe would have quietly fetched a second copy of the checkpoint instead of using the one it names. - The comparator's docstring said it compares body bytes. It does that first and falls back to the parsed answers when only the envelope differs; that fallback also makes key order and whitespace irrelevant, which is wider than "usage only" and is now stated rather than implied. - `stub_embedder.py` counted requests in a class attribute nothing read. - `apple-silicon.md` still called `compare_with_backend.py` the Laya recipe's; it is shared now and lives in `recipe/`.
This was referenced Oct 5, 2026
6 tasks
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Purpose
Closes #54
Add native Open-Jev-27B-v1.1 text inference to the existing Rust frontend, sharing the Qwen prefill implementation with Cua-S1. On one H200, the matched single-candidate JevBench comparison reduces mean warm HTTP latency from 362.209 ms with raw HF Transformers to 48.503 ms with native Rust/CUDA: 7.47× faster, or 86.61% lower latency.
A separate native CUDA Graph experiment measures 48.086→47.112 ms (2.03% lower) on warm mixed-length requests with a bounded 64-entry cache. These are separate comparisons; the HF result uses native eager execution.
The implementation:
choice,scoreandnoulanswers.CUA_S1_GRAPH=1; graph entries are cleared before scratch storage is replaced.Compatibility: CUDA ABI 4 requires rebuilding the shared CUDA library and both native workers together. Graph mode remains opt-in. See the native recipe and validation details.
Test Plan
System1-Omni Version / Commit: PR head
c25ebc198306dd42f7a3c5d22325df0d9a0b27b7; target-main snapshot and merge base10a05540a28c9af8c360ceaff0d67c36f17ba3bb. Latest main is merged; the README resolution preserves its badges, navigation and section layout alongside the Open-Jev support and news.Validate prompts, typed answers and token IDs against pinned Open-Jev fixtures; run workspace checks and reserved-GPU kernel and worker validation. Open-Jev tests and fixtures reside under
tests/open_jev/; shared Qwen configuration and CUDA tests reside undertests/qwen3_5/. Cargo registers integration tests explicitly, and private unit tests load their files from the same top-level tree. Related frontend CPU mock-worker coverage is tracked in issue #46 and PR #58.For performance comparisons, freeze sources and controls before execution, exclude one feasibility pass, and collect exactly two measured passes per configuration. Reuse one server per configuration. The mixed workload contains 74 real JevBench
noulcases, one candidate per request, 80–3399 input tokens, BF16, maximum length 16384 and concurrency 1. Downloads, export, loading, process-to-readiness, first-inference validation and warmup are outside measured execution.Warm HTTP timing includes the localhost request through the Rust frontend, tokenization, model execution and UTF8 response decoding; it excludes client body serialization and response JSON parsing. P50 is the median; P95 uses JevBench's
sorted[int(0.95*N)-1]rank. Two-pass ranges describe observed variability, not confidence intervals. Separately profiled kernel totals are not HTTP latency.Test Result
1. Step-by-step implementation and optimization comparison
Each measured row uses its own paired baseline. The sequence documents the changes and their evidence; it is not a cumulative latency waterfall. The RMSNorm experiment used GPU 5, while the later experiments used GPU 2. October 1–2 HTTP workers may have retained CUPTI instrumentation from an inactive Nsight session; October 3 HTTP measurements were unprofiled. Therefore, the gains below must not be added or multiplied across experiments.
For steps 2 and 3, the two measured pass means establish the observed spread:
Separate two-trace kernel measurements support the intended mechanism: the 107-token residual-RMSNorm family falls 1.944→0.464 ms (76.16%), and the 3399-token SiLU family falls 21.217→6.380 ms (69.93%). Original Fast's long-request SiLU total is 5.466 ms, so packing closes 94.19% of that measured kernel-family gap. These are kernel-family sums on selected requests, not incremental HTTP savings. Nsight Systems supplied the traces because Nsight Compute hardware counters were unavailable.
2. Complete backend comparison against raw HF Transformers
2026-10-03, exact H200 GPU 2 (UUID
GPU-cbf66259-f4ab-0ede-1811-82037dde5924), NUMA 0, CPUs 0–15. All three backends were newly measured through the same frozen Rust frontend, using the same model, temperature and 74 requests. Each has 148 measured responses across two passes; no Nsight launcher or trace collection was used.Native versus raw HF: 7.47× speedup and 86.61% lower mean latency. This measures the complete backend change; it does not isolate LoRA merging or individual kernels. Native's observed mean is 4.78% below original Fast, while Fast has a lower median. The two-pass budget and single-candidate subset do not establish a general backend winner. The author's 17.3 ms B300 result uses different hardware and workload.
The HF baseline uses the original Open-Jev server and
DecisionModel, its unmerged PEFT adapter, trained head and saved temperature. Full attention uses stock SDPA. Runtime checks confirm 160 unmerged LoRA modules and all 48 linear-attention layers ontorch_chunk_gated_delta_rule, stock convolution andQwen3_5RMSNormGated. Optional FLA/causal-conv1d availability checks are disabled before model imports in the same prepared environment. HF uses no custom Fast model,torch.compile, CUDA Graph replay or prefix cache. Original Fast retains its original kernel and CUDA Graph path, reporting 30 retained graphs.Native and HF agree on all 74 thresholded decisions and both score 64/74; maximum probability difference is 0.020423. Fast scores 63/74 with one different decision,
hard-opus-a-temporal_numeric-09; native/Fast maximum probability difference is 0.034353. Each backend's outputs remain exactly stable across its feasibility and measured passes. These results do not establish full numerical parity or statistical accuracy superiority. Multi-candidate prefix sharing and broader JevBench coverage remain unvalidated.Frozen inputs: native worker/frontend
202c0e1with the accepted packed-SiLU library; PR source at this measurementad1cb81; Open-Jev3308a15, Fastc52b8bb, JevBenchf8ce713; base1d4bf0f, adapter28cf730, temperature 2.5343690298472983. Environment: Torch 2.13.0+cu130, Transformers 5.10.2, PEFT 0.19.1.A collector cleanup assertion treated native's intentional SIGTERM as a failure after all HF/native results were saved. Only the remaining Fast configuration continued under a second reservation of the same exact GPU and affinity; no additional measured passes were introduced.
3. Latest native CUDA Graph comparison
The mixed workload has 57 distinct input lengths, exceeding the eight-entry cache capacity. Increasing the bounded cache to 64 allows all these lengths to remain cached after warmup. All variants use identical frozen sources, build options, dependencies, CUDA library, frontend and inputs; graph capacities differ only in the model's cache limit. Eager uses the eight-entry worker with graph mode disabled.
Graph-64 reduces warm mean latency by 3.77% for the short workload and 2.03% for the mixed workload, each against matched eager execution. Increasing capacity resolves the mixed-length regression; it does not improve the short mean over graph-eight in this run. Both graph-64 measured pass means are below eager, and the mixed aggregate narrowly exceeds the prespecified 2% improvement gate.
Controls: same exact GPU 2 and affinity as above, one reused server per configuration, first long inference validated after real readiness to preallocate maximum workload scratch storage, then short-request validation. Each workload receives one excluded feasibility pass and exactly two measured passes: 32 repeated short requests or all 74 mixed cases per pass. HTTP measurements are unprofiled. All 954 feasibility/measured responses and six first-inference validations preserve probabilities and decisions exactly (maximum delta 0.0).
After the mixed passes, graph-64's observed device memory is 122 MB above graph-eight and 142 MB above eager; these are scheduler samples, not peak measurements. The excluded graph-64 mixed feasibility mean is 110.255 ms, documenting capture cost rather than a controlled cold-latency comparison. New lengths, more than 64 active lengths or scratch growth can require recapture, so
CUA_S1_GRAPH=1remains opt-in.A separate preceding Nsight experiment confirms one graph launch, zero recaptures and zero individual runtime kernel-launch calls per warm short request, with all 834 GPU kernels retained. Node-level graph tracing shows larger gaps despite lower unprofiled HTTP latency, so those traces establish replay behavior rather than a speedup estimate. The shared Cua-S1 graph cache also increases to 64; full Cua-S1 checkpoint graph inference has not been revalidated.
4. Automated checks and evidence
The following checks passed after test relocation and the graph-cache update:
cargo fmt --all --check cargo clippy --workspace --locked --offline --all-targets -- -D warnings cargo test --workspace --locked --offline cargo build --workspace --release --locked --offlinegit diff --checkpassed. Head CI reports success for Rust, benchmarks and the documentation build; documentation deployment is skipped. Format, workspace Clippy/tests/release build, strict MkDocs, all 25 local README links and rendered navigation anchors also passed locally after this merge.61b83b3. Kernel sources and test contents are unchanged since those checks; graph-cache GPU measurements are reported separately above. SM89 compilation passed, but SM89 device execution and full Cua-S1 checkpoint inference were not checked.Local experiment archives under
profile/:jev-single-candidate-rmsnorm-register-cache-20261001/jev-single-candidate-silu-pack8-20261002/jev-hf-transformers-comparison-20261003/jev-cuda-graph-20261003-074613/jev-cuda-graph-cache64-20261003-075238/Published controls, revisions and reproduction details are in the validation page. The local archives retain complete raw evidence; they are not included in this PR.
Self-review
Contributor checklist from CONTRIBUTING.md, pending completion before requesting review: