Skip to content

perf(open-jev): pack candidate projections within requests - #102

Merged
hsliuustc0106 merged 7 commits into
mainfrom
codex-openjev-prefill-batching
Oct 6, 2026
Merged

hsliuustc0106 merged 7 commits into
mainfrom
codex-openjev-prefill-batching

Conversation

@hsliuustc0106

@hsliuustc0106 hsliuustc0106 commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor

Purpose

Open-Jev scored each candidate prompt with separate projection calls. Pack up to 16 candidates / 4096 tokens within one request for input and gate/up GEMMs, preserving independent attention, positions, convolution, GDN state and output/down projection shapes. Regroup scores by question before calibration.

Use explicit packed/sequence-local projection calls and CUDA Graph keys containing ordered sequence lengths. Preserve upstream native vision and CUDA ABI5; cache hits consume updated input embeddings.

Test Plan

System1-Omni Version / Commit: baseline 47eff9cdeda01e4847a4fb9634a43f2cab6a233f, head 35e0001ecea52bd802ea9a0ad9a9d2dc78571301. Runtime measured at 82b5e7e60363405614bccd79ba83ab774e7513f5; final runtime source hashes match. This cleanup changes documentation/evidence only.

Compare warm localhost HTTP on one scheduler-reserved H200, BF16, concurrency 1, NUMA 0 / CPUs 0–15, using the same CUDA ABI5 library, frontend, merged export, tokenizer, head, temperature and 32 MiB GEMM workspace. Reuse one server pair per arm: real readiness, validated first request and one excluded feasibility pass, then exactly two measured passes per slice. Graphs/profiling are disabled during measurements.

Separately validate checkpoint packing, independent sequences, ordering, splitting at 17 candidates, later singletons and graph-cache hits with changed token IDs. Model base: 1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0; checkpoint: 28cf73067d5b337860bbef3c85b8b82ba8730956.

Test Result

Implementation validation, preserved in the raw snapshot:

  • cargo fmt --all --check — passed.
  • cargo clippy --workspace --locked --all-targets -- -D warnings — passed.
  • cargo test --workspace --locked — 90 passed, 13 ignored.
  • cargo build --workspace --release --locked — passed.
  • Six reserved CUDA reference tests and the checkpoint packing test — passed.
  • Changed-input graph test — passed; Nsight Systems recorded four actual CUDA Graph launches.

Cleanup validation:

  • Recomputed every per-pass latency/throughput metric from raw records and the compact CSV; values agree.
  • Verified all 824 measured latencies / 3320 decisions, workload transformations, output parity and runtime source hashes.
  • Checked relative links and git diff --check; mkdocs build --strict passed.
  • No new GPU benchmark runs; runtime code and tests are unchanged by this cleanup.

Demo / evidence

Warm HTTP workload Mean latency, main → packed Reduction
Four repeated questions 98.66 → 80.86 ms 18.04%
Eight repeated questions 197.49 → 157.77 ms 20.11%
Mixed Choice/Noul/Score 257.99 → 230.47 ms 10.67%

All compared outputs match the native reference exactly: zero probability/Score drift, flips and failures. Single-question mean/p95 increased 0.21%/1.09%, within declared 2% limits. These are two-pass observations, not confidence intervals. Repeated-question workloads are synthetic transformations of real JevBench cases.

Protocol, variability and limitations · compact timings · results, controls and source hashes.

Artifact cleanup reduces 48 files to 2 and 1,523,909 to 33,668 bytes (~98% smaller); the full PR now changes 15 files instead of 61. Each CSV row holds four measured latencies for one unique case. Duplicate request/response bodies, process logs, feasibility files and rejected-attempt records are removed from the current diff. The complete raw snapshot, collector and JevBench MIT notice remain accessible, with a verified local copy.

Other architectures, higher concurrency, production traffic, peak memory, cold capture and combined packing/graph performance remain unmeasured. Fidelity is against native main, not a new full-precision accuracy claim. Task-owned GPU processes exited and the device returned available with 0 MB used. Video is N/A for this request-replay change.

Self-review

Reviewed the code, tests, documentation and evidence against the baseline/head above. No blocking findings remain. Corrected the recipe to identify the earlier graph speedup as a single-prompt result. Local profile/ archives remain untracked and excluded from the PR.

  • I have reviewed the full diff and addressed the issues I found.
  • I have checked that the change follows the project's architecture and stays focused on the stated purpose.
  • I have run the checks appropriate to this change and reported commands, results, and anything I could not verify above.
  • I have checked that the PR description, documentation, and any accuracy or performance claims match the implementation and available evidence.

Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>
…idence

Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>
Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>
Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>
Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>
Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>
Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>
@hsliuustc0106
hsliuustc0106 marked this pull request as ready for review October 6, 2026 14:01
@twu3202

twu3202 commented Oct 6, 2026

Copy link
Copy Markdown
Contributor

Ran the packed projections against per-sequence GEMMs on an RTX 6000 Ada (sm_89), with main's default handle and synthetic inputs, over 65 layouts: 2, 4 or 16 sequences of 1 to 2,047 tokens, up to 4,096 in total.

  • 27B: bit-identical in every layout whose sequences have 33 tokens or more. Below that, some layouts differ: the attention input projection between 1 and 32 tokens (up to about half of its outputs), the GDN input projection at 1 and 17 tokens, gate/up at 1.
  • 9B (open_jev: support the Open-Jev-9B checkpoint #92): about 40 of the 65 layouts differ in each projection, at short and long lengths alike, in 0.1 to 0.2% of outputs in most of them and about 40% in some, by up to 6.3e-3 of the largest output.

So the exact match should carry over to 27B on sm_89 at realistic prompt lengths, but not to 9B. If 9B packing needs to be exact, the fixed-algorithm handle in #98 keeps each row independent of M for these shapes, at the costs listed there.

@hsliuustc0106

Copy link
Copy Markdown
Contributor Author

Thanks for running these checks. This PR currently enables packing only for Open-Jev-27B. Our exact-match result is limited to the measured H200 workload; your Ada projection tests provide useful additional evidence.

We should handle 9B packing separately when integrating #92, retaining per-candidate execution until full-output parity and performance are validated. The fixed handle in #98 is worth evaluating, but its numerical and performance tradeoffs need a separate comparison.

Full-checkpoint Ada validation and coverage of the short-input cases can also follow up separately.

@hsliuustc0106
hsliuustc0106 merged commit 9986574 into main Oct 6, 2026
4 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants