Repository navigation
perf(open-jev): pack candidate projections within requests - #102
Conversation
Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>
…idence Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>
Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>
Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>
Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>
Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>
Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>
|
Ran the packed projections against per-sequence GEMMs on an RTX 6000 Ada (sm_89), with main's default handle and synthetic inputs, over 65 layouts: 2, 4 or 16 sequences of 1 to 2,047 tokens, up to 4,096 in total.
So the exact match should carry over to 27B on sm_89 at realistic prompt lengths, but not to 9B. If 9B packing needs to be exact, the fixed-algorithm handle in #98 keeps each row independent of M for these shapes, at the costs listed there. |
|
Thanks for running these checks. This PR currently enables packing only for Open-Jev-27B. Our exact-match result is limited to the measured H200 workload; your Ada projection tests provide useful additional evidence. We should handle 9B packing separately when integrating #92, retaining per-candidate execution until full-output parity and performance are validated. The fixed handle in #98 is worth evaluating, but its numerical and performance tradeoffs need a separate comparison. Full-checkpoint Ada validation and coverage of the short-input cases can also follow up separately. |
Purpose
Open-Jev scored each candidate prompt with separate projection calls. Pack up to 16 candidates / 4096 tokens within one request for input and gate/up GEMMs, preserving independent attention, positions, convolution, GDN state and output/down projection shapes. Regroup scores by question before calibration.
Use explicit packed/sequence-local projection calls and CUDA Graph keys containing ordered sequence lengths. Preserve upstream native vision and CUDA ABI5; cache hits consume updated input embeddings.
Test Plan
System1-Omni Version / Commit: baseline
47eff9cdeda01e4847a4fb9634a43f2cab6a233f, head35e0001ecea52bd802ea9a0ad9a9d2dc78571301. Runtime measured at82b5e7e60363405614bccd79ba83ab774e7513f5; final runtime source hashes match. This cleanup changes documentation/evidence only.Compare warm localhost HTTP on one scheduler-reserved H200, BF16, concurrency 1, NUMA 0 / CPUs 0–15, using the same CUDA ABI5 library, frontend, merged export, tokenizer, head, temperature and 32 MiB GEMM workspace. Reuse one server pair per arm: real readiness, validated first request and one excluded feasibility pass, then exactly two measured passes per slice. Graphs/profiling are disabled during measurements.
Separately validate checkpoint packing, independent sequences, ordering, splitting at 17 candidates, later singletons and graph-cache hits with changed token IDs. Model base:
1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0; checkpoint:28cf73067d5b337860bbef3c85b8b82ba8730956.Test Result
Implementation validation, preserved in the raw snapshot:
cargo fmt --all --check— passed.cargo clippy --workspace --locked --all-targets -- -D warnings— passed.cargo test --workspace --locked— 90 passed, 13 ignored.cargo build --workspace --release --locked— passed.Cleanup validation:
git diff --check;mkdocs build --strictpassed.Demo / evidence
All compared outputs match the native reference exactly: zero probability/Score drift, flips and failures. Single-question mean/p95 increased 0.21%/1.09%, within declared 2% limits. These are two-pass observations, not confidence intervals. Repeated-question workloads are synthetic transformations of real JevBench cases.
Protocol, variability and limitations · compact timings · results, controls and source hashes.
Artifact cleanup reduces 48 files to 2 and 1,523,909 to 33,668 bytes (~98% smaller); the full PR now changes 15 files instead of 61. Each CSV row holds four measured latencies for one unique case. Duplicate request/response bodies, process logs, feasibility files and rejected-attempt records are removed from the current diff. The complete raw snapshot, collector and JevBench MIT notice remain accessible, with a verified local copy.
Other architectures, higher concurrency, production traffic, peak memory, cold capture and combined packing/graph performance remain unmeasured. Fidelity is against native main, not a new full-precision accuracy claim. Task-owned GPU processes exited and the device returned available with 0 MB used. Video is N/A for this request-replay change.
Self-review
Reviewed the code, tests, documentation and evidence against the baseline/head above. No blocking findings remain. Corrected the recipe to identify the earlier graph speedup as a single-prompt result. Local
profile/archives remain untracked and excluded from the PR.