qwen3_5: reduce Gated DeltaNet preparation shared memory - #68
Merged
Merged
Conversation
Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>
Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Purpose
Reduce Qwen3.5/3.8 Gated DeltaNet preparation shared memory from 92 KiB to 72 KiB. Convert normalized Q/K to TF32 once, store the meaningful bits in three-byte component planes, retain the original four-term accumulation, and write U/W fragments directly as bfloat16. The float32 inverse and existing workspace/C ABI contracts are preserved.
Extend the recurrent float64 reference under
tests/qwen3_5/to 20 configurations covering tile/chunk boundaries, grouped heads, long inputs, weak decay and tiny/zero Q/K, without changing its tolerance. Merge current main after #55, resolve the README/test conflicts, preserve both branches' test cases, and use the sharedomni-qwen3-5-nativetest registration.Test Plan
System1-Omni Version / Commit:
33df9d75a575833b04bd3292e9b6ff32c7f6b760, incorporating main579fa8f2c4f75d7099179daf95695b6aef8766b9.Passed on the integrated branch:
cargo fmt --all --checkcargo clippy --workspace --locked --all-targets -- -D warningscargo test -p omni-jev --test frontend --locked— 12 frontend API tests.cargo test --workspace --lockedcargo build --workspace --release --lockedcargo test --release --locked -p omni-qwen3-5-native --test kernels --no-runmkdocs build --strictusing the existing prepared docs environment.One scheduler-reserved H200 GPU 2 job, NUMA 0 / CPUs 0–15: all six current GPU tests passed, including 20 recurrent GDN configurations. Newly built merged ABI4 worker/frontend binaries passed genuine readiness, first inference and one 74-request Open-Jev fidelity pass. Every probability, decision and token-usage field matches the pinned native reference exactly. Task-owned processes exited; the reservation was released with 0 MB remaining on GPU 2.
Integration controls, build hashes, commands and results, GPU test log, and all request results.
Test Result
GitHub CI passed for the updated head
33df9d7: Rust, benchmark tests and the strict documentation build. GitHub reports this head as conflict-free (MERGEABLE).Historical measurements from 2026-10-03, complete eager GDN wall time (two passes of 100 calls per shape, feasibility/warmups excluded):
The ≥10% medium/long kernel gate passed. Inputs are seeded synthetic tensors at real workload shapes; wall time includes host submission. Sampled U/W/QD/KD/PB/decay representations and final outputs matched the native baseline bitwise.
Mean warm HTTP latency, 74 single-candidate requests per pass: baseline 48.375 / 48.441 ms, candidate 48.105 / 48.305 ms. Aggregate 48.408 → 48.205 ms (0.42%), below the predeclared 2% target. These two repetitions do not establish a material full-model latency improvement.
The latency comparison used baseline main
58b8cbe, the same frozen #55 worker/frontend (202c0e1) for both arms, and ABI4 library sources7b93572except for candidate GDN preparation. Graph capture and shared-prefix caching were disabled. The integrated head's GDN source matches that measured candidate exactly; the 2026-10-04 run above validates integration correctness and makes no new performance comparison. HF Transformers and open-jev-fast were not remeasured here.Benchmark report, raw records and reproduction controls also document the separate installed FlashInfer comparison and rejected precision/accumulation variants. Modified-kernel execution on SM80/SM89 and hardware profiling counters remain unverified.
Self-review
Agent precheck completed with no actionable code correctness findings. Conflicts with merged #55 are resolved, test registration is shared and unique, and archived performance evidence is preserved. Contributor review/checklist is pending.
Before marking this PR ready for review or requesting maintainer review, complete the self-review checklist. Keep the PR in draft while this work is incomplete. The optional precheck-pr skill does not replace contributor or maintainer review.