Skip to content

qwen3_5: reduce Gated DeltaNet preparation shared memory - #68

Merged
hsliuustc0106 merged 2 commits into
mainfrom
codex-gdn-prefill-optimization
Oct 4, 2026
Merged

hsliuustc0106 merged 2 commits into
mainfrom
codex-gdn-prefill-optimization

Conversation

@hsliuustc0106

@hsliuustc0106 hsliuustc0106 commented Oct 3, 2026 •

Copy link
Copy Markdown
Contributor

Purpose

Reduce Qwen3.5/3.8 Gated DeltaNet preparation shared memory from 92 KiB to 72 KiB. Convert normalized Q/K to TF32 once, store the meaningful bits in three-byte component planes, retain the original four-term accumulation, and write U/W fragments directly as bfloat16. The float32 inverse and existing workspace/C ABI contracts are preserved.

Extend the recurrent float64 reference under tests/qwen3_5/ to 20 configurations covering tile/chunk boundaries, grouped heads, long inputs, weak decay and tiny/zero Q/K, without changing its tolerance. Merge current main after #55, resolve the README/test conflicts, preserve both branches' test cases, and use the shared omni-qwen3-5-native test registration.

Test Plan

System1-Omni Version / Commit: 33df9d75a575833b04bd3292e9b6ff32c7f6b760, incorporating main 579fa8f2c4f75d7099179daf95695b6aef8766b9.

Passed on the integrated branch:

  • cargo fmt --all --check
  • cargo clippy --workspace --locked --all-targets -- -D warnings
  • cargo test -p omni-jev --test frontend --locked — 12 frontend API tests.
  • cargo test --workspace --locked
  • cargo build --workspace --release --locked
  • cargo test --release --locked -p omni-qwen3-5-native --test kernels --no-run
  • CUDA 13.0 backend builds for SM80, SM89 and SM90.
  • Benchmark smoke-manifest validation and all 8 benchmark unit tests.
  • mkdocs build --strict using the existing prepared docs environment.

One scheduler-reserved H200 GPU 2 job, NUMA 0 / CPUs 0–15: all six current GPU tests passed, including 20 recurrent GDN configurations. Newly built merged ABI4 worker/frontend binaries passed genuine readiness, first inference and one 74-request Open-Jev fidelity pass. Every probability, decision and token-usage field matches the pinned native reference exactly. Task-owned processes exited; the reservation was released with 0 MB remaining on GPU 2.

Integration controls, build hashes, commands and results, GPU test log, and all request results.

Test Result

GitHub CI passed for the updated head 33df9d7: Rust, benchmark tests and the strict documentation build. GitHub reports this head as conflict-free (MERGEABLE).

Historical measurements from 2026-10-03, complete eager GDN wall time (two passes of 100 calls per shape, feasibility/warmups excluded):

Tokens Baseline pass 1 / 2 Candidate pass 1 / 2 Reduction pass 1 / 2
107 0.051291 / 0.050976 ms 0.050914 / 0.050957 ms 0.74% / 0.04%
936 0.247565 / 0.247337 ms 0.214176 / 0.214178 ms 13.49% / 13.41%
3,399 0.800775 / 0.798935 ms 0.696768 / 0.697262 ms 12.99% / 12.73%

The ≥10% medium/long kernel gate passed. Inputs are seeded synthetic tensors at real workload shapes; wall time includes host submission. Sampled U/W/QD/KD/PB/decay representations and final outputs matched the native baseline bitwise.

Mean warm HTTP latency, 74 single-candidate requests per pass: baseline 48.375 / 48.441 ms, candidate 48.105 / 48.305 ms. Aggregate 48.408 → 48.205 ms (0.42%), below the predeclared 2% target. These two repetitions do not establish a material full-model latency improvement.

The latency comparison used baseline main 58b8cbe, the same frozen #55 worker/frontend (202c0e1) for both arms, and ABI4 library sources 7b93572 except for candidate GDN preparation. Graph capture and shared-prefix caching were disabled. The integrated head's GDN source matches that measured candidate exactly; the 2026-10-04 run above validates integration correctness and makes no new performance comparison. HF Transformers and open-jev-fast were not remeasured here.

Benchmark report, raw records and reproduction controls also document the separate installed FlashInfer comparison and rejected precision/accumulation variants. Modified-kernel execution on SM80/SM89 and hardware profiling counters remain unverified.

Self-review

Agent precheck completed with no actionable code correctness findings. Conflicts with merged #55 are resolved, test registration is shared and unique, and archived performance evidence is preserved. Contributor review/checklist is pending.

Before marking this PR ready for review or requesting maintainer review, complete the self-review checklist. Keep the PR in draft while this work is incomplete. The optional precheck-pr skill does not replace contributor or maintainer review.

  • I have reviewed the full diff and addressed the issues I found.
  • I have checked that the change follows the project's architecture and stays focused on the stated purpose.
  • I have run the checks appropriate to this change and reported commands, results, and anything I could not verify above.
  • I have checked that the PR description, documentation, and any accuracy or performance claims match the implementation and available evidence.

Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>
Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>
@hsliuustc0106
hsliuustc0106 marked this pull request as ready for review October 4, 2026 17:19
@hsliuustc0106
hsliuustc0106 merged commit 07e67e1 into main Oct 4, 2026
4 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant