Skip to content

qwen3_5: add prefix continuation operations - #97

Draft
twu3202 wants to merge 1 commit into
ThinkFlowLab:mainfrom
twu3202:prefix-continuation-ops
Draft

twu3202 wants to merge 1 commit into
ThinkFlowLab:mainfrom
twu3202:prefix-continuation-ops

Conversation

@twu3202

@twu3202 twu3202 commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor

Purpose

Second PR for #85, the backend step of the plan as revised in this update. It adds CUDA operations that continue a sequence after a shared prefix: the part of the prompt shared by all of a request's candidates (request prefix), or by one question's candidates (question prefix). The shared Qwen executor calls them in #99. The workers don't use them, and their outputs are unchanged. This is a draft while the backend path is settled with #96, which adds similar operations for its JEV-VL worker; the update compares the two.

  • cs1_gdn_conv_history: the Gated DeltaNet conv, reading the conv inputs of the three positions before its first token ([3, channels]) and optionally writing those of its last three.
  • cs1_gdn_prefill_state: the chunked prefill, starting from a float32 state [H, 128, 128] and optionally writing the final one. The two may be the same buffer; with no tokens, the final state is the initial one.
  • cs1_attention_gated_cached: gated attention for Tq queries at the end of Tk positions, with the keys and values of all Tk positions. Key tiles still start at position 0, so each query reads its keys in the same tiles and order as in the unsplit call.
  • cs1_copy_rows: a queued, pitched device-to-device copy. It isn't in the plan, and nothing calls it yet. The executor reads values in rows of the fused projection output's width, and the library has no device-to-device copy, so qwen3_5: run prompts that share prefixes once per prefix #99 needs this to keep the prefix's and the branch's value rows together. Adding it there would need another ABI change.

The existing cs1_gdn_conv, cs1_gdn_prefill and cs1_attention_gated now call these operations without history, state or cached positions. Attention also skips warps whose rows are all past the last query. Those rows were never stored, so outputs don't change, and short branches over long prefixes do less wasted work.

The ABI moves from 5 (#64's vision operations) to 6. The Rust bindings in cuda.rs and the ABI notes in the backend README, the recipes and src/backends/cuda/contract.md are updated to match; the notes still said 4. #96 also moves the ABI to 6, with different entry points, so whichever lands second takes the next number. If #77 lands first, its manifest's abi_version needs the same update; the checker from #25 compares it with CS1_ABI_VERSION.

Test Plan

System1-Omni Version / Commit: 4678e3d on main 47eff9c. The kernel tests, register counts and plain-forward comparison ran on this code. The timing was measured before the rebase, at d55c46f on 7f39ac4; the conv, gated delta rule and attention sources haven't changed since on either side. src/backends/cuda/check_contract.py reports the same as on main: qwen3_5 has no manifest until #77.

CPU: the five Rust steps in CI (format, Clippy, frontend tests, workspace tests, release build), and mkdocs build --strict.

GPU, one RTX 6000 Ada (sm_89), CUDA 13.2, driver 595.91.07, library built with build.sh <dir> 89:

  • all GPU kernel tests (--ignored), including three new ones that compare each continuation with the unsplit call and one for the row copy. The comparisons use the 4B/9B and 27B head counts, prefix lengths 1, 2, 3, 31, 63, 64, 65, 127, 128, 129, 1,000 and 1,024, and branch lengths 1, 2, 64, 65 and 200: 60 splits;
  • the plain forward on main and on this branch, each with its own library, workers side by side: Cua-S1 and Open-Jev-9B;
  • timing of the plain calls on both libraries.

Test Result

Build. No compiler warnings with -Wall -Wextra. With -Xptxas -v on sm_89, attention keeps 238 registers, the state kernel goes from 96 to 94 and the conv from 21 to 36, and the new history kernel uses 10, all without spills.

Kernel tests. 10 of 10 pass (6 existing, 4 new).

  • Conv: every output is bit-identical to the unsplit conv at all 60 tested splits and at both widths. So are chains of request prefix, question prefix and branch, including chains shorter than the 4-tap conv window. The written history matches the expected inputs, with zeros before the start.
  • Gated delta rule:
    • Writing the final state leaves the prefix output bit-identical. The state after each prefix is within 3.6e-3 of the float64 recurrent reference, whose largest values are 0.14 to 1.04.
    • Splits at 64, 128 and 1,024 tokens, and a chain through 64 and 128, are bit-identical to the unsplit prefill.
    • At the other splits, outputs are within about 1.2e-2 of the float64 reference, relative to its largest magnitude, the same as the unsplit kernel in the existing test on other inputs. They are at most 1.46e-3 from the unsplit output, three BF16 steps at the output's magnitude of about 0.1.
    • In the unaligned chains (100 then 192, and 1,000 then 1,100), the question prefix's output, its state and the branch are each within the same tolerances of the float64 reference.
    • An in-place state update gives the same bits as separate buffers. No tokens copies the state, and a misaligned initial or final state pointer is rejected.
  • Cached attention: every output row is bit-identical to the unsplit call at all 60 tested splits, at both head counts, with strided values. No queries is accepted; more queries than keys is rejected.
  • Row copy: copies pitched rows and leaves the rest of the destination untouched.

Plain forward unchanged (main 47eff9c with ABI 5 against this branch with ABI 6):

Timing of the plain calls (9B shapes; one process per library, in the order main, here, here, main, main, here, three rounds; one other GPU process was listed at the start of the first run and none after; median of 9 runs each, with the range in parentheses, microseconds per call):

Tokens Conv, main / here Gated delta rule Gated attention
107 5.23 / 5.15 35.63 / 35.73 20.63 / 20.49
936 35.37 / 34.08 181.0 (165 to 188) / 188.7 (163 to 192) 114.7 / 113.7
3,399 163.9 / 159.1 826.3 (801 to 841) / 833.4 (788 to 844) 1042.3 / 1037.3

The conv is 1.5 to 3.6% faster and attention 0.5 to 0.9% faster. The gated delta rule is steadily about 0.1 µs (0.3%) slower at 107 tokens (35.70 to 35.77 against 35.59 to 35.71). At 936 and 3,399 tokens its medians are 4.3% and 0.9% higher, inside run-to-run ranges that overlap between the builds.

Demo / evidence

No demo: nothing changes for a worker yet. Kernel test output, comparison logs, timing output and scripts are kept locally.

Self-review

  • I have reviewed the full diff and addressed the issues I found.
  • I have checked that the change follows the project's architecture and stays focused on the stated purpose.
  • I have run the checks appropriate to this change and reported commands, results, and anything I could not verify above.
  • I have checked that the PR description, documentation, and any accuracy or performance claims match the implementation and available evidence.

Add CUDA operations that continue a sequence after a shared prefix,
for request-local prefix reuse (ThinkFlowLab#85): the Gated DeltaNet conv with the
inputs of the three positions before it, the chunked prefill from and
to a float32 state, gated attention for queries after cached keys, and
a pitched device-to-device row copy. The existing conv, prefill and
gated attention are these operations without history, state or cached
positions. Bump the ABI to 6 and add the Rust bindings and GPU tests
against the unsplit calls.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant