Repository navigation
Conversation
Add CUDA operations that continue a sequence after a shared prefix, for request-local prefix reuse (ThinkFlowLab#85): the Gated DeltaNet conv with the inputs of the three positions before it, the chunked prefill from and to a float32 state, gated attention for queries after cached keys, and a pitched device-to-device row copy. The existing conv, prefill and gated attention are these operations without history, state or cached positions. Bump the ABI to 6 and add the Rust bindings and GPU tests against the unsplit calls.
cuBLASLt's heuristic picks an algorithm per M, so a row's result can change with the number of rows in the call, and a prefix and its branch run separately round differently from one pass over the same tokens. Add cs1_gemm_create_fixed: one algorithm per weight shape (N, K, ldy), the heuristic's first choice at a reference M among algorithms without split-K, for every M. The default handle is unchanged. Bump the ABI to 7 and add a GPU test that each projection's rows match across M.
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Purpose
For #85: a backend step the posted plan didn't include, explained in this update. It builds on #97 and stays a draft until that is merged and the backend path is settled with #96. This PR's own change is the last commit,
f095f07.cuBLASLt's heuristic picks an algorithm per M, so a row's result depends on how many rows are in the call. In the 9B fused GDN input projection (synthetic inputs), row 0 differs in 6,768 of 12,352 outputs between M = 1 and M = 2. With prefix reuse, a prefix and its branch run as separate GEMMs, so they round differently from one pass over the same tokens. In a first shared-prefix executor, that alone moved the final hidden state by up to 1.1e-2 relative L2 against
forwardon 9B, above the plan's 1e-2 gate. That held even with chunk-aligned splits, where the conv, the gated delta rule and attention are bit-identical.This PR adds
cs1_gemm_create_fixed(workspace_bytes, reference_m). The handle keeps one algorithm per weight shape (N, K, ldy) for every M: the heuristic's first choice atreference_mrows, among algorithms without split-K. Each output row then takes the same path whatever the other rows, so its result depends neither on M nor on where the row sits. An M the algorithm can't serve returns a cuBLAS status instead of switching algorithms.The fixed handle moves the result away from the default handle's: a full pass with it differs from
forwardby up to 1.5e-2 relative L2 on 9B (#99). That's why #99 compares reuse with a full pass under the same handle, and why the update asks how to gate reuse against today's output.cs1_gemm_createis unchanged, and so are the workers: nothing creates a fixed handle yet. #99 uses it only for the shared-prefix runs, so the plain forward keeps today's algorithms and speed. On a fixed handle, #64's vision GEMMs (FP32, and BF16 with bias) also keep one algorithm per shape; nothing uses them that way, and the test coverscs1_gemmonly. The ABI moves from 6 to 7, including the note insrc/backends/cuda/contract.md.Test Plan
System1-Omni Version / Commit:
f095f07on #97 (4678e3d), onmain47eff9c. The kernel tests and the plain-forward comparison ran on the final code, in #99's tree. The timing ran at an earlier version, before the rebase onto #64. For these BF16 GEMMs it chose algorithms the same way and lacked only the split-K check.CPU: the five Rust steps in CI (format, Clippy, frontend tests, workspace tests, release build), and
mkdocs build --strict.GPU, one RTX 6000 Ada (sm_89), CUDA 13.2, driver 595.91.07:
main: Cua-S1 and Open-Jev-9B workers side by side.Test Result
CPU. All checks pass on macOS.
Kernel tests. 11 of 11 pass, the new one for all 15 shapes.
Plain forward unchanged, with the final code in #99's tree against
main47eff9c: Cua-S1 567 of 567 byte-identical (134 of them forward passes) in eager and CUDA Graph modes, and Open-Jev-9B 253 of 253 identical apart from timing.GEMM time (microseconds per call, default / fixed; mean of two rounds, each the median of three runs):
4B and 27B
Without split-K, the projections with small N take the biggest losses at small M:
At large M, the 9B gate and up projection is 22 to 28% slower from 300 rows, and the 9B and 27B down projections are 42% and 26% slower at 3,109. Other cells are faster, such as the 9B GDN input, 23 to 30% faster from 32 to 300 rows. That mix is why the fixed handle isn't the default. In a whole 9B pass, it costs 8% at 1,024 tokens and 18% at 3,072 (#99).
Demo / evidence
No demo: nothing changes for a worker. Timing output and scripts are kept locally.
Self-review