Skip to content

cuda: prefetch the next PTQ1_0 mat-vec's weights into L2 from the last CTAs - #314

Open
sb32445 wants to merge 3 commits into
PrismML-Eng:prismfrom
sb32445:pr/ptq1-l2-prefetch
Open

sb32445 wants to merge 3 commits into
PrismML-Eng:prismfrom
sb32445:pr/ptq1-l2-prefetch

Conversation

@sb32445

@sb32445 sb32445 commented Oct 4, 2026

Copy link
Copy Markdown

Overview

Every PTQ1_0 mat-vec in the decode graph pays about 2.5 us of ramp-up and ramp-down in which DRAM is idle (399 mat-vecs per verify step). The CUDA node loop now looks ahead for the next PTQ1_0 MUL_MAT (a gate/up pair that feeds one GLU counts as one op), stores its weight pointer and size in a thread-local hint (l2-hint.cuh), and the launcher of mul_mat_vec_ptq1_0_pt passes it to the kernel. The last 46 CTAs issue prefetch.global.L2 for the first 50 % (at most 16 MiB) of those weights right after their own loads. Values are not changed.

The numbers matter: more CTAs or more bytes compete with the running kernel for DRAM and are slower (184 CTAs: -2.3 %; 60 % from 92 CTAs: +0.6 %; 50 % from 46 CTAs: +2.0 %). RTX 4070, Bonsai 2 27B PTQ1_0 + MTP head (n-max 2), 8 interleaved A/B pairs of the same binary, outputs identical in all runs: greedy benchmark 108.31 -> 110.50 tok/s (+2.00 %, CI [+1.92, +2.08] %), agent-like setup (context 114688, thinking sampling, reasoning budget) 101.11 -> 103.22 tok/s (+2.10 %, CI [+2.04, +2.13] %); depth 0 / 16k / 65k +2.2 / +2.1 / +1.7 % (one run each, identical output hashes).

Additional information

  • The branch has two commits: the first contains the environment switches GGML_CUDA_L2_PREFETCH_PCT (=0 turns it off), _CTAS and _MAX_KB (defaults 50 / 46 / 16384) that were used for the tuning and measurements below, the last one fixes the measured values as constants and removes the switches. To reproduce a measurement, build the first commit.
  • Tuned on one GPU (RTX 4070, 46 SMs, 48 MB L2). Other GPUs may want other values; I am happy to gate it behind cc 8.9 or default it to off if you prefer.
  • A micro benchmark with a chain of streaming kernels showed the same pattern (1-1.6 us saved per kernel, DRAM saturates at ~94 % of peak).
  • On Hopper and newer, programmatic dependent launch would be the cleaner tool; it does not exist on cc 8.9 (ptxas: griddepcontrol requires sm_90). ggml_cuda_kernel_launch already opts into PDL there, so the idle window this PR fills may already be covered on those GPUs; I did not test whether the prefetch helps or hurts there, the default applies to all GPUs, and I can restrict it to cc 8.9.
  • The hint is a thread_local set by the node loop before each node and read by the launcher; the kernel receives the pointer, the number of lines per CTA and the number of prefetching CTAs as three extra arguments. The pointer and byte count come from the next PTQ1_0 weight tensor (at most 50 % of it, at most 16 MiB, so never beyond the tensor). The "last CTAs" are assumed to be the ones that run last, which the hardware does in practice but does not guarantee; if not, only the speed-up shrinks.
  • Weights are not modified, no value changes; outputs are identical.
  • test-backend-ops MUL_MAT, MUL_MAT_ID, MUL_MAT_VEC_FUSION, GLU, SWIGLU, RMS_NORM, FLASH_ATTN_EXT, CONCAT: 6373/6373 on CUDA0 for the full stack.

Test results

  • Hardware / software: RTX 4070 12 GB (AD104, cc 8.9, 46 SMs, 504 GB/s, 48 MB L2), Linux 6.18, NVIDIA driver 615.71, CUDA 13.4, GCC 16.2; Release build, -DGGML_CUDA=ON -DCMAKE_CUDA_ARCHITECTURES=89.
  • Base: the speed numbers were measured on prism at 88c4bc60b plus my other patches (this PR is independent of them in code); the four commits since (SYCL, WebGPU and cuda: fused FWHT quantizer for 64-wide warps (#303)) do not touch these code paths. The branch is rebased on 2459f68b5, builds, and test-backend-ops was repeated on it.
  • Model: Ternary-Bonsai-2-27B (PTQ1_0) with a community MTP draft head (Q8_0), --spec-type draft-mtp --spec-draft-n-max 2, q4_0 K/V cache, one slot.
  • Method: the same binary, one switch (GGML_CUDA_L2_PREFETCH_PCT=0 against the default, first commit of the branch) flipped through an environment variable, 4 greedy prompts x 256 tokens, alternating runs A B B A, 8 pairs, paired differences with a 95 % bootstrap interval; run-to-run noise about 0.04 to 0.12 %. The baseline drifts by up to ~1 % between sessions, only paired numbers count.
  • Tuning, greedy benchmark, 4 pairs per setting (prefetched fraction of the next weights / number of prefetching CTAs): 60 % / 92: +0.6 % (6 pairs); 100 % / 92: +0.14 % (n.s.); 60 % / 184: -2.27 %; 30 % / 92: +1.14 %; 20 % / 92: +1.19 %; 15 % / 92: +1.54 %; 30 % / 138: +0.07 % (n.s.); 30 % / 46: +1.80 %; 15 % / 46: +1.61 %; 30 % / 23: +1.76 %; 15 % / 23: +1.48 %; 30 % / 12: +1.39 %; 50 % / 46: +2.02 %; 70 % / 46: +2.00 %; 100 % / 46: +1.78 %; 50 % / 69: +1.71 %; 70 % / 34: +1.74 %. Outputs identical in all of them.
  • Confirmation with the defaults (50 % / 46 CTAs / 16 MiB), 8 pairs each: greedy benchmark 108.31 -> 110.50 tok/s, +2.00 % (95 % CI [+1.92, +2.08] %); agent-like setup (context 114688, --reasoning-format deepseek --reasoning-budget 16384, thinking sampling 1.0 / 0.95 / 20 / 0.05, fixed seed) 101.11 -> 103.22 tok/s, +2.10 % (CI [+2.04, +2.13] %). Outputs identical in all 16 runs, acceptance unchanged.
  • By depth (one run each, so only indicative; the output hashes are identical): depth 0 / 16384 / 65536: 119.6 -> 122.2 (+2.2 %), 109.0 -> 111.3 (+2.1 %), 92.6 -> 94.2 (+1.7 %) tok/s.
  • Micro benchmark (chain of dependent streaming kernels, not part of the PR): 1.0 to 1.6 us saved per kernel for 6.9 to 39 MB of weights, DRAM saturating at ~94 % of peak with prefetch, the same pattern of "fewer prefetching CTAs is better".
  • test-backend-ops test -b CUDA0 -o MUL_MAT,MUL_MAT_VEC_FUSION,GLU,SWIGLU: 2116/2116 passed on the rebased branch (CUDA0 against CPU). For the stack with all my patches MUL_MAT, MUL_MAT_ID, MUL_MAT_VEC_FUSION, GLU, SWIGLU, RMS_NORM, FLASH_ATTN_EXT, CONCAT: 6373/6373.
  • Not tested: other GPUs, multiple GPUs / tensor split, weights in host memory or unified memory (the prefetch addresses come from the weight tensor's data), several slots, models other than PTQ1_0 (only this kernel reads the hint).

Requirements

  • I have read and agree with the contributing guidelines
  • AI usage disclosure: The patches were developed with Claude Code (Anthropic's coding agent): it wrote the code, the measurement scripts and the first drafts of the commit messages and PR texts. I decided what to work on (which kernels and host paths to optimise, based on profiles of my own decode setup). The measurements and checks listed in the PR texts were run in the Claude Code sessions; I did not re-run them independently. I will maintain the changes. Commits where Claude Code was used carry a Co-Authored-By trailer.

sb32445 and others added 2 commits October 4, 2026 14:21
…t CTAs

Every PTQ1_0 mat-vec in the decode graph pays about 2.5 us of ramp-up
and ramp-down in which the DRAM is not busy. The node loop now looks
ahead for the next PTQ1_0 MUL_MAT (a gate/up pair feeding one GLU counts
as one op) and hands its weight pointer to the launcher. The last 46
CTAs of the running kernel issue prefetch.global.L2 for the first 50 %
(at most 16 MiB) of those weights right after their own loads. More
CTAs or more bytes compete with the running kernel and get slower
(184 CTAs: -2.3 %). Values are not changed.

GGML_CUDA_L2_PREFETCH_PCT=0 turns it off; _CTAS and _MAX_KB tune it.

Decode with MTP n-max 2, same binary, outputs identical: +2.00 %
(greedy benchmark), +2.10 % (Hermes setup, ctx 114688), +2.2 / +2.1 /
+1.7 % at depth 0 / 16k / 65k.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
GGML_CUDA_L2_PREFETCH_PCT, _MAX_KB and _CTAS of the previous commit were only
there to tune and measure the change. The measured values (50 % of the next
tensor, at most 16 MiB, 46 CTAs) are now constants.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
@professorpalmer

Copy link
Copy Markdown

Turing data point, since the final commit hardcodes the 4070 tuning (50% / 46 CTAs / 16 MiB): on a 4 MB-L2 card it is a loss.

RTX 2060 SUPER (sm_75, 4 MB L2, 34 SMs), stock clocks, Bonsai 2 27B PTQ1_0, q4_0 K/V, no MTP head, greedy, 2 prompts x 400 tokens per arm, server restarted per arm, three alternated runs per setting (first commit e3d9df5, so the switches were available; plus the one-column PT fix from #325):

setting 4k decode vs prefetch off 16k decode vs prefetch off
50% / 46 CTAs / 16 MiB (as in a4249f5) -2.4 / -2.7 / -2.7% -2.4 / -2.6 / -1.1%
GGML_CUDA_L2_PREFETCH_MAX_KB=1024 +1.4 / +0.6 / +2.2% +1.2 / +0.7 / +2.2%
GGML_CUDA_L2_PREFETCH_MAX_KB=512 +0.6 / +2.0% +0.6 / +1.8%

Round-to-round spread with prefetch off was under 0.1 tok/s (~0.2%). Greedy output byte-identical in every arm, as expected for a prefetch. 2048 KB and 2048 KB with 34 CTAs landed between (+0.7..1.4%).

So the byte cap wants to scale with the device: a quarter of l2CacheSize (1 MiB here) works on Turing; I have not measured what that rule does on the 4070, where you found 16 MiB best.

The cap of 16 MiB and the fixed 46 CTAs were tuned on an RTX 4070 only.
On an RTX 2060 SUPER (4 MiB L2, 34 SMs) that setting is 2.4 to 2.7 %
slower than no prefetch (measured by a reviewer, see the PR). Prefetch
with one CTA per SM, and cap the bytes at 8 MiB on Ada (cc 8.9) and at
2 MiB on every other GPU.

RTX 4070 (36 MiB L2), MTP n-max 2, same outputs in all runs, 12 pairs
against the previous setting: 8 MiB +0.02 % (n.s.), 16 MiB -0.03 %
(n.s.), 4 MiB -0.17 %, 2 MiB -0.40 %. The 2 MiB default is not measured
on other GPUs. The 4070 has 46 SMs, so one CTA per SM is the old count.
@sb32445

sb32445 commented Oct 7, 2026

Copy link
Copy Markdown
Author

@professorpalmer your 2060 numbers are right and it was my mistake: I tuned the 16 MiB cap and the 46 CTAs on the 4070 only and never tried a card with a small L2. Also, the description says "48 MB L2" for the 4070, it is 36 MiB.

I repeated the tests on the 4070 (full kTrain stack, MTP n-max 2, depth 0, 12 pairs each against the old setting, outputs identical in every run). Byte cap with one prefetching CTA per SM:

cap change 95 % CI
2 MiB -0.40 % [-0.45, -0.36]
4 MiB -0.17 % [-0.23, -0.12]
8 MiB +0.02 % [-0.04, +0.09] (n.s.)
16 MiB -0.03 % [-0.07, +0.01] (n.s.)

So on this card 8 and 16 MiB are the same and smaller caps lose a little. On the PR branch alone, small caps looked slightly better (2 MiB +0.48 % against 16 MiB); I do not know why the sign flips on the full stack, so I trust the stack numbers.

The best cap depends on the GPU (4070: 36 MiB L2, 2060 SUPER: 4 MiB), and I can only measure the 4070. So the new commit uses 8 MiB only on Ada (cc 8.9) and 2 MiB on every other GPU, with one prefetching CTA per SM (46 on the 4070, so no change there). The 2 MiB is your number and not measured by me. With this commit the 4070 is +0.02 % against the old setting (12 pairs, n.s.), test-backend-ops MUL_MAT, MUL_MAT_VEC_FUSION, GLU, SWIGLU 2116/2116. Could you run your 2060 test again? If 2 MiB is still not a gain there, I will turn the prefetch off outside Ada.

The measurements and scripts were run with Claude Code (AI assistance); I can post the exact commands if useful.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants