Repository navigation
Conversation
…t CTAs Every PTQ1_0 mat-vec in the decode graph pays about 2.5 us of ramp-up and ramp-down in which the DRAM is not busy. The node loop now looks ahead for the next PTQ1_0 MUL_MAT (a gate/up pair feeding one GLU counts as one op) and hands its weight pointer to the launcher. The last 46 CTAs of the running kernel issue prefetch.global.L2 for the first 50 % (at most 16 MiB) of those weights right after their own loads. More CTAs or more bytes compete with the running kernel and get slower (184 CTAs: -2.3 %). Values are not changed. GGML_CUDA_L2_PREFETCH_PCT=0 turns it off; _CTAS and _MAX_KB tune it. Decode with MTP n-max 2, same binary, outputs identical: +2.00 % (greedy benchmark), +2.10 % (Hermes setup, ctx 114688), +2.2 / +2.1 / +1.7 % at depth 0 / 16k / 65k. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
GGML_CUDA_L2_PREFETCH_PCT, _MAX_KB and _CTAS of the previous commit were only there to tune and measure the change. The measured values (50 % of the next tensor, at most 16 MiB, 46 CTAs) are now constants. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
|
Turing data point, since the final commit hardcodes the 4070 tuning (50% / 46 CTAs / 16 MiB): on a 4 MB-L2 card it is a loss. RTX 2060 SUPER (sm_75, 4 MB L2, 34 SMs), stock clocks, Bonsai 2 27B PTQ1_0, q4_0 K/V, no MTP head, greedy, 2 prompts x 400 tokens per arm, server restarted per arm, three alternated runs per setting (first commit
Round-to-round spread with prefetch off was under 0.1 tok/s (~0.2%). Greedy output byte-identical in every arm, as expected for a prefetch. 2048 KB and 2048 KB with 34 CTAs landed between (+0.7..1.4%). So the byte cap wants to scale with the device: a quarter of |
The cap of 16 MiB and the fixed 46 CTAs were tuned on an RTX 4070 only. On an RTX 2060 SUPER (4 MiB L2, 34 SMs) that setting is 2.4 to 2.7 % slower than no prefetch (measured by a reviewer, see the PR). Prefetch with one CTA per SM, and cap the bytes at 8 MiB on Ada (cc 8.9) and at 2 MiB on every other GPU. RTX 4070 (36 MiB L2), MTP n-max 2, same outputs in all runs, 12 pairs against the previous setting: 8 MiB +0.02 % (n.s.), 16 MiB -0.03 % (n.s.), 4 MiB -0.17 %, 2 MiB -0.40 %. The 2 MiB default is not measured on other GPUs. The 4070 has 46 SMs, so one CTA per SM is the old count.
|
@professorpalmer your 2060 numbers are right and it was my mistake: I tuned the 16 MiB cap and the 46 CTAs on the 4070 only and never tried a card with a small L2. Also, the description says "48 MB L2" for the 4070, it is 36 MiB. I repeated the tests on the 4070 (full kTrain stack, MTP n-max 2, depth 0, 12 pairs each against the old setting, outputs identical in every run). Byte cap with one prefetching CTA per SM:
So on this card 8 and 16 MiB are the same and smaller caps lose a little. On the PR branch alone, small caps looked slightly better (2 MiB +0.48 % against 16 MiB); I do not know why the sign flips on the full stack, so I trust the stack numbers. The best cap depends on the GPU (4070: 36 MiB L2, 2060 SUPER: 4 MiB), and I can only measure the 4070. So the new commit uses 8 MiB only on Ada (cc 8.9) and 2 MiB on every other GPU, with one prefetching CTA per SM (46 on the 4070, so no change there). The 2 MiB is your number and not measured by me. With this commit the 4070 is +0.02 % against the old setting (12 pairs, n.s.), The measurements and scripts were run with Claude Code (AI assistance); I can post the exact commands if useful. |
Overview
Every PTQ1_0 mat-vec in the decode graph pays about 2.5 us of ramp-up and ramp-down in which DRAM is idle (399 mat-vecs per verify step). The CUDA node loop now looks ahead for the next PTQ1_0
MUL_MAT(a gate/up pair that feeds one GLU counts as one op), stores its weight pointer and size in a thread-local hint (l2-hint.cuh), and the launcher ofmul_mat_vec_ptq1_0_ptpasses it to the kernel. The last 46 CTAs issueprefetch.global.L2for the first 50 % (at most 16 MiB) of those weights right after their own loads. Values are not changed.The numbers matter: more CTAs or more bytes compete with the running kernel for DRAM and are slower (184 CTAs: -2.3 %; 60 % from 92 CTAs: +0.6 %; 50 % from 46 CTAs: +2.0 %). RTX 4070, Bonsai 2 27B PTQ1_0 + MTP head (n-max 2), 8 interleaved A/B pairs of the same binary, outputs identical in all runs: greedy benchmark 108.31 -> 110.50 tok/s (+2.00 %, CI [+1.92, +2.08] %), agent-like setup (context 114688, thinking sampling, reasoning budget) 101.11 -> 103.22 tok/s (+2.10 %, CI [+2.04, +2.13] %); depth 0 / 16k / 65k +2.2 / +2.1 / +1.7 % (one run each, identical output hashes).
Additional information
GGML_CUDA_L2_PREFETCH_PCT(=0turns it off),_CTASand_MAX_KB(defaults 50 / 46 / 16384) that were used for the tuning and measurements below, the last one fixes the measured values as constants and removes the switches. To reproduce a measurement, build the first commit.ptxas:griddepcontrolrequires sm_90).ggml_cuda_kernel_launchalready opts into PDL there, so the idle window this PR fills may already be covered on those GPUs; I did not test whether the prefetch helps or hurts there, the default applies to all GPUs, and I can restrict it to cc 8.9.thread_localset by the node loop before each node and read by the launcher; the kernel receives the pointer, the number of lines per CTA and the number of prefetching CTAs as three extra arguments. The pointer and byte count come from the next PTQ1_0 weight tensor (at most 50 % of it, at most 16 MiB, so never beyond the tensor). The "last CTAs" are assumed to be the ones that run last, which the hardware does in practice but does not guarantee; if not, only the speed-up shrinks.Test results
-DGGML_CUDA=ON -DCMAKE_CUDA_ARCHITECTURES=89.prismat88c4bc60bplus my other patches (this PR is independent of them in code); the four commits since (SYCL, WebGPU andcuda: fused FWHT quantizer for 64-wide warps (#303)) do not touch these code paths. The branch is rebased on2459f68b5, builds, andtest-backend-opswas repeated on it.--spec-type draft-mtp --spec-draft-n-max 2, q4_0 K/V cache, one slot.GGML_CUDA_L2_PREFETCH_PCT=0against the default, first commit of the branch) flipped through an environment variable, 4 greedy prompts x 256 tokens, alternating runs A B B A, 8 pairs, paired differences with a 95 % bootstrap interval; run-to-run noise about 0.04 to 0.12 %. The baseline drifts by up to ~1 % between sessions, only paired numbers count.--reasoning-format deepseek --reasoning-budget 16384, thinking sampling 1.0 / 0.95 / 20 / 0.05, fixed seed) 101.11 -> 103.22 tok/s, +2.10 % (CI [+2.04, +2.13] %). Outputs identical in all 16 runs, acceptance unchanged.test-backend-ops test -b CUDA0 -o MUL_MAT,MUL_MAT_VEC_FUSION,GLU,SWIGLU: 2116/2116 passed on the rebased branch (CUDA0 against CPU). For the stack with all my patchesMUL_MAT, MUL_MAT_ID, MUL_MAT_VEC_FUSION, GLU, SWIGLU, RMS_NORM, FLASH_ATTN_EXT, CONCAT: 6373/6373.data), several slots, models other than PTQ1_0 (only this kernel reads the hint).Requirements