Conversation
PTQ1_0 decode on Vulkan went through the generic mul_mat_vec shader, which decodes every trit with a per-element base-3 loop; on an RX 6750 XT that ran at ~270 GFLOPS and limited Ternary-Bonsai-2-27B to 4.7 tok/s. This adds a dedicated q8_1 (mmvq) shader for PTQ1_0 and wires PQ2_0 into the Vulkan backend, which had no support for it at all. mul_mat_vecq_ptq1_0.comp: - two lanes per 128-element block, 4 rows per workgroup, no divergence - trits decoded with the multiply-by-3 recurrence on two bytes at a time in the 16-bit halves of a dword (as the CUDA vec_dot_ptq1_0_q8_1 does) - the level word is fed to dotPacked4x8AccSat as is: activations are pre-masked so the residue bytes are multiplied by zero, which removes the per-level extraction ops - the ternary -1 offset is folded into the q8_1 block sums PQ2_0: block/packed16 types, dequant shader, mul_mm loader, float and q8_1 mat-vec paths (modeled on Q2_0 with the 128-element group), get_rows. test-backend-ops: MUL_MAT / MUL_MAT_ID cases for both types at the Bonsai shapes (k = 1024..17408), odd row counts, batched B and n = 1..8, plus perf cases at Bonsai shapes. Measured on RX 6750 XT (RDNA2, AMD proprietary driver, Windows 11), Ternary-Bonsai-2-27B: PTQ1_0 decode 4.7 -> 40.8 tok/s (39 with -fa on), PQ2_0 CPU-only -> 38.5 tok/s, prefill 200 t/s. Perplexity through the new path 6.3916 vs 6.3906 with the original shader; greedy output identical to the stock build.
The PTQ1_0 q8_1 mat-vec re-ran the whole trit decode for every column of B, so batched decode (llama-server with several slots, NUM_COLS = 2..8) scaled poorly: on an RX 6750 XT 8 parallel sequences gave 66 tok/s aggregate, barely more than 4. Columns are now handled in passes of CG = 3 (a new specialization constant). Per block the lane decodes all NUM_ROWS rows level by level once and dots each packed level with the activations of every column of the pass, so an extra column costs three loads and three dot products per row and level. Integer accumulators are folded into the fp32 result whenever the level moves on to a new q8_1 sub-block scale. Passes are separate loops over the blocks, written as guarded calls, because the driver hoists the loads of everything in one block body (which spilled at 8 columns) and does not unroll a loop that contains the block loop. The single-column pipeline keeps the original loop (runtime-bounded row loop, 72 VGPRs), so single-sequence decode is unchanged. RX 6750 XT, Ternary Bonsai 2 27B, llama-batched-bench pp256/tg128, fa on, q8_0 KV, aggregate decode tok/s (before -> after): 1 seq 36.6 -> 38.1 4 seq 55.2 -> 73.3 2 seq 39.7 -> 60.8 8 seq 66.3 -> 86.4 llama-bench tg64: 40.6 (fa off) / 38.7 (fa on), same as before. Kernel time m=17408 k=5120: n=1 45 us, n=2 57, n=4 120, n=8 215. test-backend-ops: 271/271 (ptq1_0 + pq2_0, forced mmvq, n = 1..8 incl. row tail), 138/138 default path; perplexity through the 8-column path 6.2321 vs 6.2294 before (same text, fp32 summation order). Also adds perf cases with n = 2, 4, 8 for PTQ1_0 / PQ2_0 / Q4_0.
…ments - ggml_vk_get_dequantize_mul_mat_vec_id: add PTQ1_0 and PQ2_0 to the q8_1 whitelist. The id q8_1 pipelines were created but never selected, so MUL_MAT_ID always fell back to the float path. - test-backend-ops: add n = 4 to the PTQ1_0/PQ2_0 cases (a three-column pass followed by the single-column tail). - Shorten the shader comments and the coopmat2 note in vulkan-shaders-gen.cpp. test-backend-ops MUL_MAT/MUL_MAT_ID/GET_ROWS for ptq1_0 + pq2_0: 284/284 on the default path and 284/284 with GGML_VK_FORCE_MMVQ=1 (RX 6750 XT); the mul_mat_vec_id_ptq1_0_q8_1_f32 pipeline is now compiled and used.
BC-250 (gfx1013) is RDNA1: no coopmat, no accelerated integer dot. For the PQ2_0 packed-ternary format the decode path runs mul_mat_vec_pq2_0_*; the stock config used 2 rows/workgroup and let the large DMMV variant run at 4*subgroup=128 threads, underfilling the GPU on single-row decode. Gated on vendor==AMD && arch==RDNA1 && deviceID==0x13fe && !coopmat so no other device changes behavior: 4 rows/workgroup, workgroup size 32 on both DMMV variants, required subgroup size 32; RDNA1 preferred mul_mat_vec subgroup 64->32. f16-dst and all other quantized mmv paths unchanged. Measured (llama-bench, Ternary-Bonsai-2-27B-PQ2_0, 2000 MHz, -fa1 KV q4_1): tg128 22.43->23.93 t/s (+6.7%); pp512 unchanged (noise). PPL identical to last digit (corpus-ko, -c 512 --chunks 8 -b 1024). test-backend-ops MUL_MAT/MUL_MAT_ID/GET_ROWS 2310/2310 pass. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Negative control from a non-BC-250 device: the gate holds — Intel Arc B390 (Xe3 iGPU), Vulkan / WindowsRan this branch against its own base (#188) on hardware the gate is meant to exclude, to check the "no other device sees a behavioral difference" claim. It holds on this part. Device: Builds: #188 @ All four conjuncts of
|
#188 80f71b004 |
#229 d6d51ddd4 |
|
|---|---|---|
| OK | 17434 | 17434 |
| not supported | 4524 | 4524 |
| failing cases | 3 | 3 (same cases) |
| PTQ1_0 | 199 OK / 180 n-s | 199 OK / 180 n-s |
| PQ2_0 | 191 OK / 180 n-s | 191 OK / 180 n-s |
Every counter matches. (The 3 failures are the pre-existing GATED_DELTA_NET raw_gates=1 bug on prism, filed separately as #237 — unrelated to either PR.)
Throughput, Bonsai 2 27B, ABAB-interleaved
-p 512 -n 128 -ngl 99 -fa 1 -r 3, default dispatch (deliberately not GGML_VK_FORCE_MMVQ — the change is in the PQ2_0 dequant mat-vec, so the default path is where any effect would land):
| model | metric | #229 (r1 / r2) | #188 (r1 / r2) |
|---|---|---|---|
| PQ2_0 | tg128 | 10.99 / 11.09 | 11.07 / 10.93 |
| PQ2_0 | pp512 | 273.4 / 279.6 | 266.6 / 274.3 |
| PTQ1_0 | tg128 | 1.66 / 1.56 | 1.56 / 1.56 |
| PTQ1_0 | pp512 | 207.5 / 179.5 | 177.2 / 174.6 |
PQ2_0 decode — the path this PR changes — differs by 0.4% between arms, inside the <1% within-run spread, ranges overlapping. The lone 1.66 on PTQ1_0 r1 is a cold-start artifact (first load after an idle GPU); r2 agrees exactly. pp512 carries ±10-38 on both arms here and does not discriminate.
Conclusion: no measurable difference on a device outside the gate, as intended.
One scoping note
The rdna1_pipelines entry is the one change not covered by the deviceID == 0x13fe gate:
- {"argmax", 64}, {"mul_mat_vec", 64},
+ {"argmax", 64}, {"mul_mat_vec", 32},That table is keyed on RDNA1 rather than on BC-250, so it reaches every RDNA1 part, not only gfx1013. The description notes the reasoning (required_subgroup_size wins wherever subgroups are usable), so this may well be intended — but it does mean the PR body's "All changes are gated on vendor == AMD && architecture == AMD_RDNA1 && deviceID == 0x13fe" is slightly stronger than the diff. Worth stating the intended scope explicitly, since the case it would affect is an RDNA1 device that is not a BC-250 and where subgroups are not usable. I have no RDNA1 hardware, so nothing above tests that path — this result speaks only for the Intel side.
|
Follow-up from our own live BC-250 (gfx1013) service with this change deployed: on a mixed-length workload with MTP on (6 prompts, ctx 64K shared, 2 slots), generation was at parity with the previous build (33.03 vs 33.04 t/s) and output quality was unchanged. So the +6.7% is a |
|
Superseded by #281, which is the cleaned-up follow-up (single commit, dedicated PQ2_0 mat-vec + BC-250 column extension + per-column rows table). This PR has diverged from the current prism base (mergeable=false) after the upstream mat-vec restructuring, and its content is carried by #281. Closing in favor of #281. |
vulkan: tune the PQ2_0 mat-vec for AMD BC-250 (gfx1013, RDNA1)
Problem
AMD BC-250 (PCI
1002:13FE, gfx1013) is an RDNA1 part: no cooperative-matrixsupport (
matrix cores: none) and no accelerated integer dot product(
int dot: 0). For the Bonsai-2PQ2_0packed-ternary format the decode pathruns the
mul_mat_vec_pq2_0_*dequant kernels. The stock launch config used2*rm_stdq = 2rows per workgroup and let theDMMV_WG_SIZE_LARGEvariant runat
4*subgroup = 128threads, which underfills the GPU on single-row decode.Change (narrow gate: PCI 1002:13FE only)
All changes are gated on
vendor == AMD && architecture == AMD_RDNA1 && deviceID == 0x13fe && !coopmat_support,so no other device sees a behavioral difference:
mul_mat_vec_pq2_0_f32_f32: 4 rows per workgroup (was 2), workgroup size 32on both DMMV variants (was 32 / 128), required subgroup size pinned to 32.
rdna1_pipelines: preferred"mul_mat_vec"subgroup size 64 -> 32(RDNA1-scoped table; the explicit
required_subgroup_sizeargument alreadywins wherever subgroups are usable).
The f16 destination variant and every other quantized mmv path are unchanged.
The E3 mmq m-warptile variant was measured in the same sweep series and was
dropped — it did not reproduce a win.
Measurements (BC-250, RADV GFX1013, performance mode 2000 MHz)
llama-benchonTernary-Bonsai-2-27B-PQ2_0.gguf(27B),-p 512 -n 128 -r 3 -fa 1 -ctk q4_1 -ctv q4_1, warmup + median of 3, alternating.soswaps:Earlier quiet-mode (1500 MHz) sweep on the same model: tg 17.83 -> 18.71 t/s
(+4.93%), which is what selected rows=4 + subgroup 32 over r1/r2 and sg64.
Production check (2026-09-21): the same build now serves the PQ2_0-MTP GGUF as
the live server (ctx 65,536, 2 slots,
--spec-type draft-mtp): sustained30.2-30.3 t/s single-stream decode, 54,003-token prefill at ~80 t/s, zero
DeviceLost/OOM — consistent with the mmv-level gain.
Correctness
test-backend-ops -b Vulkan0 -o MUL_MAT,MUL_MAT_ID,GET_ROWS: 2310/2310tests pass, 0 FAIL, including newly added Bonsai-2 real shapes
(17408x5120, 5120x17408, 10240x5120, 5120x6144, 6144x5120) and odd
m/n tails (67, 70) for both PTQ1_0 and PQ2_0.
llama-perplexity -c 512 --chunks 8 -b 1024(corpus-ko): PPL 51.0904 basevs 51.0904 r4sg32 — identical to the last digit, far inside the ±0.002 gate.
Notes for review
(pp delta is measurement noise).
inverse-rotation fix, which is a separate model-graph change not included
in this PR.