Skip to content

vulkan: tune the PQ2_0 mat-vec for AMD BC-250 (gfx1013, RDNA1) - #229

Closed
renovys wants to merge 5 commits into
PrismML-Eng:prismfrom
renovys:bc250-pq2-mmv-tune
Closed

renovys wants to merge 5 commits into
PrismML-Eng:prismfrom
renovys:bc250-pq2-mmv-tune

Conversation

@renovys

@renovys renovys commented Sep 21, 2026 •

Copy link
Copy Markdown

Based on #188 (PQ2_0 support), which is not yet merged, so the diff below #188 shows its commits too; the net new change here is the final commit. Happy to rebase once #188 lands.

vulkan: tune the PQ2_0 mat-vec for AMD BC-250 (gfx1013, RDNA1)

Problem

AMD BC-250 (PCI 1002:13FE, gfx1013) is an RDNA1 part: no cooperative-matrix
support (matrix cores: none) and no accelerated integer dot product
(int dot: 0). For the Bonsai-2 PQ2_0 packed-ternary format the decode path
runs the mul_mat_vec_pq2_0_* dequant kernels. The stock launch config used
2*rm_stdq = 2 rows per workgroup and let the DMMV_WG_SIZE_LARGE variant run
at 4*subgroup = 128 threads, which underfills the GPU on single-row decode.

Change (narrow gate: PCI 1002:13FE only)

All changes are gated on
vendor == AMD && architecture == AMD_RDNA1 && deviceID == 0x13fe && !coopmat_support,
so no other device sees a behavioral difference:

  • mul_mat_vec_pq2_0_f32_f32: 4 rows per workgroup (was 2), workgroup size 32
    on both DMMV variants (was 32 / 128), required subgroup size pinned to 32.
  • rdna1_pipelines: preferred "mul_mat_vec" subgroup size 64 -> 32
    (RDNA1-scoped table; the explicit required_subgroup_size argument already
    wins wherever subgroups are usable).

The f16 destination variant and every other quantized mmv path are unchanged.
The E3 mmq m-warptile variant was measured in the same sweep series and was
dropped — it did not reproduce a win.

Measurements (BC-250, RADV GFX1013, performance mode 2000 MHz)

llama-bench on Ternary-Bonsai-2-27B-PQ2_0.gguf (27B), -p 512 -n 128 -r 3 -fa 1 -ctk q4_1 -ctv q4_1, warmup + median of 3, alternating .so swaps:

metric base r4sg32 delta
pp512 142.62 142.73 +0.1% (noise)
tg128 22.43 23.93 +6.7%

Earlier quiet-mode (1500 MHz) sweep on the same model: tg 17.83 -> 18.71 t/s
(+4.93%), which is what selected rows=4 + subgroup 32 over r1/r2 and sg64.

Production check (2026-09-21): the same build now serves the PQ2_0-MTP GGUF as
the live server (ctx 65,536, 2 slots, --spec-type draft-mtp): sustained
30.2-30.3 t/s single-stream decode, 54,003-token prefill at ~80 t/s, zero
DeviceLost/OOM — consistent with the mmv-level gain.

Correctness

  • test-backend-ops -b Vulkan0 -o MUL_MAT,MUL_MAT_ID,GET_ROWS: 2310/2310
    tests pass, 0 FAIL
    , including newly added Bonsai-2 real shapes
    (17408x5120, 5120x17408, 10240x5120, 5120x6144, 6144x5120) and odd
    m/n tails (67, 70) for both PTQ1_0 and PQ2_0.
  • llama-perplexity -c 512 --chunks 8 -b 1024 (corpus-ko): PPL 51.0904 base
    vs 51.0904 r4sg32 — identical to the last digit, far inside the ±0.002 gate.

Notes for review

  • The win is decode-only; pp shapes take the mmq path and are unaffected
    (pp delta is measurement noise).
  • Serving the PQ2_0-MTP model additionally needs the qwen35 MTP embedding
    inverse-rotation fix, which is a separate model-graph change not included
    in this PR.

alhnesn and others added 5 commits September 18, 2026 20:36
PTQ1_0 decode on Vulkan went through the generic mul_mat_vec shader, which
decodes every trit with a per-element base-3 loop; on an RX 6750 XT that ran
at ~270 GFLOPS and limited Ternary-Bonsai-2-27B to 4.7 tok/s. This adds a
dedicated q8_1 (mmvq) shader for PTQ1_0 and wires PQ2_0 into the Vulkan
backend, which had no support for it at all.

mul_mat_vecq_ptq1_0.comp:
- two lanes per 128-element block, 4 rows per workgroup, no divergence
- trits decoded with the multiply-by-3 recurrence on two bytes at a time in
  the 16-bit halves of a dword (as the CUDA vec_dot_ptq1_0_q8_1 does)
- the level word is fed to dotPacked4x8AccSat as is: activations are
  pre-masked so the residue bytes are multiplied by zero, which removes the
  per-level extraction ops
- the ternary -1 offset is folded into the q8_1 block sums

PQ2_0: block/packed16 types, dequant shader, mul_mm loader, float and q8_1
mat-vec paths (modeled on Q2_0 with the 128-element group), get_rows.

test-backend-ops: MUL_MAT / MUL_MAT_ID cases for both types at the Bonsai
shapes (k = 1024..17408), odd row counts, batched B and n = 1..8, plus perf
cases at Bonsai shapes.

Measured on RX 6750 XT (RDNA2, AMD proprietary driver, Windows 11),
Ternary-Bonsai-2-27B: PTQ1_0 decode 4.7 -> 40.8 tok/s (39 with -fa on),
PQ2_0 CPU-only -> 38.5 tok/s, prefill 200 t/s. Perplexity through the new
path 6.3916 vs 6.3906 with the original shader; greedy output identical to
the stock build.
The PTQ1_0 q8_1 mat-vec re-ran the whole trit decode for every column of B,
so batched decode (llama-server with several slots, NUM_COLS = 2..8) scaled
poorly: on an RX 6750 XT 8 parallel sequences gave 66 tok/s aggregate, barely
more than 4.

Columns are now handled in passes of CG = 3 (a new specialization constant).
Per block the lane decodes all NUM_ROWS rows level by level once and dots each
packed level with the activations of every column of the pass, so an extra
column costs three loads and three dot products per row and level. Integer
accumulators are folded into the fp32 result whenever the level moves on to a
new q8_1 sub-block scale. Passes are separate loops over the blocks, written
as guarded calls, because the driver hoists the loads of everything in one
block body (which spilled at 8 columns) and does not unroll a loop that
contains the block loop.

The single-column pipeline keeps the original loop (runtime-bounded row loop,
72 VGPRs), so single-sequence decode is unchanged.

RX 6750 XT, Ternary Bonsai 2 27B, llama-batched-bench pp256/tg128, fa on,
q8_0 KV, aggregate decode tok/s (before -> after):
  1 seq  36.6 -> 38.1     4 seq  55.2 -> 73.3
  2 seq  39.7 -> 60.8     8 seq  66.3 -> 86.4
llama-bench tg64: 40.6 (fa off) / 38.7 (fa on), same as before.
Kernel time m=17408 k=5120: n=1 45 us, n=2 57, n=4 120, n=8 215.

test-backend-ops: 271/271 (ptq1_0 + pq2_0, forced mmvq, n = 1..8 incl. row
tail), 138/138 default path; perplexity through the 8-column path 6.2321 vs
6.2294 before (same text, fp32 summation order).

Also adds perf cases with n = 2, 4, 8 for PTQ1_0 / PQ2_0 / Q4_0.
…ments

- ggml_vk_get_dequantize_mul_mat_vec_id: add PTQ1_0 and PQ2_0 to the q8_1
  whitelist. The id q8_1 pipelines were created but never selected, so
  MUL_MAT_ID always fell back to the float path.
- test-backend-ops: add n = 4 to the PTQ1_0/PQ2_0 cases (a three-column pass
  followed by the single-column tail).
- Shorten the shader comments and the coopmat2 note in vulkan-shaders-gen.cpp.

test-backend-ops MUL_MAT/MUL_MAT_ID/GET_ROWS for ptq1_0 + pq2_0: 284/284 on
the default path and 284/284 with GGML_VK_FORCE_MMVQ=1 (RX 6750 XT); the
mul_mat_vec_id_ptq1_0_q8_1_f32 pipeline is now compiled and used.
BC-250 (gfx1013) is RDNA1: no coopmat, no accelerated integer dot. For the
PQ2_0 packed-ternary format the decode path runs mul_mat_vec_pq2_0_*; the
stock config used 2 rows/workgroup and let the large DMMV variant run at
4*subgroup=128 threads, underfilling the GPU on single-row decode.

Gated on vendor==AMD && arch==RDNA1 && deviceID==0x13fe && !coopmat so no
other device changes behavior: 4 rows/workgroup, workgroup size 32 on both
DMMV variants, required subgroup size 32; RDNA1 preferred mul_mat_vec
subgroup 64->32. f16-dst and all other quantized mmv paths unchanged.

Measured (llama-bench, Ternary-Bonsai-2-27B-PQ2_0, 2000 MHz, -fa1 KV q4_1):
tg128 22.43->23.93 t/s (+6.7%); pp512 unchanged (noise). PPL identical to
last digit (corpus-ko, -c 512 --chunks 8 -b 1024). test-backend-ops
MUL_MAT/MUL_MAT_ID/GET_ROWS 2310/2310 pass.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
@bri-prism

Copy link
Copy Markdown
Collaborator

Negative control from a non-BC-250 device: the gate holds — Intel Arc B390 (Xe3 iGPU), Vulkan / Windows

Ran this branch against its own base (#188) on hardware the gate is meant to exclude, to check the "no other device sees a behavioral difference" claim. It holds on this part.

Device:

Intel(R) Arc(TM) B390 GPU | uma: 1 | fp16: 1 | warp size: 32 | int dot: 1 | matrix cores: KHR_coopmat

Builds: #188 @ 80f71b004 vs #229 @ d6d51ddd4, MinGW/ucrt64 g++, Ninja, Release, same flags.

All four conjuncts of bc250_pq2_mmv are false here, so mmv_pq2_rows, mmv_pq2_sg and the workgroup argument all reduce to exactly #188's values:

conjunct gate requires this device
vendor_id VK_VENDOR_ID_AMD Intel
architecture AMD_RDNA1 INTEL_XE2
deviceID 0x13fe 0xB080
!coopmat_support no coopmat has KHR_coopmat

test-backend-ops test -b Vulkan0, full and unfiltered — identical

#188 80f71b004 #229 d6d51ddd4
OK 17434 17434
not supported 4524 4524
failing cases 3 3 (same cases)
PTQ1_0 199 OK / 180 n-s 199 OK / 180 n-s
PQ2_0 191 OK / 180 n-s 191 OK / 180 n-s

Every counter matches. (The 3 failures are the pre-existing GATED_DELTA_NET raw_gates=1 bug on prism, filed separately as #237 — unrelated to either PR.)

Throughput, Bonsai 2 27B, ABAB-interleaved

-p 512 -n 128 -ngl 99 -fa 1 -r 3, default dispatch (deliberately not GGML_VK_FORCE_MMVQ — the change is in the PQ2_0 dequant mat-vec, so the default path is where any effect would land):

model metric #229 (r1 / r2) #188 (r1 / r2)
PQ2_0 tg128 10.99 / 11.09 11.07 / 10.93
PQ2_0 pp512 273.4 / 279.6 266.6 / 274.3
PTQ1_0 tg128 1.66 / 1.56 1.56 / 1.56
PTQ1_0 pp512 207.5 / 179.5 177.2 / 174.6

PQ2_0 decode — the path this PR changes — differs by 0.4% between arms, inside the <1% within-run spread, ranges overlapping. The lone 1.66 on PTQ1_0 r1 is a cold-start artifact (first load after an idle GPU); r2 agrees exactly. pp512 carries ±10-38 on both arms here and does not discriminate.

Conclusion: no measurable difference on a device outside the gate, as intended.

One scoping note

The rdna1_pipelines entry is the one change not covered by the deviceID == 0x13fe gate:

-    {"argmax", 64}, {"mul_mat_vec", 64},
+    {"argmax", 64}, {"mul_mat_vec", 32},

That table is keyed on RDNA1 rather than on BC-250, so it reaches every RDNA1 part, not only gfx1013. The description notes the reasoning (required_subgroup_size wins wherever subgroups are usable), so this may well be intended — but it does mean the PR body's "All changes are gated on vendor == AMD && architecture == AMD_RDNA1 && deviceID == 0x13fe" is slightly stronger than the diff. Worth stating the intended scope explicitly, since the case it would affect is an RDNA1 device that is not a BC-250 and where subgroups are not usable. I have no RDNA1 hardware, so nothing above tests that path — this result speaks only for the Intel side.

@renovys

renovys commented Sep 22, 2026 •

Copy link
Copy Markdown
Author

Follow-up from our own live BC-250 (gfx1013) service with this change deployed: on a mixed-length workload with MTP on (6 prompts, ctx 64K shared, 2 slots), generation was at parity with the previous build (33.03 vs 33.04 t/s) and output quality was unchanged. So the +6.7% is a llama-bench tg128 decode figure; with MTP and mixed prompts the difference is within noise.

@renovys

renovys commented Sep 27, 2026

Copy link
Copy Markdown
Author

Superseded by #281, which is the cleaned-up follow-up (single commit, dedicated PQ2_0 mat-vec + BC-250 column extension + per-column rows table). This PR has diverged from the current prism base (mergeable=false) after the upstream mat-vec restructuring, and its content is carried by #281. Closing in favor of #281.

@renovys renovys closed this Sep 27, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants