Skip to content

sycl: 2-3x prompt processing speedup by dequantizing PTQ1_0 and PQ2_0 to FP16 - #278

Merged
bri-prism merged 1 commit into
PrismML-Eng:prismfrom
dwymark:sycl-fp16-prompt-gemm
Sep 28, 2026
Merged

bri-prism merged 1 commit into
PrismML-Eng:prismfrom
dwymark:sycl-fp16-prompt-gemm

Conversation

@dwymark

@dwymark dwymark commented Sep 26, 2026

Copy link
Copy Markdown

Overview

SYCL has MMQ disabled for all types, so quantized prompt batches take the dequantize-then-GEMM path in ggml_sycl_op_mul_mat_sycl. In the default build (GGML_SYCL_F16=OFF) that path expands the weights to FP32 and runs an FP32 GEMM, which does not use the XMX matrix engines on Intel GPUs. This PR sends PTQ1_0 and PQ2_0 down the existing FP16 branch instead: half the bytes per dequantized weight, and oneDNN (or MKL with GGML_SYCL_ENABLE_DNN=0) runs an FP16 GEMM with FP32 output.

PTQ1_0 and PQ2_0 weights are exactly representable after FP16 expansion; FP16 conversion of activations and the GEMM arithmetic can still change results. This PR enables the path by default only for the two measured formats. Other quantized types retain the FP32 default and can still use the FP16 path with GGML_SYCL_F16=ON. Decode is unchanged, since single tokens and small batches go through MMVQ.

Additional information

Intel Arc 140V (Lunar Lake iGPU), Windows, oneAPI 2025.3, driver 32.0.101.9030, Release. Base prism @ adfffbe, Ternary-Bonsai-2-27B.

llama-bench -ngl 99 -fa 1 -r 1, four rounds in separate processes 30 s apart with the arm order flipped each round; median (range) in t/s:

base this PR
PQ2_0 pp512 53.8 (52.0-55.0) 119.0 (104.3-132.4) 2.2x
PQ2_0 pp2048 38.0 (32.2-43.2) 116.4 (110.7-129.8) 3.1x
PQ2_0 tg128 6.02 6.02
PTQ1_0 pp512 54.0 (52.3-55.8) 142.4 (131.0-146.3) 2.6x
PTQ1_0 pp2048 45.3 (32.2-46.7) 111.4 (106.1-137.1) 2.5x
PTQ1_0 tg128 4.72 4.94

This laptop lowers its memory clock under sustained load, which accounts for the spread.

Correctness:

  • test-backend-ops -o MUL_MAT: PQ2_0 142/142 and PTQ1_0 165/165, with oneDNN and with GGML_SYCL_ENABLE_DNN=0. 15 cases per type have n = 9, 16 or 64 and take the changed path.
  • llama-perplexity --kl-divergence, WikiText-2, -c 2048 -b 512, 8 chunks, against a base file from prism 0781925 (identical SYCL code). Unmodified base build: PPL ratio 1.000035, mean KLD 0.000000, same top 100%. This PR: PPL ratio 1.000035 +- 0.000036, mean KLD 0.000001, max KLD 0.000898, same top 99.963%. PTQ1_0 and PQ2_0 give identical numbers, as they pack the same trits.

Only tested on Xe2. Builds on #235.

Requirements

  • I have read and agree with the contributing guidelines
  • AI usage disclosure: YES. Claude Code (Anthropic) found the change while profiling prompt processing, wrote the patch, and ran the measurements above on my machine, under my direction. I reviewed the change and can explain it.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DSrpF2XYi9dav1e3hSAkaS
@dwymark dwymark changed the title sycl: dequantize PTQ1_0 and PQ2_0 prompt batches to FP16 (2-3x prompt processing) sycl: 2-3x prompt processing speedup by dequantize PTQ1_0 and PQ2_0 prompt batches to FP16 Sep 26, 2026
@dwymark dwymark changed the title sycl: 2-3x prompt processing speedup by dequantize PTQ1_0 and PQ2_0 prompt batches to FP16 sycl: 2-3x prompt processing speedup by dequantizing PTQ1_0 and PQ2_0 to FP16 Sep 26, 2026
@bri-prism

Copy link
Copy Markdown
Collaborator

Tested on an Intel Arc B390 (Panther Lake Xe3 iGPU), which is one generation newer than the 140V you measured on. Windows 11, oneAPI 2025.3 (Level Zero), default GGML_SYCL_F16=OFF, base prism @ adfffbe41 vs this PR @ 1f16fba6e. Nothing else was running.

Correctness: test-backend-ops test -b SYCL0 -o MUL_MAT -p "ptq1_0|pq2_0" gives 307/307 on both base and PR.

Precision (llama-perplexity -c 512 --chunks 8, 2B PTQ1_0, -ngl 99; the prompt batches take the new FP16 path):

comparison vs base logits max KLD same top PPL ratio
this PR 5.6e-5 99.95 % 1.0002 ± 0.0002
control: base vs itself 5.4e-5 100.00 % —

So the FP16 activation conversion is within run-to-run noise here.

Speed, Ternary-Bonsai-2-27B, llama-bench -ngl 99 -fa 1 -r 2, two rounds (t/s):

base this PR
PQ2_0 pp512 81.0 / 76.7 182.6 / 196.7 ~2.4×
PQ2_0 pp2048 75.4 / 67.7 127.6 / 168.3 ~2.1×
PTQ1_0 pp512 81.3 / 79.8 178.9 / 193.9 ~2.3×
PTQ1_0 pp2048 70.1 / 69.3 150.6 / 163.9 ~2.3×
tg64 (both formats) 8.0 – 8.9 7.5 – 9.3 unchanged

This reproduces your 2.2–3.1× on Xe3, and decode is unaffected as expected. Caveat: the arm order was base then PR in both rounds, not flipped, but the gap is far outside the round-to-round spread.

@dwymark

dwymark commented Sep 26, 2026

Copy link
Copy Markdown
Author

Great, glad to hear it. Let me know if you'd like anything else on my end. Feel free to merge whenever you are happy with the contribution.

@bri-prism

Copy link
Copy Markdown
Collaborator

Follow-up with the rest of the checks on the Arc B390 (same builds, adfffbe41 vs 1f16fba6e):

  • Full unfiltered test-backend-ops test -b SYCL0 -o MUL_MAT: 1329/1329 on both base and PR with oneDNN.
  • With GGML_SYCL_ENABLE_DNN=0: 1297/1329 on both. The 32 failures are identical in both builds (f16 x f32, n=1, non-contiguous k=1056/1057), so they predate this PR. All 307 PTQ1_0/PQ2_0 cases pass in both modes.
  • KLD on wikitext-2 this time (2B PTQ1_0, -c 512, 8 chunks): mean 1e-6, max 6.7e-5, same top 100%.

Looks good to me.

@bri-prism
bri-prism merged commit 87268f7 into PrismML-Eng:prism Sep 28, 2026
4 of 5 checks passed
@dwymark
dwymark deleted the sycl-fp16-prompt-gemm branch September 29, 2026 20:02
bri-prism pushed a commit that referenced this pull request Oct 2, 2026
* Adds an XMX (DPAS) path for PQ2_0 and PTQ1_0 matrix multiplies on Intel Xe2 GPUs (Battlemage, Lunar Lake). It covers every batch size: prompt processing, single-token decode, parallel decode and speculative verification. On Xe2 it replaces MMVQ and the dequantize-to-FP16 prompt path from #278 for these types.

The main idea: PQ2_0 packs value+1 in 2-bit fields, lowest first, which as a little-endian dword is already the DPAS 2-bit operand layout. The kernel feeds the codes to DPAS as signed 2-bit (s2) against int8 activations, converting them in registers with a borrow-free per-field subtract. Weights are never widened to 8 or 16 bits in memory.

How it works:
- **Weight layout.** On first use, each PQ2_0 weight is rewritten in place into an SoA layout (same size): all 32-byte qs blocks first, then all fp16 scales. The qs row pitch is then nb*32 bytes, so one transposed 2D block load fetches 16 rows x 128 weights directly in the DPAS B-operand layout. The 34-byte AoS blocks are only 2-byte aligned and can't be loaded that way.
- **PTQ1_0.** Base-3 has no DPAS form, so PTQ1_0 weights are expanded once into the same 2-bit layout. That is 34 bytes a block instead of 28. On Xe2 devices the buffer type reserves the room for it (get_alloc_size), so VRAM holds only the expanded copy: about +25% for these weights, and the file on disk is unchanged.
- **Activations.** They are quantized to int8 with one scale per 128 values, one PQ2_0 block. The four DPAS of a block then accumulate in int32 before a single float rescale, instead of one rescale per 32 values.
- **One kernel for all batch sizes.** Tiles are 8/16/32 tokens x 32 rows. For small batches, up to 16 threads of a work-group split K for the same tile and reduce through SLM, so a mat-vec keeps enough loads in flight.

Scope and fallbacks:
- On by default on BMG-G21/G31 and LNL-M. `GGML_SYCL_DISABLE_XMX=1` or `GGML_SYCL_ENABLE_OPT=0` turns it off.
- Only plain MUL_MAT on 2D weights in regular (non-split, non-COMPUTE) buffers with K >= 256. Everything else (op offload, split buffers, MUL_MAT_ID, other devices) takes the existing paths unchanged.
- Like the existing reorders, a tensor keeps its new layout once rewritten. The gate refuses MUL_MAT_ID expert slices and views. `get_rows` asserts if it ever sees a rewritten PQ2_0/PTQ1_0 tensor.

Files: `pq2_xe2.cpp/.hpp` (new), `ggml-sycl.cpp` (gate, alloc size, extras), `common.hpp` (layout flag), `getrows.cpp` (assert).

Intel Arc Pro B50 (16 GB), oneAPI 2026.1, Level Zero, Linux `xe` driver, `-ngl 99`, Ternary Bonsai models. "#278" is the same build with `GGML_SYCL_DISABLE_XMX=1`, i.e. the current prism code path. Both columns were measured in the same session.

`llama-bench`, t/s:

| model | test | #278 | this PR | change |
|---|---|---|---|---|
| 1.7B PQ2_0 | pp128 | 2574 | 6133 | 2.38x |
| | pp512 | 5218 | 7645 | 1.47x |
| | tg64 | 152.4 | 178.1 | +17% |
| 4B PQ2_0 | pp128 | 1068 | 3068 | 2.87x |
| | pp512 | 2298 | 3272 | 1.42x |
| | tg64 | 77.4 | 99.8 | +29% |
| 27B PQ2_0 | pp128 | 148 | 346 | 2.33x |
| | pp512 | 290 | 370 | 1.28x |
| | tg64 | 12.39 | 16.53 | +33% |
| 27B PTQ1_0 | pp128 | 127 | 348 | 2.74x |
| | pp512 | 267 | 370 | 1.39x |
| | tg64 | 12.88 | 16.54 | +28% |

Parallel decode (`llama-batched-bench -npp 128 -ntg 64`, aggregate t/s):

| sequences | 1.7B PQ2_0 | 4B PQ2_0 | 27B PQ2_0 | 27B PTQ1_0 |
|---|---|---|---|---|
| 1 | 151 -> 171 | 76 -> 97 | 12.2 -> 16.1 | 13.1 -> 16.1 |
| 2 | 250 -> 334 | 125 -> 186 | 18.9 -> 29.7 | 20.2 -> 29.7 |
| 4 | 380 -> 613 | 185 -> 355 | 25.2 -> 48.0 | 26.3 -> 47.9 |
| 8 | 478 -> 1010 | 224 -> 618 | 29.0 -> 68.6 | 31.0 -> 68.0 |
| 16 | 322 -> 1446 | 138 -> 900 | 20.2 -> 85.7 | 17.2 -> 85.8 |
| 32 | 588 -> 1769 | 262 -> 1184 | 32.5 -> 98.4 | 29.0 -> 98.4 |

On #278, batches of 9+ tokens leave MMVQ for the dequantize-to-FP16 path, which is why throughput drops at 16 sequences there.

Speculative decoding (27B PQ2_0 target, Qwen3.5-0.8B Q8_0 draft, `--spec-draft-n-max 1`, greedy, 256 tokens), t/s:

| | #278 | this PR |
|---|---|---|
| prose, target only | 13.1 | 17.6 |
| prose, with draft | 12.1 | 20.5 |
| code, target only | 12.7 | 17.5 |
| code, with draft | 12.8 | 21.9 |

With #278, drafting does not pay. With this PR, verification is cheap enough that n_max=1 gains 17-25% over the target alone. Longer drafts lose on acceptance (15-33% at n_max 2-3 with this draft model).

Where the decode gain comes from: at the 27B FFN shape (17408x5120, one token), the PQ2_0 mat-vec goes from 199 us (MMVQ, about 53% of the B50's 224 GB/s) to 88 us. End to end, 27B decode now reads about 123 GB/s. Most of the remainder is the ~2300 non-matmul kernels per token.

- `test-backend-ops -b SYCL0`: MUL_MAT 1329/1329, MUL_MAT_ID 1035/1035, GET_ROWS 119/119 (all types). Full suite, run per op: 14703/14759. The 56 failures (CONV_2D, ROLL, FLASH_ATTN_EXT, LIGHTNING_INDEXER) are identical with this path disabled, so they predate this PR. SET_ROWS aborts on current prism regardless of this PR (`set_rows.cpp:549: Unsupported tensor type!`), so it was run separately.
- Perplexity (`-c 512`, 5 chunks), #278 -> this PR:

| model | setting | #278 | this PR |
|---|---|---|---|
| 1.7B PQ2_0 | b1 | 8.4089 | 8.4335 |
| 1.7B PQ2_0 | b512 | 8.4233 | 8.4149 |
| 4B PQ2_0 | b1 | 7.0584 | 7.0565 |
| 4B PQ2_0 | b512 | 7.0545 | 7.0506 |
| 1.7B PQ2_0 | -ngl 0 (op offload) | 8.4225 | 8.4225 |
| 4B PQ2_0 | -ngl 20 (partial) | 7.0541 | 7.0535 |
| 27B PTQ1_0 | b512 | 4.2421 | 4.2399 |

- KL divergence against a CPU-only reference (1.7B, `-dev none --no-op-offload`), #278 vs this PR:
  - b1: 0.00122 vs 0.00125.
  - b512: 0.00106 vs 0.00113.
  - Same top token: 97.4-97.8% vs 97.7-97.9%.
  - The differences are within the error bars. The per-128 activation scale costs nothing measurable.

- This introduces ESIMD/XMX code to the SYCL backend, a new pattern here. I'm happy to adjust structure or naming to fit how the maintainers want it organized.
- Tuning (tile sizes, K-split thresholds) was done on the B50 only. BMG-G31 and LNL-M are gated in by architecture but untested.

- I have read and agree with the [contributing guidelines](https://github.com/ggml-org/llama.cpp/blob/master/CONTRIBUTING.md)
- AI usage disclosure: YES - the code and this description were written by Claude (Anthropic) under my direction; I tested it on my own hardware (Intel Arc Pro B50) and reviewed the changes.

* sycl: gate the PQ2_0/PTQ1_0 XMX path on 16-wide DPAS instead of a device list

The path was limited to BMG-G21/G31 and LNL-M by name. It now runs on any XMX device whose int8 DPAS is 16 wide,
as the runtime reports it through matrix_combinations: Xe-HPC, Xe2, Xe3 and later. 8-wide XMX (Xe-HPG, Arrow
Lake-H) and devices without XMX keep the existing paths.

AOT builds compile the kernels for every GGML_SYCL_DEVICE_ARCH device, and they do not build for targets without
16-wide DPAS and 2D block loads (ocloc crashes on acm-g10, and fails on tgllp, dg1, mtl and pvc-vg). pq2_xmx.cpp
is now built only when every AOT device is Xe-HPC, Xe2 or later; otherwise it compiles to stubs.

Renames pq2_xe2 to pq2_xmx to match.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants