Skip to content

sycl: explicit-SIMD decode kernels for PTQ1_0 and PQ2_0 (1.3x PQ2_0 decode) - #8

Draft
dwymark wants to merge 1 commit into
prismfrom
sycl-esimd-decode
Draft

dwymark wants to merge 1 commit into
prismfrom
sycl-esimd-decode

Conversation

@dwymark

@dwymark dwymark commented Sep 29, 2026

Copy link
Copy Markdown
Owner

Overview

On Intel GPUs, single-token decode of PTQ1_0 and PQ2_0 runs through MMVQ, where the sub-group kernels from PrismML-Eng#235 set the decode speed. This PR adds explicit-SIMD (ESIMD) MMVQ kernels for both formats: each work-item takes one weight row, spreads the row's blocks over 16 lanes, and accumulates q8_1 dot products with dp4a, four rows per work-group. They compile only under the Intel compiler and follow a new GGML_SYCL_ENABLE_ESIMD switch, default 1; setting it to 0 restores the sub-group kernels.

It also turns a one-row, single-index F32 GET_ROWS into a vectorized copy. The recurrent-state layers of the qwen35 graph issue that shape once per layer per token.

Prompt processing is unchanged: batches above the MMVQ width take the FP16 GEMM from PrismML-Eng#278.

Additional information

Intel Arc 140V (Lunar Lake iGPU), Windows, oneAPI 2025.3, driver 32.0.101.9030, Release. Base prism @ 88c4bc6, Ternary-Bonsai-2-27B.

llama-bench -ngl 99 -fa 1 -p 512 -n 128 -r 1, four rounds in separate processes 30 s apart with the arm order flipped each round; median (range) in t/s:

base this PR
PQ2_0 tg128 6.79 (6.70-6.84) 8.88 (8.80-8.98) 1.31x
PTQ1_0 tg128 5.94 (5.93-5.95) 7.03 (7.02-7.04) 1.18x
PQ2_0 pp512 139.3 (129.2-142.1) 135.4 (122.4-142.0)
PTQ1_0 pp512 147.8 (143.6-148.8) 147.6 (104.4-149.2)

With GGML_SYCL_ENABLE_ESIMD=0, this PR decodes PTQ1_0 at 6.07 and 6.08 t/s, against 6.99 with it on. This laptop lowers its memory clock under sustained load, which accounts for the spread in prompt speed.

Correctness:

  • test-backend-ops -o MUL_MAT: PQ2_0 216/216 and PTQ1_0 167/167 on base and this PR, with oneDNN, with GGML_SYCL_ENABLE_DNN=0, and with GGML_SYCL_ENABLE_ESIMD=0. 48 cases per type have n = 1 at the model's row lengths and take the new kernels. -o GET_ROWS -p type=f32: 11/11.
  • Decode logits, full vocabulary, at 5 positions over 512 prompt tokens and 32 decode tokens. PTQ1_0 is bit-identical to base. PQ2_0 against base: normalized MSE 3.0e-5, mean KL 9.9e-5, max KL 2.3e-4, same top token at every position. Against a CPU-backend reference, both builds sit at the same distance (PQ2_0 normalized MSE 1.00e-3 for this PR and 0.99e-3 for base). With GGML_SYCL_ENABLE_ESIMD=0 both formats match base exactly, so the GET_ROWS copy changes nothing.
  • llama-perplexity --kl-divergence, WikiText-2, -c 2048 -b 512, 8 chunks, against base: mean KLD 0.000000, max KLD 0.000056, same top 100% for both formats. These batches take the GEMM path, so this mainly confirms prompt processing is untouched.

Only tested on Xe2.

Requirements

  • I have read and agree with the contributing guidelines
  • AI usage disclosure: YES. Claude Code (Anthropic) wrote the kernels while profiling decode, rebased them onto current prism, and ran the measurements above on my machine, under my direction. I reviewed the change and can explain it.

On Intel GPUs the MMVQ path for both ternary formats now runs an ESIMD kernel that gives each work-item one weight row, spreads the row's blocks over 16 lanes, and accumulates q8_1 dot products with dp4a. The kernels are compiled under the Intel compiler only and follow GGML_SYCL_ENABLE_ESIMD, so setting it to 0 restores the sub-group kernels for comparison.

A one-row GET_ROWS with a single index becomes a vectorized copy, which is the shape the recurrent-state layers issue every token.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DSrpF2XYi9dav1e3hSAkaS
Claude-Session: https://claude.ai/code/session_014smgAQnqKxMyRYQbXnvBgs
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant