Repository navigation
Conversation
Replace the ESIMD kernels of the PQ2_0/PTQ1_0 XMX path with SIMT SYCL kernels ported from Intel's TernSYCL (int2_via_int2_x_int8_dpas, BSD-3-Clause): inline-vISA s2 x s8 DPAS and 2D block I/O, a mat-vec kernel with a K split reduced in SLM for up to 8 tokens, and a GEMM with 256 GRF whose tile depends on the batch size and which splits K over work-groups when a small batch would leave most of the GPU idle. - PQ2_0 layout: uint32 [K/16][N] s2 codes (one 2D block load gives a sub-group 16 columns x 128 weights in the DPAS operand layout), then the fp16 scales [K/128][N]. - PTQ1_0 stays at 1.75 bits: [K/128][7][N] dwords (qs, then qh and the scale). The kernels decode the trits through a 256-entry table in SLM, in byte-major K order so the 10-bit table entries concatenate into the s2 dwords; the activation is quantized in the same permuted order. PTQ1_0 weights no longer expand to 34 bytes a block. For large batches the GEMM decodes each weight block once per work-group and shares it through SLM. - Rows are padded to 16 in the allocation, so any row count works (e.g. 151669-token vocabs). The gate, the layout flag, the capability check and the fallbacks of ggml-org#294 are unchanged.
On top of the XMX kernels, the graph loop now fuses around PQ2_0/PTQ1_0 mat-muls: - Epilogues (TernSYCL postop 1 and 2) in the mat-vec and GEMM stores: gate + up + SWIGLU runs as two mat-muls from one quantized activation, the gate writing silu(gate) * up; mat-mul + residual ADD (also through reshapes) writes the sum. - Shared activations: a quantized activation is kept for the length of a graph compute and reused by every XMX mat-mul that reads it (q/k/v, qkv and z, gate and up). Entries live in a separate pool, as the VMM pool only frees in reverse order. - Hadamard folding (Bonsai 2): the sign flip, the 1024-wide FWHT and the int8 quantization of every Hadamard-rotated mat-mul input run in one kernel; when all users of the FWHT output are XMX mat-muls, it is never written. Otherwise one kernel does sign flip + FWHT. - A mat-mul whose only use is the gate of a later SWIGLU (z in gated delta net layers) runs at the GLU, with the SWIGLU in its store. GGML_SYCL_ENABLE_FUSION=0 turns this off.
…othing is shared PTQ1_0 decodes its trits in the kernels, so with few tokens the decode, not the weight reads, set the speed, and the 27B decoded 8 and 16 parallel sequences 3-5% slower than with the expanded weights of ggml-org#294. - The GEMM decodes in registers when a work-group has one sub-group row (WG_M == 1), without the SLM round trip and barrier that only pay off when sub-groups share the decoded block. - PTQ1_0 picks its own tiles: 16x16 and 32x16 per sub-group for 5-32 tokens (each decoded block feeds 2-4 row blocks), 32x32 with one sub-group row for 33-64 tokens with K >= 6144, and the 4-row mat-vec for 2 tokens (the 2-row variant is slower). Arc Pro B50, Bonsai 2 27B PTQ1_0, llama-batched-bench decode t/s at 8 / 16 / 32 sequences: 77.4 / 93.1 / 110.7 -> 82.2 / 105.7 / 119.7 (prism: 80.1 / 98.1 / 110.8).
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Overview
Depends on #327 (its two commits are in this branch: the first and the third, a cherry-pick of #327's follow-up; review #327 first). This PR adds the second commit: the graph loop now fuses the PQ2_0/PTQ1_0 XMX mat-muls with the ops around them, so far fewer small kernels run per token.
{mul_mat(gate), mul_mat(up), SWIGLU}runs as two mat-muls from one quantized activation, the gate writingsilu(gate) * up;mul_mat -> (reshape) -> ADD(the residual adds afterffn_down,ssm_outand the attention output) writes the sum.xmx_act_pool()), because the VMM pool only frees in reverse order.x * signs -> FWHT_1024 (hinted mul_mat) -> mul_mat. The sign flip, the FWHT and the int8 quantization now run in one kernel (the butterflies followggml_sycl_op_fwht). When every user of the FWHT output is an XMX mat-mul (checked against the graph's use counts), that output is never written; otherwise one kernel does sign flip + FWHT for the other users.zof the gated delta net layers, three nodes before its GLU) keeps its quantized activation and runs at the GLU, with the SWIGLU in its store.GGML_SYCL_ENABLE_FUSION=0turns all of this off. Other models and types take exactly the paths of #327.The fusion rules are independent of the weight type; only
ggml_sycl_pq2_xmx_eligible()(the side-effect-free half of #294's gate, split out here) decides which mat-muls take part.Results
Same setup as #327: Intel Arc Pro B50 (16 GB), oneAPI 2026.1, Level Zero 1.14.37020, Linux 7.0
xedriver,-ngl 99. "#327" is the first commit alone; all columns measured in the same session.llama-bench, t/s:-fa 1)-fa 1)Parallel decode (
llama-batched-bench -npp 128 -ntg 64), aggregate decode t/s, #327 -> this PR:Speculative decoding, Bonsai 2 27B PTQ1_0 with the grafted MTP head (sudoingx/Ternary-Bonsai-2-27B-PTQ1_0-MTP-GGUF),
llama-server, 3 chat prompts (prose, code, bash) x 256 tokens, greedy, thinking off, decode t/s:--spec-type draft-mtp --spec-draft-n-max 1--spec-type draft-mtp --spec-draft-n-max 2With this stack the B50 decodes Bonsai 2 27B at ~39 t/s (code and bash ~42, prose ~33) on one 16 GB, 70 W card, short context.
Correctness
test-backend-ops -b SYCL0, run per op over all 130 ops: MUL_MAT 1401/1401, MUL_MAT_ID 1035/1035, MUL_MAT_VEC_FUSION 506/506 (gate/up + GLU and mat-mul + bias ADD on PQ2_0 and PTQ1_0, 1-8 tokens), ADD 99/99, GET_ROWS 119/119. The ops that do not pass fully (CONV_2D, FLASH_ATTN_EXT, LIGHTNING_INDEXER, ROLL; CPY and SET_ROWS abort) do the same on prism.GGML_SYCL_DEBUG=1, that every Hadamard site of the 27B decode graph takes the fused path (257 sites per decode graph).-c 512, 5 chunks of the repository'sdocs/*.md): identical to sycl: PQ2_0/PTQ1_0 XMX kernels from TernSYCL, PTQ1_0 kept at 1.75 bpw #327 for 1.7B b1/b512, 4B b1/b512, 1.7B-ngl 0, 4B-ngl 20; 27B PQ2_0/PTQ1_0 b512 6.2356 (sycl: PQ2_0/PTQ1_0 XMX kernels from TernSYCL, PTQ1_0 kept at 1.75 bpw #327) -> 6.2367 (this PR), prism 6.2370; 27B PTQ1_0 with-ub 166.2389.Additional information
Requirements