Skip to content

ggml-cpu: add NEON, AVX2, AVX and SSSE3 vec_dot for PQ1_0 - #255

Merged
bri-prism merged 2 commits into
feat/pq1_0-g64from
feat/pq1_0-cpu-simd
Sep 29, 2026
Merged

bri-prism merged 2 commits into
feat/pq1_0-g64from
feat/pq1_0-cpu-simd

Conversation

@bri-prism

Copy link
Copy Markdown
Collaborator

Stacked on #253 (base feat/pq1_0-g64). Retarget to prism once #253 merges.

What

SIMD ggml_vec_dot_pq1_0_q8_0 for ARM NEON and x86 AVX2, AVX and SSSE3. PQ1_0 used the generic scalar dot on every CPU until now.

Why

The CPU path is the fallback for any tensor that is not offloaded, and with the scalar dot, CPU prompt processing on a PQ1_0 model ran at well under half the speed this PR gets (numbers below).

How

One PQ1_0 block holds 64 weights, so each block is dotted against two Q8_0 blocks. The bit expansion is the same one the Q1_0 kernels use: table_b2b_0 on NEON, bytes_from_bits_32 on AVX2 and AVX, and a shuffle and mask on SSSE3. Architectures without an implementation keep the generic alias in arch-fallback.h.

Testing

test-quantize-fns gives the same errors as the generic build on every path (abs 0.009694, dot 0.211317). For each x86 build I checked the disassembly of the function to confirm which path was compiled:

build path compiled test-quantize-fns
Apple M5 Pro, native NEON pass
EPYC, native (AVX-512 host) AVX2 pass
GGML_NATIVE=OFF, AVX only AVX pass
GGML_NATIVE=OFF, -mssse3 SSSE3 pass
GGML_NATIVE=OFF, GGML_SSE42=OFF generic (tail call) pass

CPU-only llama-bench (t/s), generic vs this PR:

machine, model test generic this PR
M5 Pro, MoE with almost all weights in PQ1_0 pp128 23.5 60.9
tg32 16.1 42.0
EPYC 32 threads, MoE with PQ1_0 expert down projections pp128 90.9 267.7 (AVX2)
tg32 37.6 47.3 (AVX2)

The EPYC was a shared host, so its numbers are noisy (tg32 stdev about 2.7). The AVX and SSSE3 builds also lose AVX2 for every other type, so I only checked them for correctness, not speed.

Not included

No AVX-512, RISC-V, POWER, LoongArch or s390x kernels; those keep the generic path.

AI usage disclosure: YES. Claude Code was used to help develop and test the kernels and to format this description. I reviewed every line and take full responsibility for the changes.

PQ1_0 used the generic scalar dot on every architecture. Each PQ1_0 block
(64 weights) is dotted against two Q8_0 blocks with the same bit expansion
the Q1_0 kernels use. Other architectures keep the generic alias.
@bri-prism

Copy link
Copy Markdown
Collaborator Author

Tested the x86 tiers on an Intel CPU, adding Intel/Windows coverage to your EPYC runs: Core Ultra X7 358H (Panther Lake, hybrid, AVX2 + AVX-VNNI, no AVX-512), Windows 11, MSYS2 GCC 16.2.

  • Base: prism @ 0324c6652 + ggml: add PQ1_0, the Q1_0 binary codec at group 64 #253 (generic PQ1_0).
  • PR: the same base + this PR (0e3d8aeb0).
  • Builds: each arm built four ways so that every new tier is selected: native, AVX2 (VNNI off), AVX only, and SSE4.2 only (0 ymm in that binary). Each tier is compared with the generic kernel built with the same flags.

Correctness: test-quantize-fns passes (0 failures) in every build: generic, AVX2, AVX, SSSE3.

Performance. No PQ1_0 checkpoint was available here, so I made one by requantizing Bonsai-1.7B-Q1_0.gguf with this tree's llama-quantize --allow-requantize … PQ1_0. It's fine for speed and meaningless for quality. llama-bench -ngl 0 -p 64 -n 32 -r 2, two rounds with the order reversed (t/s):

build (tier taken) threads generic pp64 PR pp64 generic tg32 PR tg32
native (AVX2) 4 12.65 / 12.26 102.11 / 94.80 (7.9x) 9.20 / 9.14 30.44 / 29.61 (3.3x)
native (AVX2) 8 17.13 / 17.06 100.54 / 101.61 (5.9x) 9.47 / 9.59 20.94 / 22.13 (2.3x)
AVX2, VNNI off 4 12.68 / 12.34 106.11 / 99.28 (8.2x) 9.12 / 9.08 30.34 / 30.04 (3.3x)
AVX only 4 12.56 / 12.30 57.07 / 53.69 (4.5x) 9.16 / 8.95 23.24 / 22.96 (2.6x)
AVX only 8 16.93 / 17.04 67.57 / 67.21 (4.0x) 9.44 / 9.62 18.93 / 18.32 (2.0x)
SSE4.2/SSSE3 4 12.19 / 12.07 45.49 / 43.11 (3.7x) 8.45 / 8.26 17.39 / 17.52 (2.1x)
SSE4.2/SSSE3 8 16.81 / 16.61 59.34 / 59.58 (3.6x) 8.98 / 8.94 16.00 / 16.07 (1.8x)

Every tier is a clear win, ordered AVX2 > AVX > SSSE3 as expected. The AVX2 path doesn't depend on VNNI (the VNNI-off build is as fast). Prompt processing gains are larger here than your 2.6–2.9x on EPYC. As with other kernels on this hybrid CPU, decode peaks at 4 threads.

Not tested: MSVC, NEON.

Tested with Claude Code.

@bri-prism

Copy link
Copy Markdown
Collaborator Author

Addendum: numerical parity, SIMD vs generic on the same file. Same requantized 1.7B PQ1_0 file, llama-perplexity -c 512 --chunks 8 -t 8, base logits from the generic (#253) native build.

comparison vs generic-native base mean KLD max KLD same top
control: generic native vs itself 0.000000 0.000051 100.000 %
control: generic built SSE-only (same PQ1_0 kernel, other ops take different ISA paths) 0.000550 0.007028 98.529 %
this PR, native (AVX2) 0.000539 0.006700 98.431 %
this PR, AVX2 (VNNI off) 0.000539 0.006700 98.431 %
this PR, AVX only 0.000527 0.008376 98.676 %
this PR, SSSE3 0.000563 0.009530 98.824 %

Reading: this collapsed requantized model (PPL 17.7) is very sensitive to summation order. Rebuilding the unchanged generic kernel with different ISA flags moves KLD by the same ~5e-4 as this PR does. So the PR's deviation is at float-reordering level, not a kernel error. That fits the code:

  • The AVX2 path's bit → weight mapping (byte_shuf + bit_masks) matches the generic loop.
  • The integer per-block dot sums are exact.
  • Only the float accumulation differs (8-lane partials with FMA and a final hsum, vs the generic serial sum).
  • I also checked the one theoretical trap: (qy ^ sm) - sm would mis-negate qy = -128, but quantize_row_q8_0 clamps to ±127, so it's unreachable.

On a non-collapsed model I'd expect this number to be much smaller. Bit-level parity per call is what test-quantize-fns covers, and that passes.

Tested with Claude Code.

@bri-prism

Copy link
Copy Markdown
Collaborator Author

NEON parity on the M5 Pro, CPU only: gpt-oss-20b PQ1_0, llama-perplexity -c 512 --chunks 8, generic (#253) vs this PR, mean KLD 0.000003 (max 0.0007), and test-quantize-fns passes on both builds. The PPL ratio reads 1.025 only because this naive pack's PPL is huge; the KLD shows the distributions match.

@bri-prism

Copy link
Copy Markdown
Collaborator Author

Follow-up on the two gaps from my earlier comment:

  • Rebased stack: this PR + ggml: add PQ1_0, the Q1_0 binary codec at group 64 #253 merge cleanly onto current prism @ 279df6644, and test-quantize-fns still passes (0 failures).
  • MSVC (previously "not tested"): MSVC 19.44, -DGGML_AVX2=ON -DGGML_AVX_VNNI=ON, 0 warnings in arch/x86/quants.c, VNNI path compiled in (425 vpdpbusd), test-quantize-fns 0 failures. PQ1_0 on the requantized 1.7B at t=4: tg32 32.0 t/s, pp64 55.0 t/s. Decode matches GCC (~30 t/s); MSVC's pp64 is lower than GCC's (~100), which is the same MSVC-vs-GCC prefill gap seen on other kernels here.

Separately, the Vulkan/SYCL behaviour of PQ1_0 (silent CPU fallback) is in my comment on #253.

…r255-20260929

# Conflicts:
#	ggml/src/ggml-cpu/arch-fallback.h
@bri-prism
bri-prism merged commit 29cddbe into feat/pq1_0-g64 Sep 29, 2026
3 of 5 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant