ggml-cpu: add NEON, AVX2, AVX and SSSE3 vec_dot for PQ1_0 - #255
Conversation
PQ1_0 used the generic scalar dot on every architecture. Each PQ1_0 block (64 weights) is dotted against two Q8_0 blocks with the same bit expansion the Q1_0 kernels use. Other architectures keep the generic alias.
|
Tested the x86 tiers on an Intel CPU, adding Intel/Windows coverage to your EPYC runs: Core Ultra X7 358H (Panther Lake, hybrid, AVX2 + AVX-VNNI, no AVX-512), Windows 11, MSYS2 GCC 16.2.
Correctness: Performance. No PQ1_0 checkpoint was available here, so I made one by requantizing
Every tier is a clear win, ordered AVX2 > AVX > SSSE3 as expected. The AVX2 path doesn't depend on VNNI (the VNNI-off build is as fast). Prompt processing gains are larger here than your 2.6–2.9x on EPYC. As with other kernels on this hybrid CPU, decode peaks at 4 threads. Not tested: MSVC, NEON. Tested with Claude Code. |
|
Addendum: numerical parity, SIMD vs generic on the same file. Same requantized 1.7B PQ1_0 file,
Reading: this collapsed requantized model (PPL 17.7) is very sensitive to summation order. Rebuilding the unchanged generic kernel with different ISA flags moves KLD by the same ~5e-4 as this PR does. So the PR's deviation is at float-reordering level, not a kernel error. That fits the code:
On a non-collapsed model I'd expect this number to be much smaller. Bit-level parity per call is what Tested with Claude Code. |
|
NEON parity on the M5 Pro, CPU only: gpt-oss-20b PQ1_0, |
|
Follow-up on the two gaps from my earlier comment:
Separately, the Vulkan/SYCL behaviour of PQ1_0 (silent CPU fallback) is in my comment on #253. |
…r255-20260929 # Conflicts: # ggml/src/ggml-cpu/arch-fallback.h
Stacked on #253 (base
feat/pq1_0-g64). Retarget toprismonce #253 merges.What
SIMD
ggml_vec_dot_pq1_0_q8_0for ARM NEON and x86 AVX2, AVX and SSSE3. PQ1_0 used the generic scalar dot on every CPU until now.Why
The CPU path is the fallback for any tensor that is not offloaded, and with the scalar dot, CPU prompt processing on a PQ1_0 model ran at well under half the speed this PR gets (numbers below).
How
One PQ1_0 block holds 64 weights, so each block is dotted against two Q8_0 blocks. The bit expansion is the same one the Q1_0 kernels use:
table_b2b_0on NEON,bytes_from_bits_32on AVX2 and AVX, and a shuffle and mask on SSSE3. Architectures without an implementation keep the generic alias inarch-fallback.h.Testing
test-quantize-fnsgives the same errors as the generic build on every path (abs 0.009694, dot 0.211317). For each x86 build I checked the disassembly of the function to confirm which path was compiled:GGML_NATIVE=OFF, AVX onlyGGML_NATIVE=OFF,-mssse3GGML_NATIVE=OFF,GGML_SSE42=OFFCPU-only
llama-bench(t/s), generic vs this PR:The EPYC was a shared host, so its numbers are noisy (tg32 stdev about 2.7). The AVX and SSSE3 builds also lose AVX2 for every other type, so I only checked them for correctness, not speed.
Not included
No AVX-512, RISC-V, POWER, LoongArch or s390x kernels; those keep the generic path.
AI usage disclosure: YES. Claude Code was used to help develop and test the kernels and to format this description. I reviewed every line and take full responsibility for the changes.