Skip to content

ggml: add PQ1_0, the Q1_0 binary codec at group 64 - #253

Open
bri-prism wants to merge 3 commits into
prismfrom
feat/pq1_0-g64
Open

bri-prism wants to merge 3 commits into
prismfrom
feat/pq1_0-g64

Conversation

@bri-prism

Copy link
Copy Markdown
Collaborator

What

Adds PQ1_0, the Q1_0 binary codec at group size 64: one sign bit per weight, LSB
first, and one fp16 scale per 64 weights.

qs[8] + fp16 d = 10 bytes / 64 weights = 1.25 bpw

Why

Q1_0 needs a tensor's contraction dim to be a multiple of 128. Checkpoints whose
hidden or FFN widths are multiples of 64 but not of 128 have no binary type for those
tensors today, and llama-quantize aborts with no tensor type fallback is defined for type q1_0. Padding would change the model's shapes, so a 64-group type is the
lossless option. With this PR, llama-quantize falls back from Q1_0 to PQ1_0
for exactly those tensors and keeps Q1_0 everywhere else.

How

The bit layout matches Q1_0 and so does the MMQ SRAM tile layout (int8 +-1 values,
one scale per 32), so the kernels are the Q1_0 kernels at group 64 rather than new
designs:

  • CPU: reference quantize/dequantize and a generic vec_dot against Q8_0 (one
    64-weight block spans two Q8_0 blocks). No SIMD path yet.
  • CUDA: vec_dot_pq1_0_q8_1 for MMVQ, a PQ1_0 MMQ tile loader with config entries
    in every arch table, dequantize and get_rows. If PQ1_0 falls through to
    dequantize + cuBLAS, that matmul is forced to fp32.
  • Metal: dequantize, mul_mv, mul_mv_id, mul_mm, mul_mm_id, get_rows and
    cpy to f32/f16.

Testing

  • test-quantize-fns passes at the Q1_0 binary thresholds.
  • test-backend-ops vs CPU for pq1_0: MUL_MAT 43/43, MUL_MAT_ID 73/73, GET_ROWS 4/4
    on CUDA sm_89, sm_90, sm_100 and sm_120, and the same plus CPY 2/2 on Metal. The
    f16-activation MUL_MAT cases are unsupported on both backends, as for Q1_0.
  • End to end, perplexity of a binary model whose weights are mostly PQ1_0 matches a
    CPU-only run to within 0.01% on every CUDA arch above.

Not included

No SIMD CPU vec_dot. No kernel tuning beyond the Q1_0 configs. ROCm, MUSA,
Vulkan and SYCL are not built or tested; the HIP config tables carry PQ1_0 entries
copied from Q1_0 but have not been run.

AI usage disclosure: YES. Claude Code was used to help develop and test the kernels and to format this description. I reviewed every line and take full responsibility for the changes.

Q1_0 stores one sign bit per weight with an fp16 scale per 128 weights, so it only
applies to tensors whose contraction dim is a multiple of 128. Checkpoints whose
hidden or FFN widths are multiples of 64 but not 128 have no binary type for those
tensors, and llama-quantize aborts ("no tensor type fallback is defined for type
q1_0").

PQ1_0 is the same codec at group 64: 8 sign bytes + fp16 d = 10 bytes per 64 weights,
1.25 bpw. It gets ggml type 144, ggml ftype 130 and llama ftype 144, all unused on
every branch of this fork. llama-quantize now falls back from Q1_0 to PQ1_0 instead
of aborting.

Kernels are the Q1_0 ones at group 64, since the bit layout and the MMQ SRAM tile
layout are the same: generic CPU vec_dot against Q8_0, CUDA MMVQ and MMQ (a PQ1_0
tile loader plus config entries in every arch table), CUDA dequantize and get_rows,
and Metal mul_mv, mul_mv_id, mul_mm, mul_mm_id, get_rows and cpy. When PQ1_0 does
reach the dequantize + cuBLAS path, that matmul runs in fp32.
@bri-prism

Copy link
Copy Markdown
Collaborator Author

Intel coverage for the backends listed as "not built or tested": Vulkan and SYCL on an Arc B390 (Panther Lake Xe3 iGPU), plus the CPU side on a Core Ultra X7 358H (AVX2 + AVX-VNNI), Windows 11. This PR and #255 merge cleanly onto current prism @ 279df6644 (after #254/#256/#261/#262/#272), and type ID 144 is still unique.

CPU: test-quantize-fns passes (0 failures) with MinGW GCC 16.2 and MSVC 19.44.

Vulkan (test-backend-ops test -b Vulkan0 -p pq1_0): 357 not supported, 3 OK, 0 fail, no aborts. supports_op declines PQ1_0 cleanly, since Vulkan's MUL_MAT type list is a whitelist.

SYCL (-b SYCL0 -p pq1_0, oneAPI 2025.3): MUL_MAT (76), MUL_MAT_ID (73), GET_ROWS (4), MUL_MAT_VEC_FUSION (36) and OUT_PROD (128) are all cleanly not supported. Two cases abort:

  • SET_ROWS → set_rows.cpp:549: Unsupported tensor type!. This is the pre-existing SYCL gap where supports_op never checks the destination type, already known from sycl: add PTQ1_0 and PQ2_0 support with MMVQ vector dot kernel #235 and not caused by this PR.
  • CPY pq1_0 → pq1_0 → ggml_sycl_cpy: unsupported type combination (pq1_0 to pq1_0). SYCL's supports_op appears to accept same-type quantized copies for any type. Also latent in SYCL rather than this PR, but PQ1_0 exposes it.

The user-facing issue: silent CPU fallback on Vulkan (and presumably SYCL)

Because the fallback is graceful, a PQ1_0 model with -ngl 99 on these backends silently runs its PQ1_0 matmuls on the CPU. The log still reports full offload. With a Bonsai-1.7B requantized to PQ1_0 (all binary tensors PQ1_0, so the worst case), same build, B390:

-ngl 99, Vulkan model buffers graph splits pp512 tg128
Bonsai-1.7B Q1_0 231 MiB Vulkan0 2 1813 75.1
same, PQ1_0 210 MiB CPU + 243 MiB Vulkan0, "offloaded 29/29 layers" 226 60.2 3.3
PQ1_0, -ngl 0 (CPU only, #255 AVX2, t=4) — — 45 17.3

So on Vulkan, GPU offload of a PQ1_0 model is ~5× slower than not offloading at all, and ~23× slower than the same weights as Q1_0. With this PR, llama-quantize falls back from Q1_0 to PQ1_0 automatically for 64-but-not-128-wide tensors, so Vulkan/SYCL users can end up here without having chosen PQ1_0. A real model would only have its odd-width tensors as PQ1_0, but each of those matmuls still round-trips through the CPU.

Suggestions, any one of which would help:

  • A Vulkan PQ1_0 dequant/mat-vec. It's the Q1_0 codec at group 64, so the Q1_0 shaders should be most of the way there.
  • Have llama-quantize log which tensors took the PQ1_0 fallback and note the backend coverage.
  • A known-issues entry until Vulkan/SYCL support lands.

…e-20260929

# Conflicts:
#	ggml/src/ggml-cpu/arch-fallback.h
PQ1_0 used the generic scalar dot on every architecture. Each PQ1_0 block
(64 weights) is dotted against two Q8_0 blocks with the same bit expansion
the Q1_0 kernels use. Other architectures keep the generic alias.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant