Conversation
Q1_0 stores one sign bit per weight with an fp16 scale per 128 weights, so it only
applies to tensors whose contraction dim is a multiple of 128. Checkpoints whose
hidden or FFN widths are multiples of 64 but not 128 have no binary type for those
tensors, and llama-quantize aborts ("no tensor type fallback is defined for type
q1_0").
PQ1_0 is the same codec at group 64: 8 sign bytes + fp16 d = 10 bytes per 64 weights,
1.25 bpw. It gets ggml type 144, ggml ftype 130 and llama ftype 144, all unused on
every branch of this fork. llama-quantize now falls back from Q1_0 to PQ1_0 instead
of aborting.
Kernels are the Q1_0 ones at group 64, since the bit layout and the MMQ SRAM tile
layout are the same: generic CPU vec_dot against Q8_0, CUDA MMVQ and MMQ (a PQ1_0
tile loader plus config entries in every arch table), CUDA dequantize and get_rows,
and Metal mul_mv, mul_mv_id, mul_mm, mul_mm_id, get_rows and cpy. When PQ1_0 does
reach the dequantize + cuBLAS path, that matmul runs in fp32.
|
Intel coverage for the backends listed as "not built or tested": Vulkan and SYCL on an Arc B390 (Panther Lake Xe3 iGPU), plus the CPU side on a Core Ultra X7 358H (AVX2 + AVX-VNNI), Windows 11. This PR and #255 merge cleanly onto current CPU: Vulkan ( SYCL (
The user-facing issue: silent CPU fallback on Vulkan (and presumably SYCL)Because the fallback is graceful, a PQ1_0 model with
So on Vulkan, GPU offload of a PQ1_0 model is ~5× slower than not offloading at all, and ~23× slower than the same weights as Q1_0. With this PR, Suggestions, any one of which would help:
|
…e-20260929 # Conflicts: # ggml/src/ggml-cpu/arch-fallback.h
PQ1_0 used the generic scalar dot on every architecture. Each PQ1_0 block (64 weights) is dotted against two Q8_0 blocks with the same bit expansion the Q1_0 kernels use. Other architectures keep the generic alias.
What
Adds
PQ1_0, theQ1_0binary codec at group size 64: one sign bit per weight, LSBfirst, and one fp16 scale per 64 weights.
Why
Q1_0needs a tensor's contraction dim to be a multiple of 128. Checkpoints whosehidden or FFN widths are multiples of 64 but not of 128 have no binary type for those
tensors today, and
llama-quantizeaborts withno tensor type fallback is defined for type q1_0. Padding would change the model's shapes, so a 64-group type is thelossless option. With this PR,
llama-quantizefalls back fromQ1_0toPQ1_0for exactly those tensors and keeps
Q1_0everywhere else.How
The bit layout matches
Q1_0and so does the MMQ SRAM tile layout (int8 +-1 values,one scale per 32), so the kernels are the
Q1_0kernels at group 64 rather than newdesigns:
vec_dotagainstQ8_0(one64-weight block spans two
Q8_0blocks). No SIMD path yet.vec_dot_pq1_0_q8_1for MMVQ, aPQ1_0MMQ tile loader with config entriesin every arch table, dequantize and
get_rows. IfPQ1_0falls through todequantize + cuBLAS, that matmul is forced to fp32.
mul_mv,mul_mv_id,mul_mm,mul_mm_id,get_rowsandcpyto f32/f16.Testing
test-quantize-fnspasses at theQ1_0binary thresholds.test-backend-opsvs CPU forpq1_0: MUL_MAT 43/43, MUL_MAT_ID 73/73, GET_ROWS 4/4on CUDA sm_89, sm_90, sm_100 and sm_120, and the same plus CPY 2/2 on Metal. The
f16-activation MUL_MAT cases are unsupported on both backends, as for
Q1_0.PQ1_0matches aCPU-only run to within 0.01% on every CUDA arch above.
Not included
No SIMD CPU
vec_dot. No kernel tuning beyond theQ1_0configs. ROCm, MUSA,Vulkan and SYCL are not built or tested; the HIP config tables carry
PQ1_0entriescopied from
Q1_0but have not been run.AI usage disclosure: YES. Claude Code was used to help develop and test the kernels and to format this description. I reviewed every line and take full responsibility for the changes.