Repository navigation
Conversation
|
Tested the current head on Apple M5 Pro. The Metal build succeeded and all 149 existing PQ2_0/F32 MUL_MAT correctness cases passed against the CPU reference, including widths 2-8 and ragged output rows. One performance issue before enabling this across all supported widths: the new eight-column path is consistently slower than the existing fallback on this device. Using the existing
The eight-column measurements were 608.70/610.45 us for the new path and 491.58/489.77 us for the fallback. The predicate currently enables the new path for every width 2-8 regardless of device, so this introduces a default regression on M5 Pro. Please retain the fallback where it wins, and measure widths 5-7 to establish the crossover before selecting the default. The M3 crossover may differ. Repro command, with and without the disable variable: test-backend-ops perf -b MTL0 -o MUL_MAT -p 'type_a=pq2_0,type_b=f32,m=17408,n=(2|3|4|8),k=5120'These are warmed operator measurements from the existing harness, which repeatedly uses the same tensors; they do not establish whole-model throughput. The correctness checks use the harness tolerance and do not independently establish the bit-identical logits claim. |
b4457d5 to
139a108
Compare
|
Thanks for testing on Apple M5 Pro. Here are the verification numbers measured on Apple Silicon M3 (10 GPU Cores, 100 GB/s) using the pure PR head (
Observations
StatusI noticed upstream commit Since this covers multi-column PQ2_0 support, please let me know if you would like this PR rebased on current |
|
Thanks for the M3 measurements. Please rebase this PR onto the current |
139a108 to
f27219d
Compare
|
Rebased onto current
M3 Benchmark (17408 x 5120, best of runs, us)
Widths 5–7 bridge the latency gap between N=4 and the GEMM crossover at N=9, cutting fallback latency by 39–49% on non-tensor devices. All unit tests passed with 0 failures ( |
|
@nogeonwoo @jasontitus, #293 and #279 both add PQ2_0 multi-column mat-vec on Metal. #279 was written against an older base, and the @nogeonwoo, a few things before merge:
|
…2..8) - Adds kernel_mul_mv_pq2_0_multicol supporting widths 2..8 (_mc_c2.._mc_c8) - Avoids the mul_mv_ext fallback cliff on pre-M5 Apple Silicon (M1-M4), yielding 1.6x-1.9x latency improvements for verify batches M=3..8 - Preserves the existing fewrow tensor path and _nr1_2_r4 dispatch on M5 (gated on !has_tensor) - Retains the upstream _nr1 dispatch for widths 2 and 3 when GGML_METAL_PQ2_0_MULTICOL_DISABLE=1 - Adds test-backend-ops evaluation cases covering widths 2..8 at Bonsai-2 projection shapes
f27219d to
42f871f
Compare
|
@bri-prism Pushed an update (
|
|
I’m comparing both changes against current prism on M2 Max and M5 Max. Preliminary M2 operator results show regressions relative to the updated baseline, so I’m checking whole-model behavior before deciding what to retain from #279. I’ll post the paired measurements and coverage results once that validation is complete. |
|
I compared both changes against current Measurements used three paired ABBA / BAAB / ABBA quartets with 8-second cooldowns. All observations were retained. Throughput ratios below are candidate / baseline; values below 1 mean the candidate is slower. #279 — consolidationVerdict: OK to consolidate, carrying over its two focused row-tail/broadcast/strided-B correctness cases.
I would retain the focused test coverage rather than keep another kernel solely for the small M5 pp2 difference. #293 — more selective default needed
Recommendation: Please preserve the existing fallback where it wins rather than enabling the new path on every The operator measurements were on battery with nominal/fair thermal states and some noisy unchanged-path controls. The two AC-powered model quartets had fair thermal states and low-power mode off. I am not claiming GPU exclusivity or a regression on every pre-M5 device. Numerical finding — M2 Max
AI assistance: Codex assisted with benchmark setup, analysis, and drafting. |
|
@jasontitus Thanks for running the numbers on M2 Max, and agree on consolidating #279. The 4–8 column drop on M2 Max matches the kernel's register pressure. On a base M3 (10 cores, 100 GB/s), saving memory traffic easily pays for the register hit. But with 38 cores and 400 GB/s, memory isn't the bottleneck and lower occupancy hurts. 2 and 3 columns look like solid wins on both setups (+19% and +8% on M2 Max, 1.39x on M3). For 4–8, selective dispatch makes sense—we can either gate 4–8 to non-Max chips via @bri-prism Let me know which approach you prefer and I'll push an update. |
Summary
Add multi-column batch verification kernel for 2-bit$M \in [2, 8]$ .
PQ2_0on Apple Silicon Metal (kernel_mul_mv_pq2_0_multicol), supporting batch sizesProblem
Previously, while 1-bit$M=4$ ),
PTQ1_0implementedkernel_mul_mv_ptq1_0_multicol(up toPQ2_0had no multi-column batch GEMV kernel.When running speculative decoding (e.g. DFlash / DSpark) or prompt evaluation with batch verification ($M \ge 2$ ),
ggml-metal-ops.cppline 2802 was forced to fall back to the scalar unpack loop inkernel_mul_mv_ext_q4_f32_disp. This caused a severe throughput bottleneck where 4-token draft verification took ~506 ms (slower than 4 sequential single-token decodes).Solution
mul_mv.metal):pq2_0_dot_multicol<nr1>andkernel_mul_mv_pq2_0_multicol<nr0, nr1>.ggml-metal-device.cpp,ggml-metal-ops.cpp):ggml_metal_pq2_0_multicol_enabled()predicate checkingggml_metal_op_mul_mat().Benchmark
Ternary-Bonsai-2-27B-PQ2_0.gguf(7.21 GB)q4_f32_dispfallback)pq2_0_multicol)Correctness
Verified output logits match bit-identically against single-token autoregressive decoding (
kernel_mul_mv_pq2_0_f32). No degradation in output tokens.