Repository navigation
Conversation
|
@karusrus @kiljoy001, thanks for already working this out on #265. Both kernels test bit-exact on an M5 Pro, including the i8mm two-row path. Our plan is to merge #265 first and then #290, which only needs its |
|
Thanks @bri-prism, sounds good. Once #265 is in, I'll rebase this on master and drop the |
PTQ1_0 had only the generic vec_dot on ARM. This adds a NEON kernel: - trits are decoded in 8-bit lanes: ((q + (q >> 1)) >> 1) >> 6 equals (3q) >> 8 for every byte value, so no widening is needed; a 128-trit block becomes eight int8x16 vectors in element order - each 32-wide Q8_0 sub-block is reduced with sdot and accumulated into a float32x4 with vmlaq_n_f32, one horizontal add per row - with __ARM_FEATURE_MATMUL_INT8, PTQ1_0 uses nrows = 2 and the nrc == 2 path computes a 2x2 tile with vmmlaq_s32, reusing decoded trits for two activation columns Snapdragon 7 Gen 4, Ternary Bonsai 2 27B, -t 5: pp64 1.22 -> 2.33 t/s, tg16 1.09 -> 1.31 t/s against the generic path. test-quantize-fns passes; greedy output matches the generic path on the same device. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
f77a146 to
9098196
Compare
|
@bri-prism rebased on prism now that #265 is in. Both ARM aliases in
|
Overview
Follow-up to #265, which dropped its PTQ1_0 NEON kernel and asked for a validated one. This adds a NEON vec_dot for PTQ1_0 x Q8_0 on ARM, with an i8mm path for prompt processing.
Results
Snapdragon 7 Gen 4 (Cortex-A720/A520, dotprod + i8mm), Ternary Bonsai 2 27B PTQ1_0, CPU only, -t 5, runs alternated:
test-backend-ops perf, MUL_MAT m=4096 k=14336: n=1 65.9 GFLOPS, n=512 153.4 GFLOPS.
Tests: with #265 merged on top, test-quantize-fns passes, including its exact packed ternary vec_dot check ("ternary packed dot products: ok (0 failures)"). That check uses nrc == 1. The nrc == 2 (i8mm) path was checked end to end: greedy 24-token output on the 27B is byte-identical to the generic path. The file also builds for armv7a (NEON, no i8mm).
Additional information
Requirements