Conversation
dwymark
marked this pull request as ready for review
September 23, 2026 16:00
Carry the explicit SIMD PQ2_0 and PTQ1_0 decode kernels, vectorized single-row state copies, the FP16-input GEMM switch, the attention decode knob, the logit probe, the comparison and perf scripts, and the model-shaped test cases onto de7a01b's 16-lane PQ2_0 kernel and gating fixes. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WskiHpEsZkidixZe6TfM5R
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WskiHpEsZkidixZe6TfM5R
Each experimental kernel now carries a short signpost naming what it does and which environment variable turns it on, and the remaining comments state present behaviour rather than the case they guard against. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WskiHpEsZkidixZe6TfM5R
dwymark
force-pushed
the
pq2-sustained-r
branch
from
September 23, 2026 19:57
290176f to
1bd1b5d
Compare
Owner
Author
|
Closing: PrismML-Eng#235 was squash-merged into prism, so the fast-forward offer no longer applies. The kernels move to a standalone pull request against prism. Written by Claude on Daniel's instruction. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
This PR adds opt-in explicit-SIMD decode kernels for PTQ1_0 and PQ2_0 to the SYCL backend, plus an FP16-input prefill path for both formats. Every path is off by default and enabled by environment variable.
On an Intel Arc 140V at 450c5c3, PQ2_0 decode goes from 6.8 to 9.0 tokens per second and prefill from 50 to 145. PTQ1_0 goes from 5.7 to 6.9 and from 48 to 135. FP32 output is bit-identical to the c5f6c97 arithmetic for both formats.
The logit probe, comparison scripts, and benchmark runner are the tooling behind those numbers.
For the upstream author
This branch is the head of PrismML-Eng#235, 42ce76d01, plus three commits. It fast-forwards onto
sycl-ptq1-pq2-mmvqwhile that commit is still the head:Push the result and the commits land in PrismML-Eng#235 automatically. Every path added here is off by default, so the merge changes no behavior on its own. There is no expectation that this is merged; cherry-pick or copy without credit as you prefer.