Conversation
PTQ1_0 decode on Vulkan went through the generic mul_mat_vec shader, which decodes every trit with a per-element base-3 loop; on an RX 6750 XT that ran at ~270 GFLOPS and limited Ternary-Bonsai-2-27B to 4.7 tok/s. This adds a dedicated q8_1 (mmvq) shader for PTQ1_0 and wires PQ2_0 into the Vulkan backend, which had no support for it at all. mul_mat_vecq_ptq1_0.comp: - two lanes per 128-element block, 4 rows per workgroup, no divergence - trits decoded with the multiply-by-3 recurrence on two bytes at a time in the 16-bit halves of a dword (as the CUDA vec_dot_ptq1_0_q8_1 does) - the level word is fed to dotPacked4x8AccSat as is: activations are pre-masked so the residue bytes are multiplied by zero, which removes the per-level extraction ops - the ternary -1 offset is folded into the q8_1 block sums PQ2_0: block/packed16 types, dequant shader, mul_mm loader, float and q8_1 mat-vec paths (modeled on Q2_0 with the 128-element group), get_rows. test-backend-ops: MUL_MAT / MUL_MAT_ID cases for both types at the Bonsai shapes (k = 1024..17408), odd row counts, batched B and n = 1..8, plus perf cases at Bonsai shapes. Measured on RX 6750 XT (RDNA2, AMD proprietary driver, Windows 11), Ternary-Bonsai-2-27B: PTQ1_0 decode 4.7 -> 40.8 tok/s (39 with -fa on), PQ2_0 CPU-only -> 38.5 tok/s, prefill 200 t/s. Perplexity through the new path 6.3916 vs 6.3906 with the original shader; greedy output identical to the stock build.
The PTQ1_0 q8_1 mat-vec re-ran the whole trit decode for every column of B, so batched decode (llama-server with several slots, NUM_COLS = 2..8) scaled poorly: on an RX 6750 XT 8 parallel sequences gave 66 tok/s aggregate, barely more than 4. Columns are now handled in passes of CG = 3 (a new specialization constant). Per block the lane decodes all NUM_ROWS rows level by level once and dots each packed level with the activations of every column of the pass, so an extra column costs three loads and three dot products per row and level. Integer accumulators are folded into the fp32 result whenever the level moves on to a new q8_1 sub-block scale. Passes are separate loops over the blocks, written as guarded calls, because the driver hoists the loads of everything in one block body (which spilled at 8 columns) and does not unroll a loop that contains the block loop. The single-column pipeline keeps the original loop (runtime-bounded row loop, 72 VGPRs), so single-sequence decode is unchanged. RX 6750 XT, Ternary Bonsai 2 27B, llama-batched-bench pp256/tg128, fa on, q8_0 KV, aggregate decode tok/s (before -> after): 1 seq 36.6 -> 38.1 4 seq 55.2 -> 73.3 2 seq 39.7 -> 60.8 8 seq 66.3 -> 86.4 llama-bench tg64: 40.6 (fa off) / 38.7 (fa on), same as before. Kernel time m=17408 k=5120: n=1 45 us, n=2 57, n=4 120, n=8 215. test-backend-ops: 271/271 (ptq1_0 + pq2_0, forced mmvq, n = 1..8 incl. row tail), 138/138 default path; perplexity through the 8-column path 6.2321 vs 6.2294 before (same text, fp32 summation order). Also adds perf cases with n = 2, 4, 8 for PTQ1_0 / PQ2_0 / Q4_0.
…ments - ggml_vk_get_dequantize_mul_mat_vec_id: add PTQ1_0 and PQ2_0 to the q8_1 whitelist. The id q8_1 pipelines were created but never selected, so MUL_MAT_ID always fell back to the float path. - test-backend-ops: add n = 4 to the PTQ1_0/PQ2_0 cases (a three-column pass followed by the single-column tail). - Shorten the shader comments and the coopmat2 note in vulkan-shaders-gen.cpp. test-backend-ops MUL_MAT/MUL_MAT_ID/GET_ROWS for ptq1_0 + pq2_0: 284/284 on the default path and 284/284 with GGML_VK_FORCE_MMVQ=1 (RX 6750 XT); the mul_mat_vec_id_ptq1_0_q8_1_f32 pipeline is now compiled and used.
For qwen35 models whose token_embd.weight is a Hadamard-latent table (prism.hadamard.inverse_weight_names) and which have no nextn.embed_tokens, graph_mtp reused model.tok_embd and consumed the raw ggml_get_rows() latent rows. llama_verify_hadamard_graph then throws on the first MTP draft-graph build, so --spec-type draft-mtp is unusable on these models. Apply the same inverse that build_inp_embd() uses for the main graph right after the embedding lookup; behavior is unchanged when the table is not latent-registered. Verified on Ternary-Bonsai-2-27B-PQ2_0-MTP (BC-250, gfx1013): draft-mtp starts and runs at 31.1 t/s (61.2% accept), output byte-identical (content sha256) to the spec-off run over 3/3 runs. test-backend-ops MUL_MAT/MUL_MAT_ID/GET_ROWS 2310/2310 pass. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Author
|
Validated the equivalent fix on a live BC-250 (gfx1013) service (via 62e04aad1 in our tree): draft-mtp runs at ~30 t/s with ~61% accept rate and no throw. |
Author
|
Closing as a duplicate to keep the queue clean: #205 has landed in |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
qwen35: apply the Hadamard inverse to the MTP token-embedding lookup
Problem —
--spec-type draft-mtpcannot start on Hadamard-latent models (patch=required)For qwen35 models whose
token_embd.weightis a Hadamard-latent table(
prism.hadamard.inverse_weight_names, e.g. Ternary-Bonsai-2-27BPQ2_0-MTP), the MTP draft graph reuses
model.tok_embdwhen the model hasno
nextn.embed_tokens(qwen35.nextn_predict_layers = 1).graph_mtpconsumed the rawggml_get_rows(tok_embd_w)output — the latentrows — without restoring the primal basis.
llama_verify_hadamard_graph(llama-context.cpp,
hadamard_verifiedper context) detects a latent tablebeing read without the inverse transform and throws:
So on the unpatched tree the first MTP draft-graph build throws, draft
context creation fails, and
--spec-type draft-mtpis unusable — not aslowdown, a startup-correctness failure. Verdict: patch=required.
Change
In
src/models/qwen35.cppgraph_mtp, right after the embedding lookup,apply the same inverse that
build_inp_embd()applies for the main graph:hadamard_inversescontains the embedding tensor, runllama_mul_mat_hadamard(ctx0, tok_embd, rot)and multiplysignswhen present.This mirrors the existing main-graph path exactly; when the table is not
latent-registered the behavior is unchanged.
Evidence (BC-250, RADV GFX1013, Ternary-Bonsai-2-27B-PQ2_0-MTP-Q8_0.gguf)
62156b08640d3047byte-identical on 3/3 runs. The patch produces correct output, not just a
faster wrong one.
measures 30.13/30.29/30.28 t/s, draft_n 134 / accepted 82 every run.
Correctness
test-backend-ops -b Vulkan0 -o MUL_MAT,MUL_MAT_ID,GET_ROWS: 2310/2310 pass,0 FAIL (includes Bonsai-2 real shapes).
were dropped or applied twice, first MTP graph build throws.
Notes for review
hadamard_inversesmap lookup is per-tensor; non-latent models andnextn.embed_tokens-carrying models take the same path as before.hadamard_verifiedflag — this fix is what letsthat verification pass instead of throwing.