Skip to content

qwen35: apply the Hadamard inverse to the MTP token-embedding lookup - #230

Closed
renovys wants to merge 5 commits into
PrismML-Eng:prismfrom
renovys:bc250-qwen35-mtp-inverse
Closed

renovys wants to merge 5 commits into
PrismML-Eng:prismfrom
renovys:bc250-qwen35-mtp-inverse

Conversation

@renovys

@renovys renovys commented Sep 21, 2026 •

Copy link
Copy Markdown

Based on #188 (PQ2_0 support), which is not yet merged, so the diff below #188 shows its commits too; the net new change here is the final commit. Happy to rebase once #188 lands.

qwen35: apply the Hadamard inverse to the MTP token-embedding lookup

Problem — --spec-type draft-mtp cannot start on Hadamard-latent models (patch=required)

For qwen35 models whose token_embd.weight is a Hadamard-latent table
(prism.hadamard.inverse_weight_names, e.g. Ternary-Bonsai-2-27B
PQ2_0-MTP), the MTP draft graph reuses model.tok_embd when the model has
no nextn.embed_tokens (qwen35.nextn_predict_layers = 1).

graph_mtp consumed the raw ggml_get_rows(tok_embd_w) output — the latent
rows — without restoring the primal basis. llama_verify_hadamard_graph
(llama-context.cpp, hadamard_verified per context) detects a latent table
being read without the inverse transform and throws:

std::runtime_error("Hadamard-latent table ... read without the inverse transform")

So on the unpatched tree the first MTP draft-graph build throws, draft
context creation fails, and --spec-type draft-mtp is unusable — not a
slowdown, a startup-correctness failure. Verdict: patch=required.

Change

In src/models/qwen35.cpp graph_mtp, right after the embedding lookup,
apply the same inverse that build_inp_embd() applies for the main graph:

  • if hadamard_inverses contains the embedding tensor, run
    llama_mul_mat_hadamard(ctx0, tok_embd, rot) and multiply signs when present.

This mirrors the existing main-graph path exactly; when the table is not
latent-registered the behavior is unchanged.

Evidence (BC-250, RADV GFX1013, Ternary-Bonsai-2-27B-PQ2_0-MTP-Q8_0.gguf)

arm binary spec result
A unpatched draft-mtp fails by construction (throw path proven above; live attempt aborted on memory-pressure guard, not needed — the throw is deterministic)
B patched draft-mtp 31.10 t/s, draft_n 134, accepted 82 (61.2%)
C patched none 22.97 t/s
  • B vs C output: same prompt, temp 0, n=128 — content sha256 62156b08640d3047
    byte-identical on 3/3 runs. The patch produces correct output, not just a
    faster wrong one.
  • Production run (2026-09-21, live server, ctx 65,536, 2 slots): 3 consecutive
    measures 30.13/30.29/30.28 t/s, draft_n 134 / accepted 82 every run.

Correctness

  • test-backend-ops -b Vulkan0 -o MUL_MAT,MUL_MAT_ID,GET_ROWS: 2310/2310 pass,
    0 FAIL (includes Bonsai-2 real shapes).
  • The Hadamard verifier itself remains the regression tripwire: if the inverse
    were dropped or applied twice, first MTP graph build throws.

Notes for review

  • hadamard_inverses map lookup is per-tensor; non-latent models and
    nextn.embed_tokens-carrying models take the same path as before.
  • Draft context uses its own hadamard_verified flag — this fix is what lets
    that verification pass instead of throwing.

alhnesn and others added 5 commits September 18, 2026 20:36
PTQ1_0 decode on Vulkan went through the generic mul_mat_vec shader, which
decodes every trit with a per-element base-3 loop; on an RX 6750 XT that ran
at ~270 GFLOPS and limited Ternary-Bonsai-2-27B to 4.7 tok/s. This adds a
dedicated q8_1 (mmvq) shader for PTQ1_0 and wires PQ2_0 into the Vulkan
backend, which had no support for it at all.

mul_mat_vecq_ptq1_0.comp:
- two lanes per 128-element block, 4 rows per workgroup, no divergence
- trits decoded with the multiply-by-3 recurrence on two bytes at a time in
  the 16-bit halves of a dword (as the CUDA vec_dot_ptq1_0_q8_1 does)
- the level word is fed to dotPacked4x8AccSat as is: activations are
  pre-masked so the residue bytes are multiplied by zero, which removes the
  per-level extraction ops
- the ternary -1 offset is folded into the q8_1 block sums

PQ2_0: block/packed16 types, dequant shader, mul_mm loader, float and q8_1
mat-vec paths (modeled on Q2_0 with the 128-element group), get_rows.

test-backend-ops: MUL_MAT / MUL_MAT_ID cases for both types at the Bonsai
shapes (k = 1024..17408), odd row counts, batched B and n = 1..8, plus perf
cases at Bonsai shapes.

Measured on RX 6750 XT (RDNA2, AMD proprietary driver, Windows 11),
Ternary-Bonsai-2-27B: PTQ1_0 decode 4.7 -> 40.8 tok/s (39 with -fa on),
PQ2_0 CPU-only -> 38.5 tok/s, prefill 200 t/s. Perplexity through the new
path 6.3916 vs 6.3906 with the original shader; greedy output identical to
the stock build.
The PTQ1_0 q8_1 mat-vec re-ran the whole trit decode for every column of B,
so batched decode (llama-server with several slots, NUM_COLS = 2..8) scaled
poorly: on an RX 6750 XT 8 parallel sequences gave 66 tok/s aggregate, barely
more than 4.

Columns are now handled in passes of CG = 3 (a new specialization constant).
Per block the lane decodes all NUM_ROWS rows level by level once and dots each
packed level with the activations of every column of the pass, so an extra
column costs three loads and three dot products per row and level. Integer
accumulators are folded into the fp32 result whenever the level moves on to a
new q8_1 sub-block scale. Passes are separate loops over the blocks, written
as guarded calls, because the driver hoists the loads of everything in one
block body (which spilled at 8 columns) and does not unroll a loop that
contains the block loop.

The single-column pipeline keeps the original loop (runtime-bounded row loop,
72 VGPRs), so single-sequence decode is unchanged.

RX 6750 XT, Ternary Bonsai 2 27B, llama-batched-bench pp256/tg128, fa on,
q8_0 KV, aggregate decode tok/s (before -> after):
  1 seq  36.6 -> 38.1     4 seq  55.2 -> 73.3
  2 seq  39.7 -> 60.8     8 seq  66.3 -> 86.4
llama-bench tg64: 40.6 (fa off) / 38.7 (fa on), same as before.
Kernel time m=17408 k=5120: n=1 45 us, n=2 57, n=4 120, n=8 215.

test-backend-ops: 271/271 (ptq1_0 + pq2_0, forced mmvq, n = 1..8 incl. row
tail), 138/138 default path; perplexity through the 8-column path 6.2321 vs
6.2294 before (same text, fp32 summation order).

Also adds perf cases with n = 2, 4, 8 for PTQ1_0 / PQ2_0 / Q4_0.
…ments

- ggml_vk_get_dequantize_mul_mat_vec_id: add PTQ1_0 and PQ2_0 to the q8_1
  whitelist. The id q8_1 pipelines were created but never selected, so
  MUL_MAT_ID always fell back to the float path.
- test-backend-ops: add n = 4 to the PTQ1_0/PQ2_0 cases (a three-column pass
  followed by the single-column tail).
- Shorten the shader comments and the coopmat2 note in vulkan-shaders-gen.cpp.

test-backend-ops MUL_MAT/MUL_MAT_ID/GET_ROWS for ptq1_0 + pq2_0: 284/284 on
the default path and 284/284 with GGML_VK_FORCE_MMVQ=1 (RX 6750 XT); the
mul_mat_vec_id_ptq1_0_q8_1_f32 pipeline is now compiled and used.
For qwen35 models whose token_embd.weight is a Hadamard-latent table
(prism.hadamard.inverse_weight_names) and which have no nextn.embed_tokens,
graph_mtp reused model.tok_embd and consumed the raw ggml_get_rows() latent
rows. llama_verify_hadamard_graph then throws on the first MTP draft-graph
build, so --spec-type draft-mtp is unusable on these models.

Apply the same inverse that build_inp_embd() uses for the main graph right
after the embedding lookup; behavior is unchanged when the table is not
latent-registered. Verified on Ternary-Bonsai-2-27B-PQ2_0-MTP (BC-250,
gfx1013): draft-mtp starts and runs at 31.1 t/s (61.2% accept), output
byte-identical (content sha256) to the spec-off run over 3/3 runs.
test-backend-ops MUL_MAT/MUL_MAT_ID/GET_ROWS 2310/2310 pass.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
@renovys

renovys commented Sep 22, 2026 •

Copy link
Copy Markdown
Author

Validated the equivalent fix on a live BC-250 (gfx1013) service (via 62e04aad1 in our tree): draft-mtp runs at ~30 t/s with ~61% accept rate and no throw.

@renovys

renovys commented Sep 26, 2026

Copy link
Copy Markdown
Author

Closing as a duplicate to keep the queue clean: #205 has landed in prism (422590f, merge of zhaoyilun's mtp-hadamard-embedding) carrying the same inverse-Hadamard on the token embeddings the qwen35 MTP draft graph reads. We verified the equivalent on our BC-250 service, so folding this copy into that decision. Thanks all for the reviews and cross-references.

@renovys renovys closed this Sep 26, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants