From 518ad108f0b72bac4f397a486695a5786e33a58d Mon Sep 17 00:00:00 2001 From: zhaoyilun Date: Sat, 19 Sep 2026 15:13:13 +0800 Subject: [PATCH] qwen35: apply the Hadamard inverse to the MTP token-embedding lookup The MTP graph in qwen35 does its own ggml_get_rows() on the token embedding table but never restores the primal basis, while llm_graph_context::build_inp_embd() does exactly that for the trunk (llama-graph.cpp, "a Hadamard-latent embedding table stores rotated rows; restore the primal basis right after the lookup"). On a Hadamard-folded model the draft head therefore consumes embeddings in the rotated basis, and llama_verify_hadamard_graph (correctly) refuses to build the graph, so in-file MTP cannot start at all: W llama_verify_hadamard_graph: latent lookup 'mtp_tok_embd-64' consumed by op=RMS_NORM name='norm-64' E llama_init_from_model: failed to initialize the context: Hadamard-latent table 'token_embd.weight' is read without the inverse transform E common_speculative_init_result: failed to create MTP context Reproduced with prism-b10683-d8f26ee (newest release with published assets, see #193) and still present on b10709-9a9394a, on ProCreations/Ternary-Bonsai-2-27B-MTP (the PQ2_0 target with one MTP layer in the same file) with: llama-server -m Ternary-Bonsai-2-27B-PQ2_0-MTP-Q8_0.gguf --spec-type draft-mtp Fix: apply the same rot + signs inverse used by the trunk path, keyed on the table that was actually looked up (layer.nextn.embed_tokens when present, otherwise model.tok_embd). The same change ships as runtime/bonsai-mtp-embedding.patch in ProCreations/Ternary-Bonsai-2-27B-MTP (and their patched source archive); this is the equivalent for current prism. Credit for finding it belongs there. --- src/models/qwen35.cpp | 14 ++++++++++++++ 1 file changed, 14 insertions(+) diff --git a/src/models/qwen35.cpp b/src/models/qwen35.cpp index c5f816e96268..2f143d367a4d 100644 --- a/src/models/qwen35.cpp +++ b/src/models/qwen35.cpp @@ -634,6 +634,20 @@ llama_model_qwen35::graph_mtp::graph_mtp(const llama_model & model, const llm_gr ggml_tensor * tok_embd_w = layer.nextn.embed_tokens ? layer.nextn.embed_tokens : model.tok_embd; tok_embd = ggml_get_rows(ctx0, tok_embd_w, inp->tokens); + + // a Hadamard-latent embedding table stores rotated rows; restore the primal + // basis right after the lookup (h = s * (H z)), exactly like the trunk path in + // llm_graph_context::build_inp_embd. Without this the draft head consumes + // embeddings in the rotated basis and llama_verify_hadamard_graph refuses the graph. + if (hadamard_inverses) { + const auto it = hadamard_inverses->find(tok_embd_w); + if (it != hadamard_inverses->end()) { + tok_embd = llama_mul_mat_hadamard(ctx0, tok_embd, it->second.rot); + if (it->second.signs) { + tok_embd = ggml_mul(ctx0, tok_embd, it->second.signs); + } + } + } } else { tok_embd = inp->embd; }