Load GLM-5-Next GGUFs converted with the earlier glm5next arch name - #239
danielhanchen wants to merge 1 commit into
Conversation
GLM-5.3-Flash GGUFs and mmproj files made before upstream GLM-5-Next support differ from the current ones only in metadata: the arch name glm5next (and so the KV prefix glm5next.*), the projector name glm5next, clip.vision.swiglu_limit instead of swiglu_clamp, and no image pixel budget keys. Tensor names and shapes are identical. Accept glm5next as LLM_ARCH_GLM5_NEXT and read KVs under the prefix the file was written with. For the projector, map glm5next to GLM5V, read swiglu_limit when swiglu_clamp is absent, and default the pixel budget to what the current converter writes (12544 / 6272000). Files with the current names load exactly as before.
|
You have reached your Codex usage limits for security reviews. Please try again later. |
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: bec2164445
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| if (name == "glm5next") { // pre-upstream converter | ||
| return LLM_ARCH_GLM5_NEXT; |
There was a problem hiding this comment.
Add regression tests for the legacy metadata aliases
This compatibility fix has no automated regression coverage for either affected path: loading a model whose architecture and hyperparameter prefix are glm5next, or loading an mmproj that uses projector type glm5next, swiglu_limit, and no pixel-budget keys. Existing architecture and mtmd tests can still pass if either alias path stops working, so add synthetic legacy-metadata cases that fail on the parent commit and verify both loaders.
Useful? React with 👍 / 👎.
| if (name == "glm5next") { // pre-upstream converter | ||
| return LLM_ARCH_GLM5_NEXT; |
There was a problem hiding this comment.
Canonicalize the alias before creating architecture filters
When test-llama-archs is invoked with --arch glm5next, this new branch makes llm_arch_from_string() report a known architecture, so the argument parser builds the exact filter ^glm5next$. The test cases are matched against llm_arch_name(LLM_ARCH_GLM5_NEXT), which is glm5-next, so both normal and --save modes select zero cases and exit successfully. Return or expose the canonical name to exact-filter callers instead of making the alias appear canonical.
Useful? React with 👍 / 👎.
Summary
Upstream merged GLM-5-Next as
glm5-next(ggml-org#27773), superseding our ggml-org#27754, which usedglm5next. Every GLM-5.3-Flash GGUF and mmproj onunsloth/GLM-5.3-Flash-GGUFwas converted with the old names, so mainline refuses them:The repo also has
Shard_Rewrite/replacements (new shard 1 + mmproj). This PR makes both the original files and the rewritten ones load, so users who already downloaded the old shards keep working and the swap stays optional.What differs between the old and new files
Checked on all 12 quants (BF16, Q8_0, UD-IQ1_S ... UD-Q6_K_XL) by reading the GGUF headers:
glm5next.*->glm5-next.*), plus one key upstream never reads (index_share_mtp). Tensor names and shapes in shards 2..N are identical to what upstream expects.projector_typeglm5next->glm5v,clip.vision.swiglu_limit->clip.vision.swiglu_clamp, and the rewritten file addsimage_min_pixels = 12544/image_max_pixels = 6272000. Same 348 tensors.Changes
llm_arch_from_string:glm5nextmaps toLLM_ARCH_GLM5_NEXT.LLM_KVtakes an optional prefix; the loader passes the arch name as written in the file, soglm5next.*keys resolve. Files with the current name are unaffected.clip_projector_type_from_string:glm5nextmaps toPROJECTOR_TYPE_GLM5V(outsidePROJECTOR_TYPE_NAMES, so the registry test from mtmd: test that every projector is registered and uniquely named #176 still sees unique names).swiglu_clampreadsswiglu_limitand gets the default pixel budget. Files withswiglu_clampkeep the existing required keys (legacy is detected by key presence, so an explicitswiglu_clamp = 0still means unclamped).Testing
GLM-5.3-Flash UD-IQ1_S on one B200, greedy,
NVIDIA_TF32_OVERRIDE=0:glm5nextllama-perplexityKLD over 10 x 4096 wikitext-2 chunks: original vs rewritten files on this PR score identically (mean KLD, PPL and top-p all equal), i.e. the compat path adds nothing numerically.test-llama-archs(133/133) andtest-mtmd-implpass.Base is b11368 (
base/upstream-1fb7ef3e3). The full nightly pin set merges cleanly with this PR on top.