Skip to content

Load GLM-5-Next GGUFs converted with the earlier glm5next arch name - #239

Open
danielhanchen wants to merge 1 commit into
base/upstream-1fb7ef3e3from
compat/glm5next-legacy-ggufs
Open

danielhanchen wants to merge 1 commit into
base/upstream-1fb7ef3e3from
compat/glm5next-legacy-ggufs

Conversation

@danielhanchen

Copy link
Copy Markdown
Member

Summary

Upstream merged GLM-5-Next as glm5-next (ggml-org#27773), superseding our ggml-org#27754, which used glm5next. Every GLM-5.3-Flash GGUF and mmproj on unsloth/GLM-5.3-Flash-GGUF was converted with the old names, so mainline refuses them:

llama_model_load: error loading model: unknown model architecture: 'glm5next'

The repo also has Shard_Rewrite/ replacements (new shard 1 + mmproj). This PR makes both the original files and the rewritten ones load, so users who already downloaded the old shards keep working and the swap stays optional.

What differs between the old and new files

Checked on all 12 quants (BF16, Q8_0, UD-IQ1_S ... UD-Q6_K_XL) by reading the GGUF headers:

  • Shard 1 holds no tensors. Old vs rewritten shard 1 differ only in the arch name and the KV prefix (glm5next.* -> glm5-next.*), plus one key upstream never reads (index_share_mtp). Tensor names and shapes in shards 2..N are identical to what upstream expects.
  • mmproj: projector_type glm5next -> glm5v, clip.vision.swiglu_limit -> clip.vision.swiglu_clamp, and the rewritten file adds image_min_pixels = 12544 / image_max_pixels = 6272000. Same 348 tensors.

Changes

  • llm_arch_from_string: glm5next maps to LLM_ARCH_GLM5_NEXT.
  • LLM_KV takes an optional prefix; the loader passes the arch name as written in the file, so glm5next.* keys resolve. Files with the current name are unaffected.
  • clip_projector_type_from_string: glm5next maps to PROJECTOR_TYPE_GLM5V (outside PROJECTOR_TYPE_NAMES, so the registry test from mtmd: test that every projector is registered and uniquely named #176 still sees unique names).
  • GLM5V hparams: a file without swiglu_clamp reads swiglu_limit and gets the default pixel budget. Files with swiglu_clamp keep the existing required keys (legacy is detected by key presence, so an explicit swiglu_clamp = 0 still means unclamped).

Testing

GLM-5.3-Flash UD-IQ1_S on one B200, greedy, NVIDIA_TF32_OVERRIDE=0:

build original shard 1 + mmproj rewritten shard 1 + mmproj
upstream b11368 fails: unknown arch glm5next loads
this PR loads loads
  • Text and image output are byte-identical between original and rewritten files on this PR, and identical to upstream b11368 running the rewritten files. The image test (red circle, blue square, "UNSLOTH 42") is read correctly; old and new mmproj give the same answer.
  • llama-perplexity KLD over 10 x 4096 wikitext-2 chunks: original vs rewritten files on this PR score identically (mean KLD, PPL and top-p all equal), i.e. the compat path adds nothing numerically.
  • test-llama-archs (133/133) and test-mtmd-impl pass.

Base is b11368 (base/upstream-1fb7ef3e3). The full nightly pin set merges cleanly with this PR on top.

GLM-5.3-Flash GGUFs and mmproj files made before upstream GLM-5-Next
support differ from the current ones only in metadata: the arch name
glm5next (and so the KV prefix glm5next.*), the projector name glm5next,
clip.vision.swiglu_limit instead of swiglu_clamp, and no image pixel
budget keys. Tensor names and shapes are identical.

Accept glm5next as LLM_ARCH_GLM5_NEXT and read KVs under the prefix the
file was written with. For the projector, map glm5next to GLM5V, read
swiglu_limit when swiglu_clamp is absent, and default the pixel budget to
what the current converter writes (12544 / 6272000). Files with the
current names load exactly as before.
@danielhanchen
danielhanchen requested a review from CISC as a code owner October 4, 2026 05:35
@chatgpt-codex-connector

Copy link
Copy Markdown

You have reached your Codex usage limits for security reviews. Please try again later.

@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Oct 4, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-10-04T05:41:33.070377Z bec2164 PR opened
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: bec2164445

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread src/llama-arch.cpp
Comment on lines +1072 to +1073
if (name == "glm5next") { // pre-upstream converter
return LLM_ARCH_GLM5_NEXT;

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Add regression tests for the legacy metadata aliases

This compatibility fix has no automated regression coverage for either affected path: loading a model whose architecture and hyperparameter prefix are glm5next, or loading an mmproj that uses projector type glm5next, swiglu_limit, and no pixel-budget keys. Existing architecture and mtmd tests can still pass if either alias path stops working, so add synthetic legacy-metadata cases that fail on the parent commit and verify both loaders.

Useful? React with 👍 / 👎.

Comment thread src/llama-arch.cpp
Comment on lines +1072 to +1073
if (name == "glm5next") { // pre-upstream converter
return LLM_ARCH_GLM5_NEXT;

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Canonicalize the alias before creating architecture filters

When test-llama-archs is invoked with --arch glm5next, this new branch makes llm_arch_from_string() report a known architecture, so the argument parser builds the exact filter ^glm5next$. The test cases are matched against llm_arch_name(LLM_ARCH_GLM5_NEXT), which is glm5-next, so both normal and --save modes select zero cases and exit successfully. Return or expose the canonical name to exact-filter callers instead of making the alias appear canonical.

Useful? React with 👍 / 👎.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant