Repository navigation
Rename llama_mxint4.toml to llama_mxint8.toml and use dtype= for from_pretrained - #4
Open
Shreyas8612 wants to merge 3 commits into
Open
Shreyas8612 wants to merge 3 commits into
Shreyas8612 wants to merge 3 commits into
Conversation
added 3 commits
September 15, 2026 20:58
The default LLaMA quantisation config was named llama_mxint4.toml and its header said "MXInt4", but every block in the file sets weight_width = 8 and data_in_width = 8 (block size 32), i.e. it is an MXInt8 W8A8 config. The misleading name made it easy to assume results produced with the default config were 4-bit. Rename the file to llama_mxint8.toml, rewrite the header comment to describe the actual widths, and update every default and usage example that pointed at the old path: the eval_* CLIs, quant_hf_serve.py, search_rotation.py, QUANT_HF_SERVE.md and run_lm_eval_phase.sh. The config contents are unchanged. Verified with py_compile on every touched module, `bash -n` on the script, `--help` on the CLIs (defaults now show llama_mxint8.toml) and a repo-wide grep confirming no remaining llama_mxint4 references.
transformers 5 renamed the `torch_dtype` keyword of `from_pretrained` to `dtype`; passing the old name still works but logs "`torch_dtype` is deprecated! Use `dtype` instead!" on every model load. Pass `dtype=` in setup_model so the warning goes away and the call keeps working when the alias is removed. Also add .venv/, .pytest_cache/ and *.vcd to .gitignore so a local virtual environment, pytest cache and simulator waveform dumps are not picked up by `git status`. Verified with py_compile on quant_eval/utils.py and by checking the installed transformers modeling_utils, which accepts `dtype` and only emits the deprecation warning for `torch_dtype`.
quant_hf_serve.py and the LLaDA evaluator still passed torch_dtype= to from_pretrained, so every model load in those paths logged the transformers 5 deprecation warning that the previous commit removed from setup_model. Switch the five remaining call sites to dtype=; the value passed is unchanged. Verified with py_compile on both modules.
There was a problem hiding this comment.
🟡 Changes recommended
dtype= is incompatible with the supported Transformers 4.x dependency range without compatibility handling or a raised dependency floor.
Get a fresh assessment by requesting another Copilot review.
Pull request overview
Renames the default LLaMA quantization configuration to accurately reflect MXInt8 W8A8 settings and updates model-loading arguments.
Changes:
- Renames configuration references to
llama_mxint8.toml. - Replaces
torch_dtype=withdtype=. - Updates CLI defaults, scripts, and documentation.
- Adds common development artifacts to
.gitignore.
File summaries
| File | Description |
|---|---|
quant_eval/utils.py |
Updates shared model loading. |
quant_eval/scripts/run_lm_eval_phase.sh |
Updates the default configuration path. |
quant_eval/eval/llada/eval_llada.py |
Updates LLaDA model loading. |
quant_eval/configs/llama_mxint8.toml |
Renames and clarifies the configuration. |
quant_eval/cli/search_rotation.py |
Updates configuration references. |
quant_eval/cli/quant_hf_serve.py |
Updates serving loads and examples. |
quant_eval/cli/QUANT_HF_SERVE.md |
Updates documentation paths. |
quant_eval/cli/eval_ppl.py |
Updates configuration references. |
quant_eval/cli/eval_phase_lm.py |
Updates configuration references. |
quant_eval/cli/eval_phase_bfcl.py |
Updates configuration references. |
quant_eval/cli/eval_osworld.py |
Updates configuration references. |
quant_eval/cli/eval_lm.py |
Updates configuration references. |
quant_eval/cli/eval_llada.py |
Updates configuration references. |
quant_eval/cli/eval_evalplus.py |
Updates configuration references. |
quant_eval/cli/eval_dllm.py |
Updates configuration references. |
.gitignore |
Adds development artifact exclusions. |
Review details
- Files reviewed: 15/16 changed files
- Comments generated: 3
- Review effort level: Lite
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
| model = AutoModelForCausalLM.from_pretrained( | ||
| model_name, | ||
| torch_dtype=torch_dtype, | ||
| dtype=torch_dtype, |
| self.model = LLaDAModelLM.from_pretrained( | ||
| model_path, trust_remote_code=True, | ||
| torch_dtype=torch.bfloat16, config=config, **model_kwargs, | ||
| dtype=torch.bfloat16, config=config, **model_kwargs, |
| model = AutoModelForCausalLM.from_pretrained( | ||
| model_name, | ||
| torch_dtype=dtype, | ||
| dtype=dtype, |
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The default LLaMA quantisation config was named llama_mxint4.toml and its
header said "MXInt4", but every block in it sets weight_width = 8 and
data_in_width = 8 with block size 32, i.e. it is an MXInt8 W8A8 config.
The name made it easy to misreport results produced with the default.
the actual widths. Config contents are unchanged.
(eval_* CLIs, quant_hf_serve.py, search_rotation.py, QUANT_HF_SERVE.md,
run_lm_eval_phase.sh). A repo-wide grep finds no remaining references.
torch_dtypekwarg of from_pretrained todtypeand warns on every model load; switch setup_model, the serveCLI and the LLaDA evaluator to
dtype=. Values passed are unchanged.Verified with py_compile on all changed modules, bash -n on the script,
and --help on the CLIs (defaults now show llama_mxint8.toml). Two CLIs
(eval_evalplus, eval_phase_bfcl) need optional extras to import at all;
their defaults were checked in the source.