Skip to content

Rename llama_mxint4.toml to llama_mxint8.toml and use dtype= for from_pretrained - #4

Open
Shreyas8612 wants to merge 3 commits into
mainfrom
fix/mxint8-config-name-and-dtype-arg
Open

Shreyas8612 wants to merge 3 commits into
mainfrom
fix/mxint8-config-name-and-dtype-arg

Conversation

@Shreyas8612

Copy link
Copy Markdown
Collaborator

The default LLaMA quantisation config was named llama_mxint4.toml and its
header said "MXInt4", but every block in it sets weight_width = 8 and
data_in_width = 8 with block size 32, i.e. it is an MXInt8 W8A8 config.
The name made it easy to misreport results produced with the default.

  • Rename the file to llama_mxint8.toml and rewrite the header to describe
    the actual widths. Config contents are unchanged.
  • Update every default and usage example pointing at the old path
    (eval_* CLIs, quant_hf_serve.py, search_rotation.py, QUANT_HF_SERVE.md,
    run_lm_eval_phase.sh). A repo-wide grep finds no remaining references.
  • transformers 5 renamed the torch_dtype kwarg of from_pretrained to
    dtype and warns on every model load; switch setup_model, the serve
    CLI and the LLaDA evaluator to dtype=. Values passed are unchanged.
  • Add .venv/, .pytest_cache/ and *.vcd to .gitignore.

Verified with py_compile on all changed modules, bash -n on the script,
and --help on the CLIs (defaults now show llama_mxint8.toml). Two CLIs
(eval_evalplus, eval_phase_bfcl) need optional extras to import at all;
their defaults were checked in the source.

Shreyas8612 added 3 commits September 15, 2026 20:58
The default LLaMA quantisation config was named llama_mxint4.toml and
its header said "MXInt4", but every block in the file sets
weight_width = 8 and data_in_width = 8 (block size 32), i.e. it is an
MXInt8 W8A8 config. The misleading name made it easy to assume results
produced with the default config were 4-bit.

Rename the file to llama_mxint8.toml, rewrite the header comment to
describe the actual widths, and update every default and usage example
that pointed at the old path: the eval_* CLIs, quant_hf_serve.py,
search_rotation.py, QUANT_HF_SERVE.md and run_lm_eval_phase.sh. The
config contents are unchanged.

Verified with py_compile on every touched module, `bash -n` on the
script, `--help` on the CLIs (defaults now show llama_mxint8.toml) and
a repo-wide grep confirming no remaining llama_mxint4 references.
transformers 5 renamed the `torch_dtype` keyword of `from_pretrained`
to `dtype`; passing the old name still works but logs
"`torch_dtype` is deprecated! Use `dtype` instead!" on every model
load. Pass `dtype=` in setup_model so the warning goes away and the
call keeps working when the alias is removed.

Also add .venv/, .pytest_cache/ and *.vcd to .gitignore so a local
virtual environment, pytest cache and simulator waveform dumps are not
picked up by `git status`.

Verified with py_compile on quant_eval/utils.py and by checking the
installed transformers modeling_utils, which accepts `dtype` and only
emits the deprecation warning for `torch_dtype`.
quant_hf_serve.py and the LLaDA evaluator still passed torch_dtype= to
from_pretrained, so every model load in those paths logged the
transformers 5 deprecation warning that the previous commit removed from
setup_model. Switch the five remaining call sites to dtype=; the value
passed is unchanged.

Verified with py_compile on both modules.
Copilot AI lite review requested due to automatic review settings September 17, 2026 01:44

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Changes recommended

dtype= is incompatible with the supported Transformers 4.x dependency range without compatibility handling or a raised dependency floor.

Get a fresh assessment by requesting another Copilot review.

Pull request overview

Renames the default LLaMA quantization configuration to accurately reflect MXInt8 W8A8 settings and updates model-loading arguments.

Changes:

  • Renames configuration references to llama_mxint8.toml.
  • Replaces torch_dtype= with dtype=.
  • Updates CLI defaults, scripts, and documentation.
  • Adds common development artifacts to .gitignore.
File summaries
File Description
quant_eval/utils.py Updates shared model loading.
quant_eval/scripts/run_lm_eval_phase.sh Updates the default configuration path.
quant_eval/eval/llada/eval_llada.py Updates LLaDA model loading.
quant_eval/configs/llama_mxint8.toml Renames and clarifies the configuration.
quant_eval/cli/search_rotation.py Updates configuration references.
quant_eval/cli/quant_hf_serve.py Updates serving loads and examples.
quant_eval/cli/QUANT_HF_SERVE.md Updates documentation paths.
quant_eval/cli/eval_ppl.py Updates configuration references.
quant_eval/cli/eval_phase_lm.py Updates configuration references.
quant_eval/cli/eval_phase_bfcl.py Updates configuration references.
quant_eval/cli/eval_osworld.py Updates configuration references.
quant_eval/cli/eval_lm.py Updates configuration references.
quant_eval/cli/eval_llada.py Updates configuration references.
quant_eval/cli/eval_evalplus.py Updates configuration references.
quant_eval/cli/eval_dllm.py Updates configuration references.
.gitignore Adds development artifact exclusions.
Review details
  • Files reviewed: 15/16 changed files
  • Comments generated: 3
  • Review effort level: Lite

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

model = AutoModelForCausalLM.from_pretrained(
model_name,
torch_dtype=torch_dtype,
dtype=torch_dtype,
self.model = LLaDAModelLM.from_pretrained(
model_path, trust_remote_code=True,
torch_dtype=torch.bfloat16, config=config, **model_kwargs,
dtype=torch.bfloat16, config=config, **model_kwargs,
Comment thread quant_eval/utils.py
model = AutoModelForCausalLM.from_pretrained(
model_name,
torch_dtype=dtype,
dtype=dtype,

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants