Skip to content

[MLPerf 6.1][GRPO] Offline evaluation for Qwen3.5-397B GRPO - #906

Merged
ShriyaRishab merged 3 commits into
mlcommons:masterfrom
jepio:qwen35-offline-eval
Sep 28, 2026
Merged

ShriyaRishab merged 3 commits into
mlcommons:masterfrom
jepio:qwen35-offline-eval

Conversation

@jepio

@jepio jepio commented Sep 14, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Stacked on #904: this PR contains that PR's commit plus the two offline-evaluation commits at the tip; the diff reduces to those two once #904 merges.

The benchmark rules (mlcommons/training_policies#596) require offline evaluation for qwen35_397b_grpo: training saves a checkpoint after every step from the first-evaluation step CEIL(2.5 + 3840 / global_batch_size) until it stops, each checkpoint records the timestamp of its latest weight update, and the saved checkpoints are evaluated post-run until one reaches the target. run_stop is emitted with the passing checkpoint's weight-update timestamp, so checkpoint writing and evaluation time are excluded from the measured score.

This PR advances the NeMo-RL pin to f987c0596 (NVIDIA-NeMo/RL), which implements the requirement in the reference launcher, and updates the documentation accordingly.

Implementation (in the pinned NeMo-RL commit)

  • DEFERRED_OFFLINE_EVAL=1 disables inline validation and saves weights-only checkpoints every step from grpo.val_start_at through the final step; each checkpoint records its weight-update end time in training_info.json.
  • After training stops, a second driver in the same allocation and Ray cluster restores each checkpoint without optimizer state, refits vLLM, and validates in step order, stopping at the first checkpoint that reaches the target.
  • The evaluator appends its events to the training mllog and emits the single run_stop backdated to the passing checkpoint's weight-update timestamp (aborted at the final checkpoint's timestamp when none crosses).
  • Deferred checkpoints are removed at teardown unless RETAIN_GRPO_CHECKPOINTS=1.

Signed-off-by: Jeremi Piotrowski <jpiotrowski@nvidia.com>
Advance the NeMo-RL pin to f987c0596 (mlperf-training-qwen35-next-v2), which
adds deferred (offline) checkpoint evaluation to the reference launcher.
The rules require offline evaluation for qwen35_397b_grpo: checkpoints every
step from the first-evaluation step through the end of training, evaluated
post-run until one reaches the target; run_stop is backdated to the passing
checkpoint's weight-update timestamp. Describe how the pinned reference
implements this, and correct the submodule path in the checkout instructions.
@jepio
jepio requested review from a team as code owners September 14, 2026 15:33
@github-actions

Copy link
Copy Markdown

MLCommons CLA bot All contributors have signed the MLCommons CLA ✍️ ✅

@jepio

jepio commented Sep 15, 2026 •

Copy link
Copy Markdown
Contributor Author

260913135524191054819_1_mllog.log

here's a reference run mllog output with offline validation

@RissyRan

Copy link
Copy Markdown

260913135524191054819_1_mllog.log

here's a reference run mllog output with offline validation

I am still a little bit confused about offline eval rule. I see eval before run_stop. I thought we moved it totally to be offline, and we just need to store the checkpoint before it ends. Could you help me understand more?

Does this eval add back in post-processing step to the original logging? Thank you very much!

:::MLLOG {"namespace": "", "time_ms": 1789340490400, "event_type": "INTERVAL_END", "key": "block_stop", "value": null, "metadata": {"file": "/opt/nemo-rl/nemo_rl/algorithms/mlperf_grpo_logging.py", "lineno": 134, "samples_count": 4864, "step": 19}}
:::MLLOG {"namespace": "", "time_ms": 1789341041466, "event_type": "INTERVAL_START", "key": "eval_start", "value": null, "metadata": {"file": "/opt/nemo-rl/nemo_rl/algorithms/mlperf_grpo_deferred.py", "lineno": 301, "samples_count": 4608, "step": 18}}
:::MLLOG {"namespace": "", "time_ms": 1789342000941, "event_type": "POINT_IN_TIME", "key": "eval_accuracy", "value": 0.7410358786582947, "metadata": {"file": "/opt/nemo-rl/nemo_rl/algorithms/mlperf_grpo_deferred.py", "lineno": 315, "samples_count": 4608}}
:::MLLOG {"namespace": "", "time_ms": 1789342000944, "event_type": "INTERVAL_END", "key": "eval_stop", "value": null, "metadata": {"file": "/opt/nemo-rl/nemo_rl/algorithms/mlperf_grpo_deferred.py", "lineno": 318, "samples_count": 4608, "step": 18}}
:::MLLOG {"namespace": "", "time_ms": 1789340105384, "event_type": "INTERVAL_END", "key": "run_stop", "value": null, "metadata": {"file": "/opt/nemo-rl/nemo_rl/algorithms/mlperf_grpo_deferred.py", "lineno": 373, "status": "success", "samples_count": 4608}}

@jepio

jepio commented Sep 17, 2026

Copy link
Copy Markdown
Contributor Author

@RissyRan

I am still a little bit confused about offline eval rule. I see eval before run_stop. I thought we moved it totally to be offline, and we just need to store the checkpoint before it ends. Could you help me understand more?

The reference implementation runs one python process for training which emits block_stop when it finishes the last step and captures the checkpoint. This marks the end of training. The training process doesn't emit run_stop before it terminates. Then a second python process starts which runs validation over the checkpoints. Once the correct checkpoint is found, the run_stop event is emitted. The run_stop has an adjusted timestamp matching the step that reached the target accuracy (run stop timestamp 1789340105384 is smaller than block stop timestamp 1789340490400). This is for scoring purposes.

Evaluation is offline here, in the sense that it doesn't happen during training but only afterwards. For convenience we have both in the same launcher, but you could separate them even more. But all events from the same run (training + the offline validation of checkpoints) should end up in the same log file.

Does this eval add back in post-processing step to the original logging? Thank you very much!

In our case both parts log to the same log file, but you can add the entries back in post-processing I think.

@ShriyaRishab
ShriyaRishab merged commit 7bf4fbd into mlcommons:master Sep 28, 2026
4 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants