[MLPerf 6.1][GRPO] Offline evaluation for Qwen3.5-397B GRPO - #906
Conversation
Signed-off-by: Jeremi Piotrowski <jpiotrowski@nvidia.com>
Advance the NeMo-RL pin to f987c0596 (mlperf-training-qwen35-next-v2), which adds deferred (offline) checkpoint evaluation to the reference launcher.
The rules require offline evaluation for qwen35_397b_grpo: checkpoints every step from the first-evaluation step through the end of training, evaluated post-run until one reaches the target; run_stop is backdated to the passing checkpoint's weight-update timestamp. Describe how the pinned reference implements this, and correct the submodule path in the checkout instructions.
|
MLCommons CLA bot All contributors have signed the MLCommons CLA ✍️ ✅ |
|
260913135524191054819_1_mllog.log here's a reference run mllog output with offline validation |
I am still a little bit confused about offline eval rule. I see eval before run_stop. I thought we moved it totally to be offline, and we just need to store the checkpoint before it ends. Could you help me understand more? Does this eval add back in post-processing step to the original logging? Thank you very much! |
The reference implementation runs one python process for training which emits Evaluation is offline here, in the sense that it doesn't happen during training but only afterwards. For convenience we have both in the same launcher, but you could separate them even more. But all events from the same run (training + the offline validation of checkpoints) should end up in the same log file.
In our case both parts log to the same log file, but you can add the entries back in post-processing I think. |
Summary
Stacked on #904: this PR contains that PR's commit plus the two offline-evaluation commits at the tip; the diff reduces to those two once #904 merges.
The benchmark rules (mlcommons/training_policies#596) require offline evaluation for
qwen35_397b_grpo: training saves a checkpoint after every step from the first-evaluation stepCEIL(2.5 + 3840 / global_batch_size)until it stops, each checkpoint records the timestamp of its latest weight update, and the saved checkpoints are evaluated post-run until one reaches the target.run_stopis emitted with the passing checkpoint's weight-update timestamp, so checkpoint writing and evaluation time are excluded from the measured score.This PR advances the NeMo-RL pin to
f987c0596(NVIDIA-NeMo/RL), which implements the requirement in the reference launcher, and updates the documentation accordingly.Implementation (in the pinned NeMo-RL commit)
DEFERRED_OFFLINE_EVAL=1disables inline validation and saves weights-only checkpoints every step fromgrpo.val_start_atthrough the final step; each checkpoint records its weight-update end time intraining_info.json.run_stopbackdated to the passing checkpoint's weight-update timestamp (abortedat the final checkpoint's timestamp when none crosses).RETAIN_GRPO_CHECKPOINTS=1.