Skip to content

[MLPerf RL] Implement RCP logging and deferred offline evaluation for DeepSWE - #2425

Merged
tianshub merged 1 commit into
atwigg/mlperffrom
lewu/rcp-logging-eval
Sep 29, 2026
Merged

tianshub merged 1 commit into
atwigg/mlperffrom
lewu/rcp-logging-eval

Conversation

@uwlei

@uwlei uwlei commented Sep 24, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Implements MLPerf v6.1.0 (qwen35_397b_grpo) Reference Convergence Point (RCP) logging and deferred offline evaluation for the distributed DeepSWE training and evaluation pipeline (b/565792423, per rco_logging_eval_plan.md):

  1. Training Checkpoint Manifest & block_stop (mlperf_35b_128_v5p.sh, rl_program.py, trainer_worker.py, run_deepswe_dist.py, maxtext_utils.py, mllog_utils.py):

    • Computes val_start_at = ceil(2.5 + 3840 / global_batch_size) (18 at gbs=256, with optional --val_start_at / VAL_START_AT override) and gates checkpoint saving on optimizer_step >= val_start_step.
    • Captures timestamp_ms = time.time_ns() // 1_000_000 immediately after the policy weight update (apply_optimizer=True) and before save_checkpoint() so checkpoint serialization and offline evaluation time are excluded from time-to-train.
    • Saves native scanned MaxText checkpoints during post-training (avoiding post-training unscan overhead; scanned-to-unscanned conversion is performed during offline evaluation via google/tunix#2504).
    • TrainerWorker.save_checkpoint returns the resolved checkpoint_path in Response.metadata; rl_program._maybe_save_checkpoint passes it to on_checkpoint_saved, which upserts {step, checkpoint_path, timestamp_ms, samples_count, mllog_file, ...} into ${METRIC_LOGGER_DIR}/eval_checkpoints.jsonl.
    • mllog_utils.train_stop() emits only block_stop backdated to last_step_timestamp_ms and no longer emits run_stop.
  2. Deferred Offline Evaluation & Backdated run_stop (eval_deepswe.py, mllog_utils.py, mlperf_base.sh, mlperf_35b_eval.sh, k8s_launcher.sh):

    • When CHECKPOINT_MANIFEST_FILE is set, mlperf_base.sh validates manifest step contiguity and evaluates checkpoints sequentially in step order on 64 TPU v5p chips (16 workers x 2x2x1), using pass@4 (pass_at_k["4"]) as eval_accuracy.
    • mllog_utils.configure_logger downloads the existing training MLLOG (seed_<seed>.out) from GCS before appending eval_start, tracked_stats (validation_time), eval_accuracy, and eval_stop for each evaluated checkpoint. Fails fast if --rcp_logging is enabled without mlperf_logging installed.
    • Stops at the first checkpoint reaching eval_accuracy >= target_accuracy (0.69) and emits run_stop(status="success", samples_count=..., time_ms=checkpoint_timestamp_ms) backdated to that checkpoint's weight-update timestamp; otherwise emits run_stop(status="aborted") on the final checkpoint.
    • Without CHECKPOINT_MANIFEST_FILE, mlperf_35b_eval.sh runs standalone single-checkpoint evaluation (and supports mock RCP verification when RCP_LOGGING=true, writing to ${EVAL_OUTPUT_DIR}/mllog).

Notes vs. plan

  • log_offline_eval_step does not emit a duplicate train_samples event because init_print already logs train_samples (EXACTLY_ONE in MLPerf v6.1.0 compliance rules).
  • Manifest rows record mllog_file so offline evaluation appends to the exact training MLLOG file.
  • The eval container image (sanbao/tunix_stack:eval) needs mlperf_logging installed; eval_samples=251 matches the plan and upstream mlcommons/logging master (whereas the 6.1.0-rc1 tag still checks 256).

Verification

1. Unit Tests (CPU)

  • maxtext_utils_test (45/45), rl_program_test (139/139), deepswe_mllog_utils_test (18/18), eval_deepswe_test (12/12), grpo_recipe_wiring_test (41/41), distributed_rl_engine_test (73/73), run_trainer_node_test (53/53), trainer_worker_test (19/19).
  • Added unit tests covering:
    • TrainerWorker._resolve_checkpoint_path resolving <checkpoint_dir>/<step>/model_params from trainer.checkpoint_dir and rl_program passing checkpoint_path from save_checkpoint Response.metadata to on_checkpoint_saved.
    • compute_val_start_step, manifest serialization/contiguity validation, multi-checkpoint offline evaluation stopping at the first passing checkpoint (status="success"), non-converging runs (status="aborted"), and fail-fast when --rcp_logging is set without mlperf_logging.

2. Live GKE Cluster E2E Verification (bodaborg-v5p-nap, TPU v5p)

  1. Training with Scanned Base Checkpoint -> Unscanned Save Hook (128 v5p chips):

    • Ran mlperf_35b_128_v5p.sh with MAX_STEPS=2 VAL_START_AT=1 DEFERRED_OFFLINE_EVAL=1 RCP_LOGGING=true starting from the scanned base checkpoint gs://sanbao-europe/qwen3.5-35B-A3B-scanned/base/0/items.
    • install_eval_checkpoint_unscan_hook saved unscanned bfloat16 model_params checkpoints at steps 1 and 2 (~16.7 s per save), wrote both entries to eval_checkpoints.jsonl, and terminated training MLLOG at block_stop backdated to step 2's weight update (1790653536423) with no run_stop.
  2. Sequential Offline Evaluation on Hook-Produced Checkpoints (64 v5p chips, TASKS_LIMIT=64):

    • Executed mlperf_35b_eval.sh against the manifest (16 workers x 2x2x1, mesh_fsdp=2, mesh_tp=2, NUM_GENERATIONS=4, TEMPERATURE=0.1, TOP_P=0.95, TASKS_LIMIT=64 -> 256 attempts/step).
    • (Note: the training run predated the Response.metadata["checkpoint_path"] fix, so the manifest paths were normalized to absolute gs:// URIs; the eval pods loaded this branch + mlperf_logging via bootstrap overlay because sanbao/tunix_stack:eval does not yet bundle mlperf_logging.)

Manifest (gs://sanbao-europe/mlperf/qwen35_35b/trellis/lewu_rcp_eval/hook_e2e_v2/mllog/eval_checkpoints_abs.jsonl):

{"step": 1, "checkpoint_path": "gs://sanbao-europe/mlperf/qwen35_35b/trellis/lewu_rcp_eval/hook_e2e_v2/lewu-rcp-train/checkpoints/1/model_params", "timestamp_ms": 1790653351142, "samples_count": 256, "global_batch_size": 256, "batch_size": 16, "num_generations": 16, "val_start_at": 1, "max_steps": 2, "target_accuracy": 0.69, "seed": 42, "mllog_file": "gs://sanbao-europe/mlperf/qwen35_35b/trellis/lewu_rcp_eval/hook_e2e_v2/mllog/seed_42.out"}
{"step": 2, "checkpoint_path": "gs://sanbao-europe/mlperf/qwen35_35b/trellis/lewu_rcp_eval/hook_e2e_v2/lewu-rcp-train/checkpoints/2/model_params", "timestamp_ms": 1790653536423, "samples_count": 512, "global_batch_size": 256, "batch_size": 16, "num_generations": 16, "val_start_at": 1, "max_steps": 2, "target_accuracy": 0.69, "seed": 42, "mllog_file": "gs://sanbao-europe/mlperf/qwen35_35b/trellis/lewu_rcp_eval/hook_e2e_v2/mllog/seed_42.out"}

Orbax Unscanned Checkpoint Restore Across All 16 Eval Workers (scan_layers=False):

  • Step 1: all 16 workers (lewu-rcp-eval-0..15) restored checkpoints/1/model_params in 33.0–38.7 s with zero shape or sharding errors.
  • Step 2: all 16 workers (lewu-rcp-eval-0..15) restored checkpoints/2/model_params in 31.8–51.9 s with zero shape or sharding errors.
2026-09-29 06:10:30,649 - absl - INFO - [process=0] [sync] Finished load in 36.16 seconds @ gs://sanbao-europe/mlperf/qwen35_35b/trellis/lewu_rcp_eval/hook_e2e_v2/lewu-rcp-train/checkpoints/1/model_params
2026-09-29 06:53:20,738 - absl - INFO - [process=0] [sync] Finished load in 50.64 seconds @ gs://sanbao-europe/mlperf/qwen35_35b/trellis/lewu_rcp_eval/hook_e2e_v2/lewu-rcp-train/checkpoints/2/model_params

Step 1 GCS summary.json (.../eval_results_v3/step_1/20260929T060823Z-f604fb47/summary.json):

{
  "instances": 64,
  "attempts_per_instance": 4,
  "expected_attempts": 256,
  "completed_attempts": 256,
  "missing_attempts": 0,
  "error_attempts": 7,
  "resolved_attempts": 69,
  "avg_at_k": 0.26953125,
  "pass_at_k": {"1": 0.26953125, "4": 0.609375},
  "mean_reward": 0.26953125,
  "status_counts": {"SUCCEEDED": 176, "MAX_CONTEXT_LIMIT_REACHED": 73, "ERROR": 7},
  "complete": true,
  "fatal_error": null,
  "rcp_logged": true,
  "target_accuracy": 0.69,
  "target_reached": false,
  "checkpoint_step": 1,
  "checkpoint_timestamp_ms": 1790653351142,
  "samples_count": 256
}

Step 2 GCS summary.json (.../eval_results_v3/step_2/20260929T065059Z-01527e40/summary.json):

{
  "instances": 64,
  "attempts_per_instance": 4,
  "expected_attempts": 256,
  "completed_attempts": 256,
  "missing_attempts": 0,
  "error_attempts": 0,
  "resolved_attempts": 91,
  "avg_at_k": 0.35546875,
  "pass_at_k": {"1": 0.35546875, "4": 0.625},
  "mean_reward": 0.35546875,
  "status_counts": {"SUCCEEDED": 181, "MAX_CONTEXT_LIMIT_REACHED": 75},
  "complete": true,
  "fatal_error": null,
  "rcp_logged": true,
  "target_accuracy": 0.69,
  "target_reached": false,
  "checkpoint_step": 2,
  "checkpoint_timestamp_ms": 1790653536423,
  "samples_count": 512
}

Appended MLLOG Events (gs://sanbao-europe/mlperf/qwen35_35b/trellis/lewu_rcp_eval/hook_e2e_v2/mllog/seed_42.out):

:::MLLOG {"namespace": "", "time_ms": 1790653536423, "event_type": "INTERVAL_END", "key": "block_stop", "value": null, "metadata": {"file": "/app/tunix/utils/mllog_utils.py", "lineno": 274, "step": 2, "samples_count": 512}}
:::MLLOG {"namespace": "", "time_ms": 1790662429020, "event_type": "INTERVAL_START", "key": "eval_start", "value": null, "metadata": {"file": "/app/tunix/utils/mllog_utils.py", "lineno": 328, "step": 1, "samples_count": 256}}
:::MLLOG {"namespace": "", "time_ms": 1790664126781, "event_type": "POINT_IN_TIME", "key": "tracked_stats", "value": {"validation_time": 1697.759601957}, "metadata": {"file": "/app/tunix/utils/mllog_utils.py", "lineno": 350, "step": 1}}
:::MLLOG {"namespace": "", "time_ms": 1790664126781, "event_type": "POINT_IN_TIME", "key": "eval_accuracy", "value": 0.609375, "metadata": {"file": "/app/tunix/utils/mllog_utils.py", "lineno": 361, "samples_count": 256}}
:::MLLOG {"namespace": "", "time_ms": 1790664126782, "event_type": "INTERVAL_END", "key": "eval_stop", "value": null, "metadata": {"file": "/app/tunix/utils/mllog_utils.py", "lineno": 367, "step": 1, "samples_count": 256}}
:::MLLOG {"namespace": "", "time_ms": 1790665045803, "event_type": "INTERVAL_START", "key": "eval_start", "value": null, "metadata": {"file": "/app/tunix/utils/mllog_utils.py", "lineno": 328, "step": 2, "samples_count": 512}}
:::MLLOG {"namespace": "", "time_ms": 1790666751105, "event_type": "POINT_IN_TIME", "key": "tracked_stats", "value": {"validation_time": 1705.300022161}, "metadata": {"file": "/app/tunix/utils/mllog_utils.py", "lineno": 350, "step": 2}}
:::MLLOG {"namespace": "", "time_ms": 1790666751105, "event_type": "POINT_IN_TIME", "key": "eval_accuracy", "value": 0.625, "metadata": {"file": "/app/tunix/utils/mllog_utils.py", "lineno": 361, "samples_count": 512}}
:::MLLOG {"namespace": "", "time_ms": 1790666751105, "event_type": "INTERVAL_END", "key": "eval_stop", "value": null, "metadata": {"file": "/app/tunix/utils/mllog_utils.py", "lineno": 367, "step": 2, "samples_count": 512}}
:::MLLOG {"namespace": "", "time_ms": 1790653536423, "event_type": "INTERVAL_END", "key": "run_stop", "value": null, "metadata": {"file": "/app/tunix/utils/mllog_utils.py", "lineno": 1034, "status": "aborted", "samples_count": 512}}

3. MLPerf v6.1.0 Checker Results

  • On the 2-step TPU test MLLOG (VAL_START_AT=1, MAX_STEPS=2): compliance_checker validates all structural/disclosure/interval checks and flags only the two expected test-config conditions (samples_count 256/512 vs. default validation_start_samples=4608 because VAL_START_AT=1, and eval_accuracy=0.625 <= 0.69 because 2 steps do not converge).
  • Full recipe cadence (val_start_at=18, samples_count=4608..5120):
    • Converging run (step 19 reaches pass@4 >= 0.69, backdated run_stop(status="success")):

Compliance Checker (python3 -m mlperf_logging.compliance_checker --usage training --ruleset 6.1.0):

INFO - Running compliance on file: .../seed_42.out
INFO -  Compliance checks: training_6.1.0/common.yaml
INFO -  Compliance checks: training_6.1.0/closed_common.yaml
INFO -  Compliance checks: training_6.1.0/closed_qwen35_397b_grpo.yaml
INFO - SUCCESS

RCP Checker (python3 -m mlperf_logging.rcp_checker --rcp_usage training --rcp_version 6.1.0 --verbose):

INFO -  Running RCP Checker, pass: pruned_rcps
INFO -  RCP Record: {'Benchmark': 'qwen35_397b_grpo', 'BS': 256, 'Hyperparams': {'opt_base_learning_rate': 1e-06, 'opt_gradient_clip_norm': 0.125, 'num_prompts_per_step': 16, 'num_generations_per_prompt': 16, 'gradient_accumulation_steps': 32}, 'Epochs to converge': [4608, 4608, 4608, 4608, 4608, 4608], 'RCP Mean': np.float64(4608.0), 'RCP Stdev': np.float64(0.0), 'Max Speedup': np.float64(1.0), 'Min Epochs': np.float64(4608.0)}
INFO -  Submission mean epochs: 4608.0000
INFO - RCP found, RCP test PASSED

References

Checklist

  • I have added all the necessary unit tests for my change.
  • I have verified that my change does not break existing code and all unit tests pass.
  • I have added all appropriate doc-strings/documentation.
  • My PR is based on the latest changes of the target branch (atwigg/mlperf).
  • I have signed the Contributor License Agreement.
  • I have followed Contribution Guidelines.

@google-cla

google-cla Bot commented Sep 24, 2026

Copy link
Copy Markdown

Thanks for your pull request! It looks like this may be your first contribution to a Google open source project. Before we can look at your pull request, you'll need to sign a Contributor License Agreement (CLA).

View this failed invocation of the CLA check for more information.

For the most up to date status, view the checks section at the bottom of the pull request.

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces support for the OpenHands/CodeAct agent scaffold, enabling multi-turn tool use, IPython cell execution, and robust sandbox stashing/restoring of R2E grading tests. It also adds cluster resource reaping, GCS weight-sync verification, and extensive configuration plumbing for MaxText training. The code review identified critical issues: the pinned JAX and Flax versions in the Dockerfile do not exist on PyPI and will fail the build, and a hardcoded topology ID in the YAML generator will cause Kueue scheduling failures on smaller TPU slices. Additionally, minor PEP 8 line-length violations and redundant import paths were flagged for cleanup.

Comment thread Dockerfile
Comment thread Dockerfile
Comment thread tunix/experimental/distributed/deployment/yaml_generator.py
Comment thread examples/deepswe/openhands_utils.py
Comment thread examples/deepswe/template.py
Comment thread examples/deepswe/train_deepswe_nb.py
@uwlei
uwlei changed the base branch from main to atwigg/mlperf September 24, 2026 21:07
Comment thread tunix/experimental/examples/deepswe_dist/eval_deepswe.py Outdated
Comment thread tunix/experimental/examples/deepswe_dist/run_deepswe_dist.py Outdated
Comment thread tunix/experimental/examples/deepswe_dist/run_deepswe_dist.py Outdated
Comment thread tunix/experimental/examples/deepswe_dist/run_deepswe_dist.py Outdated
Comment thread tunix/experimental/examples/deepswe_dist/run_deepswe_dist.py Outdated
Comment thread tunix/experimental/orchestrator/rl_program.py Outdated
@uwlei
uwlei force-pushed the lewu/rcp-logging-eval branch 2 times, most recently from bede958 to 4a04b5d Compare September 25, 2026 20:21
@uwlei uwlei changed the title Lewu/rcp logging eval [MLPerf RL]RCP logging for eval pipeline Sep 25, 2026
@uwlei
uwlei marked this pull request as ready for review September 25, 2026 21:08
@uwlei uwlei changed the title [MLPerf RL]RCP logging for eval pipeline [MLPerf RL] Implement RCP logging and deferred offline evaluation for DeepSWE Sep 25, 2026
@uwlei
uwlei force-pushed the lewu/rcp-logging-eval branch 2 times, most recently from 3fa1c0a to 4d345c1 Compare September 25, 2026 22:28
@uwlei
uwlei requested a review from tianshub September 25, 2026 22:59
@uwlei
uwlei requested review from hgao327 and s-noghabi September 25, 2026 22:59
Comment thread tunix/experimental/examples/deepswe_dist/eval_deepswe.py Outdated
Comment thread tunix/experimental/examples/deepswe_dist/eval_deepswe.py Outdated
Comment thread tunix/utils/mllog_utils.py Outdated
@uwlei
uwlei force-pushed the lewu/rcp-logging-eval branch 3 times, most recently from aac4ed4 to 8c5db00 Compare September 29, 2026 08:01
Comment thread tunix/experimental/worker/trainer_worker.py Outdated
@tianshub

Copy link
Copy Markdown
Collaborator

/gemini review

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces MLPerf RCP (mllog) compliance logging and deferred offline evaluation capabilities, including saving unscanned bfloat16 checkpoints, managing checkpoint manifests, and backdating run stops. Feedback on these changes focuses on resolving a potential JSON parsing failure in mlperf_base.sh when handling multiple matching summary files, as well as addressing several style guide violations. Specifically, the code should be refactored to avoid speculative getattr and dict.get calls on structured schemas (violating the 'One well-lit path' and 'Fail loud' rules) and to catch specific exceptions rather than a broad Exception when importing optional packages.

Comment thread tunix/experimental/examples/recipes/mlperf_base.sh Outdated
Comment thread tunix/experimental/examples/deepswe_dist/eval_deepswe.py
Comment thread tunix/experimental/orchestrator/rl_program.py Outdated
Comment thread tunix/experimental/worker/trainer_worker.py Outdated
Comment thread tunix/utils/mllog_utils.py
Comment thread tunix/utils/mllog_utils.py Outdated
@uwlei
uwlei force-pushed the lewu/rcp-logging-eval branch from 8c5db00 to b6f40bc Compare September 29, 2026 18:16
Comment thread tunix/utils/maxtext_utils.py Outdated
getattr(constants, "PIPELINE_PARALLELISM", "pipeline_parallelism"): 1,
getattr(constants, "CONTEXT_PARALLELISM", "context_parallelism"): train_sp,
getattr(constants, "EXPERT_PARALLELISM", "expert_parallelism"): getattr(args, "train_mesh_expert", 1),
# Mandatory v6.1 precision and run-config disclosures.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

They are in training_6.1.0/common.yaml (lines 147–192), which mlperf_logging.compliance_checker --usage training --ruleset 6.1.0 runs on every training submission:

Local copy (used in our test runs):
common.yaml
Upstream mlcommons/logging (6.1.0-rc1): mlperf_logging/compliance_checker/training_6.1.0/common.yaml#L147-L192

@uwlei
uwlei force-pushed the lewu/rcp-logging-eval branch from b6f40bc to e4cc533 Compare September 29, 2026 20:05
@uwlei

uwlei commented Sep 29, 2026

Copy link
Copy Markdown
Collaborator Author

/gemini review

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces MLPerf RCP (mllog) compliance logging support for offline evaluation, including a new offline evaluation script (mlperf_35b_eval.sh), checkpoint manifest generation, and several helper utilities in mllog_utils.py. Feedback on the changes highlights a potential issue in mlperf_base.sh where mapfile can process an empty trailing row, and a performance concern in run_deepswe_dist.py where synchronous GCS/file I/O operations are called within the asyncio event loop. Additionally, several instances of speculative .get() and getattr() in run_deepswe_dist.py, maxtext_utils.py, and rl_program.py violate the repository's 'One well-lit path' style guide rule and should be replaced with direct attribute/key access or explicit checks.

Comment thread tunix/experimental/examples/recipes/mlperf_base.sh
Comment thread tunix/experimental/examples/deepswe_dist/run_deepswe_dist.py Outdated
Comment thread tunix/utils/maxtext_utils.py Outdated
Comment thread tunix/experimental/orchestrator/rl_program.py Outdated
Comment thread tunix/experimental/examples/deepswe_dist/run_deepswe_dist.py
@uwlei
uwlei force-pushed the lewu/rcp-logging-eval branch from e4cc533 to 6336d4c Compare September 29, 2026 20:55
@tianshub
tianshub merged commit 37aa73e into atwigg/mlperf Sep 29, 2026
2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants