Conversation
… stale heartbeat The manager terminated a worker once its ping was older than UNRESPONSIVE_THRESHOLD (120 s). That heartbeat is one thread inside the worker process and it stalls while the job is healthy: on 2026-09-10, -14 and -15 four pods were killed mid-step with their training logs still advancing, each exactly two minutes after the last ping, and the manager fetched those logs live from the pod at the moment it terminated it. Eight more went the same way within 35 minutes on 2026-09-03. Now a stale ping makes the manager read the pod's log endpoint (only then, so the normal path is unchanged). A worker whose log grew within PROGRESS_GRACE (600 s) is left alone with a warning; one whose log is flat, unreachable or absent is reaped as before; and past UNRESPONSIVE_HARD_LIMIT (3600 s) the worker is reaped regardless, so a broken heartbeat cannot hold a pod. All three are environment-configurable (OW_UNRESPONSIVE_THRESHOLD, OW_PROGRESS_GRACE, OW_UNRESPONSIVE_HARD_LIMIT). The decision rule is a pure function in cluster/liveness.py with unit tests. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…to diagnose a reaped run Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
The cluster manager reaps a worker once its
pingis older thanUNRESPONSIVE_THRESHOLD(120 s), terminating the pod and requeueing the job from scratch. That heartbeat is written by one thread inside the worker process, and it stalls while the job is healthy.Four pods of the "Vili CLR" org, RL training jobs on 2x H200:
…-da5fed94, 2026-09-10…-57d84cad, 2026-09-10…-6877f752, 2026-09-14…-d6142678, 2026-09-15In each case
fetch_and_save_worker_logsfetched the log live from the pod's :10101 endpoint at the moment of termination and succeeded, so the pod was up and reachable when it was killed, and the saved log shows the job progressing normally with no error, then silence at the last ping. Eight more workers of the same org were reaped within 35 minutes on 2026-09-03, which looks like one heartbeat-side incident hitting every worker at once. Across 91 runs that is 14 healthy pods killed, each costing a 2-3 h restart from step zero.I could not see why the worker's heartbeat thread goes quiet: its
logging.erroroutput is not in the pod log. If the manager logs for the timestamps above are retrievable, "hasn't pinged for N seconds" lines and anything around them would help.Change
A stale ping no longer terminates a pod by itself. When a worker's ping is past the threshold, the manager reads the pod's
/logsendpoint (only then, so the normal path is unchanged) and:PROGRESS_GRACE(default 600 s);UNRESPONSIVE_HARD_LIMIT(default 3600 s), so a broken heartbeat cannot hold a pod indefinitely (the pod-side TTL remains the backstop beyond that).All three thresholds are environment-configurable (
OW_UNRESPONSIVE_THRESHOLD,OW_PROGRESS_GRACE,OW_UNRESPONSIVE_HARD_LIMIT). The decision rule is a pure function inopenweights/cluster/liveness.pywith unit tests intests/test_worker_liveness.py;org_manager.pykeeps a per-worker(log length, time it last grew)only for workers whose ping has gone stale.Cost of the change for a genuinely dead worker: it is reaped after
PROGRESS_GRACEplus one poll instead of after 120 s, about ten minutes of idle pod.Test
tests/test_worker_liveness.pycovers fresh ping, stale ping with unreadable log, stale ping with growing log, stale ping with flat log, and the hard limit.org_managerimports cleanly with the new constants.🤖 Generated with Claude Code