Skip to content

Cluster manager: consult the pod's log before reaping a worker with a stale heartbeat - #82

Open
vohonen wants to merge 2 commits into
longtermrisk:mainfrom
vohonen:unresponsive-worker-progress-check
Open

vohonen wants to merge 2 commits into
longtermrisk:mainfrom
vohonen:unresponsive-worker-progress-check

Conversation

@vohonen

@vohonen vohonen commented Sep 15, 2026

Copy link
Copy Markdown
Contributor

Problem

The cluster manager reaps a worker once its ping is older than UNRESPONSIVE_THRESHOLD (120 s), terminating the pod and requeueing the job from scratch. That heartbeat is written by one thread inside the worker process, and it stalls while the job is healthy.

Four pods of the "Vili CLR" org, RL training jobs on 2x H200:

run last ping (UTC) worker marked terminated job state in the log at termination
worker …-da5fed94, 2026-09-10 07:42:00 07:44:04 training step 96, mid "Evaluating responses"
worker …-57d84cad, 2026-09-10 13:43:23 13:43:25 step 161 of 200
worker …-6877f752, 2026-09-14 10:45:10 10:47:10 step 142
worker …-d6142678, 2026-09-15 09:07:27 09:09:32 step 72

In each case fetch_and_save_worker_logs fetched the log live from the pod's :10101 endpoint at the moment of termination and succeeded, so the pod was up and reachable when it was killed, and the saved log shows the job progressing normally with no error, then silence at the last ping. Eight more workers of the same org were reaped within 35 minutes on 2026-09-03, which looks like one heartbeat-side incident hitting every worker at once. Across 91 runs that is 14 healthy pods killed, each costing a 2-3 h restart from step zero.

I could not see why the worker's heartbeat thread goes quiet: its logging.error output is not in the pod log. If the manager logs for the timestamps above are retrievable, "hasn't pinged for N seconds" lines and anything around them would help.

Change

A stale ping no longer terminates a pod by itself. When a worker's ping is past the threshold, the manager reads the pod's /logs endpoint (only then, so the normal path is unchanged) and:

  • leaves the worker alone, with a warning, if its log grew within PROGRESS_GRACE (default 600 s);
  • reaps it as before if the log is flat for longer than that, unreachable, or the worker has no pod;
  • reaps it regardless once the ping is older than UNRESPONSIVE_HARD_LIMIT (default 3600 s), so a broken heartbeat cannot hold a pod indefinitely (the pod-side TTL remains the backstop beyond that).

All three thresholds are environment-configurable (OW_UNRESPONSIVE_THRESHOLD, OW_PROGRESS_GRACE, OW_UNRESPONSIVE_HARD_LIMIT). The decision rule is a pure function in openweights/cluster/liveness.py with unit tests in tests/test_worker_liveness.py; org_manager.py keeps a per-worker (log length, time it last grew) only for workers whose ping has gone stale.

Cost of the change for a genuinely dead worker: it is reaped after PROGRESS_GRACE plus one poll instead of after 120 s, about ten minutes of idle pod.

Test

tests/test_worker_liveness.py covers fresh ping, stale ping with unreadable log, stale ping with growing log, stale ping with flat log, and the hard limit. org_manager imports cleanly with the new constants.

🤖 Generated with Claude Code

Vili Kohonen and others added 2 commits September 15, 2026 13:37
… stale heartbeat

The manager terminated a worker once its ping was older than UNRESPONSIVE_THRESHOLD
(120 s). That heartbeat is one thread inside the worker process and it stalls while the
job is healthy: on 2026-09-10, -14 and -15 four pods were killed mid-step with their
training logs still advancing, each exactly two minutes after the last ping, and the
manager fetched those logs live from the pod at the moment it terminated it. Eight more
went the same way within 35 minutes on 2026-09-03.

Now a stale ping makes the manager read the pod's log endpoint (only then, so the normal
path is unchanged). A worker whose log grew within PROGRESS_GRACE (600 s) is left alone
with a warning; one whose log is flat, unreachable or absent is reaped as before; and past
UNRESPONSIVE_HARD_LIMIT (3600 s) the worker is reaped regardless, so a broken heartbeat
cannot hold a pod. All three are environment-configurable (OW_UNRESPONSIVE_THRESHOLD,
OW_PROGRESS_GRACE, OW_UNRESPONSIVE_HARD_LIMIT). The decision rule is a pure function in
cluster/liveness.py with unit tests.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…to diagnose a reaped run

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant