Skip to content

rstudio-pm: use /__cluster-health__ for the readiness probe - #959

Merged
jonyoder merged 1 commit into
mainfrom
ppm-readiness-cluster-health
Oct 8, 2026
Merged

jonyoder merged 1 commit into
mainfrom
ppm-readiness-cluster-health

Conversation

@jonyoder

@jonyoder jonyoder commented Oct 6, 2026

Copy link
Copy Markdown
Contributor

Change the rstudio-pm readiness probe default from /__ping__ to /__cluster-health__. After rstudio/package-manager#21124, /__ping__ returns 200 during startup and in offline mode, so it no longer shows readiness. Liveness and startup probes are unchanged.

/__cluster-health__ needs Package Manager 2026.06.0 or later (the chart default is 2026.09.0). On older images, set readinessProbe.httpGet.path back to /__ping__. Merge this with or before the Package Manager release that contains #21124.


Testing

  • ct lint on the changed charts passed (version bump detected, 0.20.5 to 0.20.6), and helm lint --strict passed on all lint/*.yaml values files.
  • helm unittest: 4/4 pass, including the new tests/probe_test.yaml.
  • README regenerated with helm-docs 1.13.1.
  • kind, chart defaults, no license, posit/package-manager:2026.09.0-ubuntu-24.04: the pod became Ready with 0 restarts and no probe failures. kubectl describe shows Readiness: http-get http://:4242/__cluster-health__ delay=3s timeout=1s period=3s.
  • Earlier, in a kind cluster on a build that includes #21124, running rspm offline returned __ping__=200 __cluster-health__=503, and the pod went NotReady until rspm online.

🤖 Generated with Claude Code

Package Manager's /__ping__ now answers during startup and offline (rstudio/package-manager#21124), so it no longer shows readiness.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@jonyoder
jonyoder requested review from a team, jmwoliver and jstruzik October 6, 2026 20:58
@jonyoder

jonyoder commented Oct 7, 2026

Copy link
Copy Markdown
Contributor Author

More testing: 2 replicas, a license, and a build of current package-manager main

I ran this chart at f59ef3a in kind with 2 replicas, Postgres in the cluster, a shared PVC, and liveness on /__ping__. The image was rstudio-pm and rspm built from package-manager main at 4e1be9f8e4, which includes all of rstudio/package-manager#21124. No published image had that yet.

Test Result
Both replicas start Pass. Both were Ready in about 10s, and /__cluster-health__ returned 200 healthy.
Delete one pod (4 times) Pass. Only that pod left the Service endpoints, within 1s. The other stayed Ready on every 1s sample, and the replacement was Ready about 4s after it was created.
Leader fails its cluster checks (HealthCheckTimeout=10s, StepDownTimeout=0s, a fake stale node row) Pass. The leader returned 503 degraded and went NotReady, with no restart. The follower stayed Ready. The leader was Ready again about 4s after the row was removed.
rspm offline on one pod, then rspm online Pass. /__ping__ returned 200 and /__cluster-health__ 503, and that pod went NotReady about 6s later while the other stayed Ready. It was Ready again about 4s after rspm online.
Licensed (dev license, Advanced tier) Pass. Both replicas Ready, no restarts.

Things worth knowing:

  • One replica (the default): the only pod is the leader. A Postgres outage longer than about 60s makes it NotReady, which leaves the Service with no endpoints until the database comes back.
  • Postgres outage, 2 replicas: only the leader went NotReady (after about 70s); the follower stayed Ready. The leader logged Leader stepping down but kept returning degraded until Postgres was back. That may be a PPM bug, and I haven't confirmed the cause.
  • Rolling restarts: each pod restart made the leader fail its cluster check for about 15 to 25s while the old pod's node row aged out. The 60s default covers this, but a HealthCheckTimeout under about 45s would make the leader NotReady on routine rollouts.
  • No draining 503 seen on delete: an idle node went from 200 straight to connection refused. Kubernetes drops the pod from the endpoints because it is terminating, not because of the probe.

🤖 Generated with Claude Code

@jonyoder
jonyoder requested a review from CDRayn October 7, 2026 20:35
@jonyoder
jonyoder merged commit 9a3a26d into main Oct 8, 2026
9 checks passed
@jonyoder
jonyoder deleted the ppm-readiness-cluster-health branch October 8, 2026 10:19
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants