Skip to content

Stop serving stack pods tolerating every taint - #456

Merged
haarchri merged 1 commit into
modelplaneai:mainfrom
haarchri:fix/serving-stack-tolerations-aicr
Sep 29, 2026
Merged

haarchri merged 1 commit into
modelplaneai:mainfrom
haarchri:fix/serving-stack-tolerations-aicr

Conversation

@haarchri

@haarchri haarchri commented Sep 15, 2026 •

Copy link
Copy Markdown
Collaborator

Description of your changes

aicr's bundler stamps a wildcard toleration {operator: Exists} on every pod it renders, Deployments included. On the generated clouds that put cert-manager, the Prometheus stack and the rest on tainted GPU nodes, where they squat on accelerated capacity and their eviction delays autoscaler scale-down.

The generator now rewrites every wildcard the bundle ships: node agents that must reach GPU nodes tolerate exactly the GPU taint the cluster compositions apply, everything else tolerates nothing and schedules on the untainted system pool. Like ALLOW/DROP the tables fail closed, so an aicr bump can't put a new control-plane pod back on GPU nodes. A new unit test pins every cloud, chart values and manifests alike, to keyed tolerations.

The hand-written clouds never stated a wildcard, but one shipped at runtime anyway: the node-exporter chart's default toleration is keyless, and Nebius, Vultr and Existing didn't override it. The unit test can only see stated values, so catching this class needs a rendered cluster. Those three halves now scope node-exporter to the GPU taint, mirroring the generated clouds.

Fixes #

I have:

  • Read and followed Modelplane's contribution process.
  • Run nix flake check (or ./nix.sh flake check) and made sure it passes.
  • Added or updated tests covering any composition function changes.
  • Signed off every commit with git commit -s.

@dennis-upbound dennis-upbound left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The unanchored patterns are the one fail-open path here.

A renamed upstream workload stops matching the allowlist and the audit goes red, which is the safe direction. But k8s-ephemeral-storage-metrics and kubelet-plugin match anywhere in a pod name, so a future workload that happens to contain either substring gets a wildcard toleration allowed silently — the exact regression this PR exists to catch.

Anchor the ones that can be? The comment says they're unanchored because the generated and hand-written halves name the same workload differently, so if that's only true for some of them, the rest could take ^.

@negz negz left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The fix LGTM but I'm a bit hesitant to introduce the new E2E test coverage - at least in the same PR.

Comment thread e2e/clouds/chainsaw/chainsaw-test.yaml Outdated
…olerating every taint

Signed-off-by: Christopher Haar <christopher.haar@upbound.io>
@haarchri
haarchri force-pushed the fix/serving-stack-tolerations-aicr branch from 928ad6d to 9574ef0 Compare September 29, 2026 18:26
@haarchri
haarchri merged commit 662c2b9 into modelplaneai:main Sep 29, 2026
6 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

Status: In Review

Development

Successfully merging this pull request may close these issues.

3 participants