Skip to content

Allow scaling a ModelDeployment to zero replicas - #454

Merged
haarchri merged 2 commits into
modelplaneai:mainfrom
haarchri:fix/deployment-scale-to-zero
Sep 29, 2026
Merged

haarchri merged 2 commits into
modelplaneai:mainfrom
haarchri:fix/deployment-scale-to-zero

Conversation

@haarchri

Copy link
Copy Markdown
Collaborator

Description of your changes

Allow scaling a ModelDeployment to zero replicas

spec.replicas carried a minimum of 1, so kubectl scale --replicas=0 was rejected at admission. The docs name KEDA and the scale subresource in the same paragraph, and scale-to-zero is a main reason to put KEDA in front of an expensive GPU workload, so the floor broke an expectation the API itself set. There was also no way to park a deployment: withdrawing its endpoints while keeping the object in place meant tainting the hosting cluster.

Dropping the floor alone would misreport the parked state. With zero desired, an empty schedule reads as InsufficientCapacity 0 of 0 replicas scheduled, and a control plane with no clusters reads as
NoClusters, both leaving the XR not ready, so a deliberately parked deployment would look permanently broken.

Zero desired now takes a dedicated path before input resolution: nothing is composed so the existing ModelReplicas and ModelEndpoints are pruned, status.replicas reports 0 for the scale subresource, and ReplicasScheduled and ReplicasReady read True with a new ScaledToZero reason, the way a Deployment at zero replicas reports Available. The docs' Scaling section now states the 0 to 10 range and the parked semantics.

Fixes #438

I have:

  • Read and followed Modelplane's contribution process.
  • Run nix flake check (or ./nix.sh flake check) and made sure it passes.
  • Added or updated tests covering any composition function changes.
  • Signed off every commit with git commit -s.

@dennis-upbound dennis-upbound left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Does parking actually release the GPUs?

The getting-started cast says minNodeCount: 1 # keep >=1; the autoscaler can't scale a GPU pool up from 0 for DRA pods. So on a pool sized that way, parking removes the pods but the floor keeps the nodes, and you keep paying for them. Set minNodeCount: 0 and the pool can empty, but then nothing brings it back for DRA pods — parking becomes one-way.

Savings are real in between (8 nodes down to 1), so this isn't an objection to the feature. But the docs say "a KEDA scaler can idle an expensive deployment and later restore it", and restore only holds while the floor is 1 or more. Worth saying which of the two it is.

@negz negz left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Few nits but LGTM overall.

Comment thread apis/modeldeployments/definition.yaml Outdated
Comment thread docs/content/models/model-deployment.md Outdated
Comment thread functions/compose-model-deployment/function/fn.py Outdated
@negz

negz commented Sep 28, 2026

Copy link
Copy Markdown
Collaborator

Does parking actually release the GPUs?

It releases them from this MD, making them available to others. It doesn't release them from the cluster. I think that's worthwhile.

Worth saying which of the two it is.

Not following what needs to be clarified.

Signed-off-by: Christopher Haar <christopher.haar@upbound.io>
Signed-off-by: Christopher Haar <christopher.haar@upbound.io>
@haarchri
haarchri force-pushed the fix/deployment-scale-to-zero branch from 7af3ec4 to ad4789e Compare September 29, 2026 18:09
@haarchri
haarchri merged commit d4f6226 into modelplaneai:main Sep 29, 2026
6 checks passed
@haarchri
haarchri deleted the fix/deployment-scale-to-zero branch September 29, 2026 18:34
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

Status: In Review

Development

Successfully merging this pull request may close these issues.

spec.replicas is bounded (minimum 1, maximum 10) but neither bound is documented

3 participants