Allow scaling a ModelDeployment to zero replicas - #454
Conversation
dennis-upbound
left a comment
There was a problem hiding this comment.
Does parking actually release the GPUs?
The getting-started cast says minNodeCount: 1 # keep >=1; the autoscaler can't scale a GPU pool up from 0 for DRA pods. So on a pool sized that way, parking removes the pods but the floor keeps the nodes, and you keep paying for them. Set minNodeCount: 0 and the pool can empty, but then nothing brings it back for DRA pods — parking becomes one-way.
Savings are real in between (8 nodes down to 1), so this isn't an objection to the feature. But the docs say "a KEDA scaler can idle an expensive deployment and later restore it", and restore only holds while the floor is 1 or more. Worth saying which of the two it is.
negz
left a comment
There was a problem hiding this comment.
Few nits but LGTM overall.
It releases them from this MD, making them available to others. It doesn't release them from the cluster. I think that's worthwhile.
Not following what needs to be clarified. |
Signed-off-by: Christopher Haar <christopher.haar@upbound.io>
Signed-off-by: Christopher Haar <christopher.haar@upbound.io>
7af3ec4 to
ad4789e
Compare
Description of your changes
Allow scaling a ModelDeployment to zero replicas
spec.replicascarried a minimum of 1, sokubectl scale --replicas=0was rejected at admission. The docs name KEDA and the scale subresource in the same paragraph, and scale-to-zero is a main reason to put KEDA in front of an expensive GPU workload, so the floor broke an expectation the API itself set. There was also no way to park a deployment: withdrawing its endpoints while keeping the object in place meant tainting the hosting cluster.Dropping the floor alone would misreport the parked state. With zero desired, an empty schedule reads as
InsufficientCapacity0 of 0 replicas scheduled, and a control plane with no clusters reads asNoClusters, both leaving the XR not ready, so a deliberately parked deployment would look permanently broken.
Zero desired now takes a dedicated path before input resolution: nothing is composed so the existing
ModelReplicasandModelEndpointsare pruned,status.replicasreports 0 for the scale subresource, andReplicasScheduledandReplicasReadyread True with a newScaledToZeroreason, the way a Deployment at zero replicas reports Available. The docs' Scaling section now states the 0 to 10 range and the parked semantics.Fixes #438
I have:
nix flake check(or./nix.sh flake check) and made sure it passes.git commit -s.