Repository navigation
Conversation
Hetzner answers 423 locked while an action is still running on the balancer, and the operator skipped the target on that answer as if it were permanently invalid. A node added right after another change could stay out of the balancer until the next reconcile. Tell temporary rejections (locked, conflict, robot_unavailable, 5xx) from permanent ones and retry the former up to twice, after 1s and 2s. A target that still fails is skipped as before, and a 429 is still returned at once for the rate limit gate. Once one target has used up its retries, the rest of the run does not retry: the lock is on the whole balancer, and waiting per target would stall the service for minutes on a large balancer. Assisted-by: LLM Signed-off-by: Aleksei Sviridkin <f@lex.la>
s3rius
left a comment
There was a problem hiding this comment.
I'm not sure if we even need retries, because kubernetes handles reconcilidation itself.
So in any case the retry will happen on kubernetes side. Maybe we don't need this PR, it adds functionality that already exists but on a platform level.
| }; | ||
|
|
||
| /// Retries after the first attempt, sleeping 1 unit, 2 units, ... between them. | ||
| const ADD_TARGET_RETRIES: u32 = 2; |
There was a problem hiding this comment.
We can add exponential backoff.
Basically increase number of retries, and on each retry increase delay by factor of 2.
Something like, 6 retries, would result in retries in
1, 2, 4, 8, 16, 32 seconds.
This way we can have stronger guarantees that the request will succeeds.
There was a problem hiding this comment.
@s3rius I'd like to keep this as it is for now. A requeue does retry, but it runs the whole reconcile again, and here only the call that hit the lock is repeated. A longer backoff inside one reconcile would hold the service for more than a minute.
Every retry also costs API budget, and robotlb can't see how much is left: the generated client drops Hetzner's rate limit headers. I opened HenningHolmDE/hcloud-rust#44 to expose them. Once it's merged I'll rework the retry around the real budget.
A target added while another action is still running on the balancer gets skipped as if the IP were invalid (#54). Hetzner answers 423
lockedin that window.The add-target loop now retries a temporary rejection (
locked,conflict,robot_unavailableor a 5xx) twice, after 1s and 2s. A permanent one, like an IP outside the vSwitch subnet, is skipped as before. A 429 is still returned at once for the rate limit gate. After one target has used up its retries, the other targets in that run are not retried, because the lock is on the whole balancer.I did not add waiting for the running action to finish. That costs extra API calls per target, and the retry plus the existing 30 second requeue covers the short window. Other balancer calls (removing targets, changing services) still fail on
lockedas before.tokio::timerelies on the tokiotimefeature; #52 declares it explicitly.Stacked on #49.
Closes #54