Conversation
added 2 commits
September 27, 2026 09:00
PickRoute remembers the region a route last served from (`lastRegion`, keyed by provider|region|model) so a balanced pool does not flip regions as rotation advances. The latch was only ever written, never re-evaluated, so a change in pool composition underneath a route left it pinned to whatever region it picked first. In production a migration added 78 cn WorkBuddy accounts to a pool that held 2 global ones. The route had latched onto global, so every request kept landing on those two accounts while the 78 new cn accounts received nothing: 285 requests averaged 48s TTFB (39% over 15s, worst 528s) against global, while cn sat idle. Restarting the process was the only way to clear it. Drop the latch once its region is left behind by a factor of two, so the route re-seats on the largest region. The hysteresis keeps a roughly balanced pool on its current region instead of thrashing between two comparable regions. TestPickRouteReseatsStaleRegionLatch fails on the previous code with exactly the production symptom (map[global:12]); TestPickRouteKeepsRegionLatchWhenBalanced pins the balanced case.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
Pool.PickRouteremembers the region a route last served from (lastRegion, keyed byprovider|region|model) so a balanced pool does not flip regions as rotation advances. The latch was only ever written, never re-evaluated, so a change in pool composition underneath a route left it pinned to whatever region it happened to pick first.In production a migration added 78
cnWorkBuddy accounts to a pool that held 2globalones. The route had latched ontoglobal, so every request kept landing on those two accounts while the 78 newcnaccounts received nothing:Restarting the process was the only way to clear it, and the same trap awaits anyone who adds or migrates accounts into a running instance.
Fix
Drop the latch once its region is left behind by a factor of two, so the route re-seats on the largest region:
The hysteresis keeps a roughly balanced pool on its current region instead of thrashing between two comparable regions, preserving the original intent for the common case.
Tests
TestPickRouteReseatsStaleRegionLatchfails on the previous code with exactly the production symptom:TestPickRouteKeepsRegionLatchWhenBalancedpins the balanced case so the hysteresis cannot be widened into region thrashing.Validation
Only
internal/executor/pool.gochanges; no migration, no API, no console change.scripts/release-notes.py validatepasses with the added bilingual fragment.Deployed verification
Patched locally against a live instance: 12 fresh requests spread across 23
cnaccounts at 1.0–2.4s TTFB, where the same instance previously pinned 100% of traffic to 2globalaccounts.