Skip to content

fix(orchestration): hold queue claims until startup completes, tweak limiter limits and Reclaim stale jobs from running to pending without needing a process restart - #45

Merged
JohnDuprey merged 3 commits into
mainfrom
dev
Sep 28, 2026

Conversation

@Zacgoose

Copy link
Copy Markdown
Contributor

No description provided.

The limiter throttled background work to a fixed 2 as soon as half the HTTP
pool was busy. On the hosted default profile (6 HTTP / 8 BG) routine API-client
traffic tripped that for hours, holding scheduled cache runs at 2 concurrent.

- BackgroundHttpPressureThreshold now defaults to two thirds of HttpPoolSize
  (4 of 6) instead of half.
- New BackgroundHttpPressureConcurrency sets the throttled BG level; defaults
  to half the ceiling (4 of 8), clamped to [min(2, ceiling), ceiling].
ResolveTaskWorkAsync dropped any task whose status was Running, on the
assumption that a worker in this process held it. But Running also reaches
the live graph from storage: it is the pre-invoke durable marker another
process writes. That happens when a run is rehydrated from storage (e.g. an
old container outliving the new one's startup recovery during an App Service
swap) or when the pump's rehydrated copy wins the _activeRuns race against
recovery at startup. The dropped job was marked Skipped, its queue row
deleted, and nothing re-drives Running, so the run never finalized until the
next restart (observed: a 508-task planner stuck at 505/508 for ~31h).

Track ownership with a non-persisted OwnedHere flag set wherever this process
marks a task Running. A Running task not owned here is an interrupted attempt
from a dead owner: count it and run it, or fail it terminally on the third
attempt, exactly as crash recovery does.
JobQueuePump started claiming at host start, while SchedulerService only ran
ResumeInterruptedRunsAsync once the worker pool and storage were ready. A
claim in that window rehydrated its run (with the previous process's stale
Running markers) into _activeRuns, and recovery's reset then landed on a copy
that lost the TryAdd race.

OrchestratorService now exposes a RecoveryDone gate, opened from a finally
around SchedulerService's startup sequence so it opens whether recovery
succeeds, throws, is skipped, or the host stops mid-startup. The pump awaits
it before its first claim. HTTP, health and row enqueueing are unaffected;
rows written meanwhile are claimed on the first cycle after.
@JohnDuprey
JohnDuprey merged commit dc5075d into main Sep 28, 2026
7 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants