Skip to content

fix(remote): serialise starts of one environment with a lock - #263

Draft
outofcoffee wants to merge 3 commits into
mainfrom
fix/wake-lock
Draft

outofcoffee wants to merge 3 commits into
mainfrom
fix/wake-lock

Conversation

@outofcoffee

Copy link
Copy Markdown
Collaborator

Two starts for one remote environment can no longer both launch an instance, whichever client they come from.

Closes #224

Summary

  • A start takes a per-environment lock before it checks the weights or looks up the instance, and releases it on every way a start can end. A start that finds the lock held launches nothing and replies with a retryable 503 starting ("another start is in progress"), which the CLI and gateway already retry.
  • The lock is an SSM parameter created only if absent, /cloud-vm-llm/<env>/wake-lock. It expires when the invocation that took it would time out, so a start killed before it could release the lock blocks its environment for at most about 15 minutes. An expired or unreadable lock is taken over.
  • A start whose instance is stopped, stopping or terminated under it now ends at once with a retryable reply and lets the lock go. Before, it polled until its deadline.
  • status shows a start in progress: the state is starting while the lock is held and the instance is absent, stopped or pending, and every reply carries start_in_progress. spinloop status prints it with no client change.
  • The start Lambda's role gains ssm:DeleteParameter on the lock parameter only.
  • OpenSpec change wake-lock (a delta to endpoint-lifecycle), with docs and tests.

Implementation details

  • The lock covers the whole run of a start, not only the launch. The tag lookup that decides whether to launch lags a launch, so releasing at launch would leave the race open.
  • Stops do not take the lock. A lock held for up to 15 minutes would refuse the stop that gets you out of a hung start, and the stop Lambda's 120 second limit could not wait for it. The start ends itself when it sees its instance stopped, which is what keeps a stop mid-start from blocking the environment.
  • A retryable reply after a stop means a client still waiting on its start will wake the instance again. That matches what happened before once the deadline passed, only sooner.
  • SSM has no compare-and-set, so two starts taking over the same abandoned lock within milliseconds can still both proceed. That needs a killed start plus a near-simultaneous pair. A DynamoDB conditional write would close it, at the cost of a new table; the design records the trade-off.
  • A start waiting for GPU capacity holds no lock between attempts, so status does not show it as starting.
  • Existing deployments need the stack redeployed (pnpm run deploy or spinloop remote bootstrap) for the new grant.

The OpenSpec change is not archived yet.

spinloop-agent added 3 commits October 5, 2026 23:41
A start took no lock, so two starts close together could each miss the
other's not-yet-visible instance and launch one apiece. Starts now take a
per-environment SSM lock that expires with the invocation, a refused start
replies with a retryable 503, and a start whose instance is stopped under it
ends at once and releases the lock. Status reports a start in progress.

Closes #224
@outofcoffee outofcoffee added the bug Something isn't working label Oct 5, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bug Something isn't working

Projects

None yet

Development

Successfully merging this pull request may close these issues.

bug: the remote control plane's wake is not idempotent against a concurrent start

1 participant