Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
11 changes: 11 additions & 0 deletions docs/commands/remote.md
Original file line number Diff line number Diff line change
Expand Up @@ -254,6 +254,17 @@ output degrades rather than goes blank. `--format=table` is a key-value
table, and `--format=json` is the raw reply. A stopped engine keeps its
readings: the sparkline runs to the stop, ending at it.

While a start is under way, `status` shows **`starting`** for the endpoint,
whichever machine began the start, until the instance is running. Once it is
running but still loading the model it reads `running` and not ready, as
before. Two starts for one endpoint never launch two instances: the second is
told "another start is in progress" and retries on its own, so running
`start` twice, or from two machines, is safe. A start that is waiting for GPU
capacity is between attempts, not in progress, and `status` does not show it.
Stopping an endpoint during a start is always allowed; the start ends and
reports that the instance was stopped, and a client that keeps retrying will
wake it again.

Both report **`active`** — how long since the endpoint's engine last did
any work. It comes from the activity the on-instance daemon tracks, so it is
one answer decided on the box rather than something each command re-derives
Expand Down
2 changes: 2 additions & 0 deletions docs/maintainer/internals.md
Original file line number Diff line number Diff line change
Expand Up @@ -35,6 +35,8 @@ These are mistakes already made here; each was silent rather than loud, which is

**An IAM user's inline policies are capped at 2,048 characters in aggregate.** The control plane's seven functions each take a `grantInvokeUrl` pair — two actions, the auth-type conditions, the function's ARN — and with the log-reading, stack-discovery, pricing and self-service statements the document far exceeds that; the first deploy of the `RemoteCliUser` inline policy failed with `ServiceLimitExceeded`, which CDK does not warn about ahead of time. It is now a stack-owned `AWS::IAM::ManagedPolicy` (`RemoteCliPolicy`, 6,144 cap; the deployed document measures ~2 KB, so the grant list has room to grow). Keep it managed rather than re-inlining it, and keep the iam self-service ARN built from the `AWS::Partition`/`AWS::AccountId` pseudo parameters instead of the user's `Arn`: the policy attaches to that user, so referencing the user from inside it is a dependency cycle.

**A start holds a per-environment lock for its whole run, and it is an SSM parameter, not a conditional write.** `wake()` in the start Lambda takes `/cloud-vm-llm/<env>/wake-lock` (created with `Overwrite: false`) after the read-only checks and before the weights check, and releases it in a `finally`. It has to cover the whole poll and not only the launch, because the tag lookup that decides whether to launch lags a launch: a second start that got in after the lock was released could still miss the first one's instance. The value is `{owner, expiresAt}` with the expiry set to the invocation's remaining time plus 30 seconds, so a start killed before its `finally` blocks the environment for no longer than it could have run. A refused start replies 503 `starting` with 15 seconds to retry, which the CLI and gateway already retry. Taking over an expired lock is read, re-read, delete, create-if-absent; SSM has no compare-and-set, so two starts taking over the same abandoned lock within milliseconds can still both proceed. That needs a killed start plus a near-simultaneous pair, and a DynamoDB conditional write would remove it. Releasing needs `ssm:DeleteParameter`, granted on `parameter/cloud-vm-llm/*/wake-lock` only. Stops do not take the lock: a lock held for up to 15 minutes would refuse the stop that gets you out of a hung start, and the stop Lambda's 120 second limit could not wait for it. Instead each polling loop in a start checks the instance state and ends the start, releasing the lock, if the instance is stopped, stopping or terminated. The status read reports the lock without taking it: `starting` when the lock is held and the instance is absent, stopped or pending, and `start_in_progress` in every reply. A start waiting for capacity has released its lock between attempts, so it is not shown.

**The file credential store's index is non-secret by design.** OS keystores offer no way to list entries, so the file store — used where no keystore is reachable or `SPINLOOP_REMOTE_KEYSTORE=file` — keeps a plain-text index of the stored regions beside the `0600` per-region files under `<config>/keystore/`. The report (`spinloop remote auth`) reads the index, so a file added, removed or renamed by hand is reported wrong until the index matches; and a corrupt index is reported, not silently reset, because a report that misleads about what is stored misleads about which access keys exist on the AWS side.

**The two local model caches are separate.** A model `llama-server` downloaded sits in llama.cpp's cache (`$LLAMA_CACHE`, else the platform's user cache directory) as flat filenames; one fetched with the hub's tools sits in the Hugging Face hub cache (`$HF_HUB_CACHE`, else `$HF_HOME/hub`, else `~/.cache/huggingface/hub`) as `models--<owner>--<name>/snapshots/<sha>/` of symlinks into a content-addressed blob store. Neither tool looks in the other's, so a model already on the machine is "already on the machine" in only one of the two senses. `spinloop hf` therefore resolves both roots up front (`hf.ResolveRoots`) and checks both before touching the network; a cache-aware lookup that consults only one side re-downloads what is already there. The hub-cache shape has two more traps: a snapshot entry whose symlink dangles is an interrupted or half-finished download and must count as absent, and `refs/<revision>` holds a commit sha, so a revision name is only a snapshot once it has been resolved through `refs/`.
Expand Down
6 changes: 6 additions & 0 deletions docs/troubleshooting.md
Original file line number Diff line number Diff line change
Expand Up @@ -86,6 +86,12 @@ scaled; see [`spinloop serve`](commands/serve.md#parallelism).
stderr; `--timeout` (default 15m) bounds the wait. `status` and `logs`
answer while it boots and after it is gone — logs are readable even from a
terminated instance.
- **`start` says another start is in progress.** Another command, machine or
gateway is already starting that endpoint, and only one start works on an
endpoint at a time. `start` retries on its own and finishes when the other
one does. If nothing is really starting, the lock a crashed start left
behind expires within about fifteen minutes. `status` shows `starting`
meanwhile.
- **Quota.** Bootstrap needs enough GPU vCPU quota for a later launch; a launch
that can't get an instance reports the AWS error.

Expand Down
2 changes: 2 additions & 0 deletions openspec/changes/wake-lock/.openspec.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,2 @@
schema: spec-driven
created: 2026-10-05
53 changes: 53 additions & 0 deletions openspec/changes/wake-lock/design.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,53 @@
## Context

See proposal.md for why. `wake()` in `remote/lambda/start/index.ts` reads the deploy config, finds the environment's Elastic IP and security group, runs the weights check (which can start a seed), looks the instance up by tag, then launches or re-wakes it and polls until the model answers. That poll can run for the Lambda's whole 15-minute limit. The tag lookup is `DescribeInstances`, which is eventually consistent, so it is not a safe way to tell whether another start has already launched.

The Lambda's role already has `ssm:GetParameter` and `ssm:PutParameter` on `/cloud-vm-llm/*`. It has no `ssm:DeleteParameter`. Issue 224 says no new grant is needed; releasing a lock by deleting the parameter does need one.

## Goals / Non-Goals

**Goals:**
- Two starts for one environment never launch two instances, from any caller.
- A start that cannot take the lock does nothing but reply that it should retry.
- A crashed start cannot block an environment for longer than the time that start could have run.

**Non-Goals:**
- Serialising stops, pauses or the idle sweep against starts. Stop is the way out of a start that has hung, so a lock held for up to 15 minutes would refuse it exactly when it is needed. The stop Lambda times out at 120 seconds, so it could not wait for the lock either. It would also need get, put and delete on the lock parameter, and the 5-minute sweep would make extra SSM calls for every environment. Stops race a start in a different way, and the start now handles that by ending promptly when it sees the stop (see Decisions).
- Making the lock safe against two starts that both find the same expired lock in the same few milliseconds. See Risks.
- A queue, so that a refused start waits its turn on the server. The caller retries.
- Any change to the CLI or gateway.

## Decisions

**An SSM parameter as the lock.** `/cloud-vm-llm/<env>/wake-lock`, created with `Overwrite: false`, so creation fails with `ParameterAlreadyExists` when another start holds it. This needs no new infrastructure. DynamoDB with a conditional put would give true compare-and-set but adds a table, a construct and grants to a stack that has none today; the SSM lock is enough for the case that matters (starts racing within seconds of each other, where creation is atomic).

**What the lock holds.** A JSON value `{"owner": <Lambda request id>, "expiresAt": <ISO time>}`. The owner lets a release remove only a lock this start created. `expiresAt` is now plus the time remaining on this invocation plus a 30 second margin, so a lock lives exactly as long as the start that took it could run. A fixed TTL longer than the 900 second limit would leave a killed start's environment blocked for longer than needed.

**Where it is taken and released.** `wake()` keeps its signature and becomes a wrapper: it runs the existing read-only checks that reply "unconfigured" and "undeployed" (they change nothing, so they need no lock), takes the lock, runs the rest of the existing body as `wakeLocked()`, and releases the lock in a `finally`. The lock is taken before the weights check, because that check can start a seed and two starts racing there could start two. The lock is held through the whole poll, not only the launch: the lookup cannot be trusted to see a just-launched instance for a while, so releasing at launch would reopen the race.

**What a refused start returns.** HTTP 503 with state `starting`, a message saying another start for the environment is in progress, and `retry_after_seconds: 15` with the matching `retry-after` header. The Go client already retries any 503 using `retry_after_seconds` and prints "instance starting; retrying in 15s", and the gateway wake path treats a 503 the same way, so neither changes. When the holder finishes, the retry finds the running instance and returns ready.

**Taking over an expired lock.** When creation fails because the parameter exists, the start reads it. If `expiresAt` is in the future, it replies as above. If it has passed, the start reads the lock again, checks the owner and expiry are unchanged, deletes it, and tries the create-if-absent again; whichever start's create succeeds holds the lock, and the other replies as above. An unreadable or malformed lock value counts as expired, since nothing else could ever clear it.

**Ending a start whose instance was stopped.** Today, after a start has issued its own start command, an instance seen `stopped` has no branch in the first poll loop, and the later loops (agent online, daemon answering, health) never look at the instance state, so a stop mid-wake leaves the start polling until its deadline. With the lock that would also block the environment for those minutes. Each iteration of those loops now checks the instance state with `getInstance`, tolerating `InvalidInstanceID.NotFound` as the first loop already does, and when it sees `stopped`, `stopping`, `shutting-down` or `terminated` after the start command was issued it returns 503 with the observed state, a message that the instance was stopped while starting, and `retry_after_seconds: 15`. The `finally` in `wake()` releases the lock. The existing first-loop behaviours are unchanged: a `stopped` instance seen before the start command was issued is re-woken, and a dying instance in that loop keeps its existing reply.

**Retrying is the client's choice, and it can undo a deliberate stop.** The reply is retryable, so a `spinloop remote start` that is still waiting will ask again and re-wake the instance a user just paused. That matches what happens today once the deadline passes, only sooner. A non-retryable reply would honour the stop but would make a scheduled 18:00 stop that lands during a slow start fail the start with an error rather than being a hiccup. This change keeps the reply retryable and does not try to decide whose intent wins.

**Showing a start in progress.** The GET status branch of the start Lambda reads the lock (`wakeLockHeld`: the parameter exists, parses, and has not expired) and adds `start_in_progress` to its reply. When the lock is held and the instance is absent, `stopped` or `pending`, `state` is `starting`, so `spinloop status`, which prints the state string as it is, shows it with no change to the Go client. When the instance is already `running` the state stays `running` with `healthy: false` and only the flag marks the start, so nothing that keys off `running` changes. A failure to read the lock is logged and treated as no start in progress, because a status read that fails over the lock is worse than one that omits it. Reads take no lock and write nothing.

**What status does not show.** A start waiting for capacity holds no lock between attempts (it released the lock when it replied no-capacity), so an environment in that wait looks idle. Showing it would need a record of the last start result, which is a separate feature.

**Failures other than "already exists".** Any other error creating the lock is logged and answered with a 503 `starting`, retryable, with nothing launched. A failure releasing the lock is logged and ignored: the reply to the caller is already decided, and the lock expires on its own.

**IAM.** One statement on the start Lambda's role: `ssm:DeleteParameter` on `parameter/cloud-vm-llm/*/wake-lock`. Create and read use the existing grants.

## Risks / Trade-offs

- **Two starts taking over the same expired lock together.** SSM has no compare-and-set, so between one start's re-read and its delete, another can delete and re-create the lock, and the first then deletes the new one. This needs an abandoned lock (a killed start) and two starts arriving within milliseconds after its expiry. The result is the old race for that one start, no worse than today. A DynamoDB lock would remove it; this change accepts it and records it.
- **A killed start blocks the environment until the lock expires.** At most the remaining time of that invocation plus 30 seconds, which is at most about 15 minutes. Callers see "another start is in progress" for that time.
- **A refused start does not wait on the server.** Callers poll every 15 seconds, so a start can learn the instance is ready up to 15 seconds late.
- **Deployments that have not taken the new grant.** The start cannot delete its lock, so a start would hold the environment until expiry. The release failure is logged. Deploying the stack fixes it; the proposal says so.

## Open Questions

None that change what gets built.
33 changes: 33 additions & 0 deletions openspec/changes/wake-lock/proposal.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,33 @@
## Why

The start Lambda decides whether to launch an instance by looking one up by tag. That lookup is eventually consistent, so two starts for the same environment that arrive close together can each miss the other's new instance and each launch one: two GPU instances for one environment, both billed (issue 224).

The gateway already coalesces concurrent wakes inside one gateway process, but the control plane is open to every other caller: a second gateway, `spinloop remote start` racing a gateway, and a scheduled start (issue 178) firing while someone starts the environment by hand.

## What Changes

- The start Lambda takes a per-environment lock before it looks anything up and releases it on every way a start can end.
- A start that finds the lock held does not look up, launch or re-wake anything. It replies with a retryable 503 saying another start is in progress, which the CLI and gateway already retry.
- A lock left behind by a start that never finished (the Lambda was killed) expires when that start would have timed out, and the next start takes it over.
- The lock is an SSM parameter created only if absent, `/cloud-vm-llm/<env>/wake-lock`. The start Lambda's role gains `ssm:DeleteParameter` on that parameter name only.
- Locks are per environment: starts for different environments do not wait for each other.
- The status read shows a start in progress: while a start holds the lock, the state is `starting` (until the instance is running) and the reply carries `start_in_progress: true`, so a second client can see that another client began a start. A start waiting for GPU capacity holds no lock between attempts and is not shown.
- A start that sees its instance stopped, stopping or terminated after it has issued its start command now ends at once with a retryable reply, instead of polling until its deadline. Without this, a stop mid-start would leave the environment locked for up to 15 minutes. Stops themselves are not locked and never wait for a start.

## Capabilities

### New Capabilities

None.

### Modified Capabilities

- `endpoint-lifecycle`: adds requirements that concurrent starts of one environment are serialised, that a held or abandoned lock is handled as above, that a start ends promptly when its instance is stopped under it, and that status shows a start in progress. The existing "Starting on demand" behaviour is otherwise unchanged.

## Impact

- `remote/lambda/start/index.ts`: `wake()` becomes a lock wrapper around the existing body; a new shared module holds the lock.
- `remote/lib/llm-stack.ts`: one IAM statement on the start Lambda's role.
- `remote/test/`: new tests for the lock and for `wake()` under contention.
- `docs/maintainer/internals.md` and `remote/README.md`: a short note on the lock.
- No CLI, config or API change. Existing deployments need a `spinloop remote bootstrap` re-run (or `pnpm run deploy`) to get the new grant; until then a start fails at the lock instead of launching unlocked.
Loading
Loading