Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion .github/workflows/remote.yml
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
name: Remote deployment

# The CDK project under remote/ that `spinloop remote` drives. It is TypeScript
# The CDK project under remote/ that `spinloop cloud` drives. It is TypeScript
# rather than Go, so it has its own workflow; the paths filter keeps it off
# pull requests that do not touch it.
on:
Expand Down
10 changes: 5 additions & 5 deletions AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -18,8 +18,8 @@ This file provides guidance to coding agents, such as Claude Code, when working

- **Harness configuration** — deep-merges provider settings into a coding agent's config. Three harnesses are supported: [opencode](https://opencode.ai) (config under `${XDG_CONFIG_HOME:-$HOME/.config}/opencode`), [Pi](https://github.com/earendil-works/pi) (`~/.pi/agent/models.json`), and [lucinate](https://github.com/lucinate-ai/lucinate) (`~/.lucinate/connections.json`). The harness is chosen at runtime — never baked into a Spinloop file — so the same selection applies to any of them.
- **Local hosting** — `spinloop serve` runs an inference engine (llama.cpp, oMLX) directly; `spinloop daemon` supervises one instead, exposing an HTTP control API for starting/stopping it and reading its status, metrics, and logs.
- **Cloud deployment** — `spinloop remote` drives a scale-to-zero GPU instance through its own AWS control plane (`remote/`): starts it on demand, deploys a model to it, and stops it when idle.
- **Fleet monitoring and control** — `spinloop fleet` observes and drives every engine you run — local daemons and remote environments alike — across machines from one place, including an interactive dashboard, and can route a harness launch to whichever node already has (or can load) the wanted model.
- **Cloud deployment** — `spinloop cloud` drives a scale-to-zero GPU instance through its own AWS control plane (`remote/`): starts it on demand, deploys a model to it, and stops it when idle.
- **Fleet monitoring and control** — `spinloop fleet` observes and drives every engine you run — local daemons and cloud environments alike — across machines from one place, including an interactive dashboard, and can route a harness launch to whichever node already has (or can load) the wanted model.
- **Fleet gateway** — `spinloop gateway` serves a fleet under one OpenAI-compatible endpoint: each request is answered by the fleet's own selector, and a stopped node is started — or the request refused — the way the fleet file's wake policy says.
- **Fleet orchestration** — `spinloop orchestrator` works a backlog of work items against the fleet, through the gateway, at the pace the fleet file declares: each admitted item runs as a one-shot agent of the active harness, its inference going through the gateway.

Expand Down Expand Up @@ -65,18 +65,18 @@ The binary lives under `cmd/`; domain logic is split into `internal/` packages s
- `internal/pi` — Pi's `models.json` IO: deep-merge of one managed provider, preserving siblings and unknown fields. (`pi-integration`)
- `internal/lucinate` — lucinate's `connections.json` IO: one managed connection, no secret ever written to disk. (`lucinate-integration`)
- `cmd/spinloop/fleet.go`, `metrics_render.go`, `status_render.go`, `fleet_dashboard.go`, `dashboard_*.go` — the `fleet` command group and its Bubble Tea dashboard (the CLI's only TUI). (`fleet-client`, `fleet-config`)
- `internal/fleet` — the fleet client: the `Node` interface (`daemonNode`/`remoteNode`), concurrent fan-out (each reading stamped with the time its call returned), the `StartPhase` a start reports and the one function that renders it, and routing/waking a node for a launch. (`fleet-client`, `fleet-config`, `fleet-routing`, `remote-node`)
- `internal/fleet` — the fleet client: the `Node` interface (`daemonNode`/`cloudNode`), concurrent fan-out (each reading stamped with the time its call returned), the `StartPhase` a start reports and the one function that renders it, and routing/waking a node for a launch. (`fleet-client`, `fleet-config`, `fleet-routing`, `remote-node`)
- `cmd/spinloop/gateway.go` + `internal/gateway` — the `gateway` command and the OpenAI-compatible front of a fleet: caller authentication, the cached fan-out behind the model list and routing, the wake behind a request where the file allows it, and the `/v1/fleet` topology the orchestrator reads. (`fleet-gateway`)
- `cmd/spinloop/orchestrator.go` + `internal/orchestrator` — the `orchestrator` command and the loop that works an items backlog against the fleet: the items file, matching and admission against the gateway's topology under the file's declared concurrency, the state kept beside the items file, and the dispatch that runs each admitted item as a one-shot harness agent through the gateway. (`fleet-orchestrator`, `fleet-config`)
- `examples/fleet-docker/` — a runnable multi-node fleet that doubles as the fleet integration test, run per PR by CI. (`fleet-docker-example`)
- `cmd/spinloop/remote.go` + `internal/remote` — the `remote` command group and the scale-to-zero cloud GPU control plane (SigV4-signed Lambda Function URL calls — the repo's only AWS/network dependency). (`remote-environments`, `endpoint-lifecycle`, `endpoint-provisioning`, `remote-endpoint`, `remote-seed`, `weight-seeding`, `remote-keep`, `remote-start-probe`)
- `cmd/spinloop/cloud.go` + `internal/cloud` — the `cloud` command group and the scale-to-zero cloud GPU control plane (SigV4-signed Lambda Function URL calls — the repo's only AWS/network dependency). (`remote-environments`, `endpoint-lifecycle`, `endpoint-provisioning`, `remote-endpoint`, `remote-seed`, `weight-seeding`, `remote-keep`, `remote-start-probe`)
- `internal/daemon` — the engine supervisor and the HTTP control API. `Routes()` in `api.go` is checked against `docs/openapi.yaml` by `openapi_test.go` — keep them in sync when adding an endpoint. Depends on no cloud package: what to serve comes from `internal/inference`. (`daemon-api`, `daemon-api-contract`, `engine-activity`, `engine-metrics`, `api-logging`, `serve-daemon`)
- `internal/inference` — `DeployConfig`, the runner-neutral description of what an engine should serve, shared by every node kind that runs one. A leaf: standard library only, so no node kind has to depend on how another is reached.
- `internal/contextsize` — parses human-friendly sizes (`128k`, `1.5m`) for `CONTEXT`/`OUTPUT`.
- `internal/preset` — parses llama.cpp-style preset `.ini` files, dialect-aware (LlamaCpp vs. OMLX). (`inference-runners`)
- `internal/catalog/providers.yaml` — externalised provider plumbing (URLs, key env vars, npm packages) — no model ids. Add providers here, not in Go. (`provider-catalog`)
- `examples/` — runnable guides, each a directory with a README and a `Spinloop`.
- `remote/` — the TypeScript CDK project `spinloop remote` drives (Lambdas, EC2 Image Builder, S3 weights), built and tested by pnpm with its own CI job. **Public repo:** nothing identifying a deployment (account ids, ARNs, hosts, bucket names) may be committed — enforced by `scripts/check-no-cloud-identifiers.sh`.
- `remote/` — the TypeScript CDK project `spinloop cloud` drives (Lambdas, EC2 Image Builder, S3 weights), built and tested by pnpm with its own CI job. **Public repo:** nothing identifying a deployment (account ids, ARNs, hosts, bucket names) may be committed — enforced by `scripts/check-no-cloud-identifiers.sh`.

## Invariants to keep

Expand Down
52 changes: 26 additions & 26 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -86,20 +86,20 @@ Start the whole fleet with `spinloop up` or just one with `spinloop fleet start

### 3. On a cloud GPU, for as long as you need one

Nothing on your desk with a big enough card? `spinloop remote` drives a
Nothing on your desk with a big enough card? `spinloop cloud` drives a
scale-to-zero instance in your own AWS account: it boots when you ask, loads the
model, and stops itself once you stop using it.

```sh
spinloop remote start # boot, wait for the model to load, print the endpoint
spinloop remote keep 4h # hold it against the idle sweep while you work
spinloop remote stop # terminate now, rather than waiting for the idle timer
spinloop cloud start # boot, wait for the model to load, print the endpoint
spinloop cloud keep 4h # hold it against the idle sweep while you work
spinloop cloud stop # terminate now, rather than waiting for the idle timer
```

A remote is also just another node: give it `kind: remote` in `fleet.yaml` and it
A cloud environment is also just another node: give it `kind: cloud` in `fleet.yaml` and it
sits on the same board as the machines you own. That is `vllm-1` in the picture
above — configured, no instance running, costing nothing until someone presses
`s`. → [Remote inference instance](#remote-inference-instance)
`s`. → [Cloud inference instance](#cloud-inference-instance)

### Then point your agent at it

Expand Down Expand Up @@ -302,8 +302,8 @@ spinloop harness open [<spinloop>] [-H <name>] [--spinloop[=<path>]] [args...]
# launch the harness (a leading Spinloop or alias is
# applied first)
spinloop completion <shell> # tab completion (bash, zsh, powershell)
spinloop remote <bootstrap|bake|start|pause|stop|restart|status|metrics|logs|deploy|env|ls|keep|seed> [path]
# control the remote GPU inference instance
spinloop cloud <bootstrap|bake|start|pause|stop|restart|status|metrics|logs|deploy|env|ls|keep|seed> [path]
# control the cloud GPU inference instance
# (bootstrap does the once-per-account setup;
# bake bakes the runner AMI(s) it launches from;
# deploy sets what it serves, from the Spinloop;
Expand Down Expand Up @@ -531,7 +531,7 @@ nodes:
engine:
port: 18080 # only when the daemon cannot report the engine's address
- name: qwen
kind: remote # a `spinloop remote` environment, driven as a fleet node
kind: cloud # a `spinloop cloud` environment, driven as a fleet node
```

A node's `host`/`port` are the **daemon's**, not the model server's — those are
Expand Down Expand Up @@ -615,14 +615,14 @@ Writing a client? [`docs/openapi.yaml`](docs/openapi.yaml) is the full
contract, and it ships with every release. See
[`docs/http-api.md`](docs/http-api.md) for the endpoints in prose.

## Remote inference instance
## Cloud inference instance

Running a model on your own cloud GPU box? [`remote/`](remote/) deploys one.
`spinloop remote` drives its scale-to-zero lifecycle: the instance only exists
`spinloop cloud` drives its scale-to-zero lifecycle: the instance only exists
while you are using it, and stops itself after a period of idleness.

```sh
spinloop remote start --env dev-2 --print-env # boot the instance, wait for the
spinloop cloud start --env dev-2 --print-env # boot the instance, wait for the
# model to load, then print OPENAI_BASE_URL /
# OPENAI_API_KEY exports for eval
spinloop status --env <name> --env dev-2 # instance state, endpoint health,
Expand All @@ -631,27 +631,27 @@ spinloop metrics --env <name> --env dev-2 # tokens, GPU, CPU and RAM
# the same last-active
spinloop logs --env <name> --env dev-2 # what the engine (or the boot)
# said, even after it's gone
spinloop remote pause --env dev-2 # stop now, but keep it re-wakeable
spinloop remote restart --env dev-2 # fresh engine, same address: stop
spinloop cloud pause --env dev-2 # stop now, but keep it re-wakeable
spinloop cloud restart --env dev-2 # fresh engine, same address: stop
# it, then wake it
spinloop remote keep 4h --env dev-2 # hold it against the idle sweep
spinloop cloud keep 4h --env dev-2 # hold it against the idle sweep
# for 4 hours (start --keep does the same at wake time)
spinloop remote stop --env dev-2 # terminate now instead of waiting
spinloop cloud stop --env dev-2 # terminate now instead of waiting
# for the idle timer
```

Instances ship their engine and boot output to CloudWatch, so `spinloop remote
Instances ship their engine and boot output to CloudWatch, so `spinloop cloud
logs` still works once the instance has terminated — including for a start that
failed before the engine came up (`--source boot`). See
[docs/commands/remote.md](docs/commands/remote.md#reading-the-logs).
[docs/commands/cloud.md](docs/commands/cloud.md#reading-the-logs).

Configuration lives in a `remote.json` per **environment**, named with the
Configuration lives in a `cloud.json` per **environment**, named with the
`--env` flag on every command above (`--env dev-2` selects the file at
`remotes/dev-2/remote.json` under spinloop's config directory,
`clouds/dev-2/cloud.json` under spinloop's config directory,
`${SPINLOOP_CONFIG_DIR:-${XDG_CONFIG_HOME:-~/.config}/spinloop}`); with no
`--env`, the `default` environment is used. The Spinloop itself says only what
the environment serves — the name is a machine-local choice, so it stays out of
the file. `spinloop remote deploy --env dev-2` writes the file for you when it
the file. `spinloop cloud deploy --env dev-2` writes the file for you when it
registers the environment; deploying [`remote/`](remote/) yourself prints the
same values:

Expand All @@ -669,21 +669,21 @@ machine? `spinloop harness open --env dev-2` on its own — no Spinloop at all
configures the harness straight from what is deployed there:

```sh
spinloop remote deploy path/to/Spinloop --env dev-2 # from wherever you deployed it
spinloop cloud deploy path/to/Spinloop --env dev-2 # from wherever you deployed it
spinloop harness open --env dev-2 --prompt "..." # from anywhere with dev-2 registered
```

Every URL and the region can be overridden with the matching
[`SPINLOOP_REMOTE_*`](docs/env-vars.md) environment variable. The commands
[`SPINLOOP_CLOUD_*`](docs/env-vars.md) environment variable. The commands
sign with an AWS credential resolved per region: explicit environment
credentials or a named profile first, then the stored control-plane credential
from [`spinloop remote auth --store`](docs/commands/remote.md#credentials),
from [`spinloop cloud auth --store`](docs/commands/cloud.md#credentials),
then the standard chain (config files, SSO sessions, instance metadata). The
credential needs `lambda:InvokeFunctionUrl` allowed. A cold `start` takes a
few minutes while the instance boots and loads the model; `--timeout`
(default 15m) caps the wait.

The AWS credentials, region and `SPINLOOP_REMOTE_*` overrides can all travel
The AWS credentials, region and `SPINLOOP_CLOUD_*` overrides can all travel
with the Spinloop, in the `.env` beside it. A value already set in your shell wins over the `.env`. To pin a value
in the Spinloop itself, add an `ENV` line (`ENV AWS_PROFILE=prod`) — it may repeat
and overrides both the `.env` and your shell. `ENV` applies only on your
Expand All @@ -700,7 +700,7 @@ Bedrock authenticates through your AWS credentials.

`spinloop harness open` carries that same local environment to the agent it launches:
the whole `.env` beside the active Spinloop fills gaps, and the Spinloop's `ENV` lines
override both your shell and the `.env` — the same precedence the `spinloop remote`
override both your shell and the `.env` — the same precedence the `spinloop cloud`
commands use. These variables shape only the launched agent; `spinloop` never
changes its own environment.

Expand Down
26 changes: 13 additions & 13 deletions cmd/spinloop/alias_test.go
Original file line number Diff line number Diff line change
Expand Up @@ -340,7 +340,7 @@ func TestApply_ByAlias(t *testing.T) {
registerSpinloop(t, "PROVIDER llamacpp\nMODEL gemma\nALIAS q3\n")
t.Chdir(t.TempDir()) // somewhere else entirely

// The alias line goes to stderr, so `spinloop remote env` can be eval'd.
// The alias line goes to stderr, so `spinloop cloud env` can be eval'd.
var stdout string
stderr := captureStderr(t, func() {
stdout = captureStdout(t, func() {
Expand Down Expand Up @@ -538,7 +538,7 @@ func TestEnvAlias_SuppliesTheSpinloop(t *testing.T) {
t.Chdir(t.TempDir()) // no Spinloop here
t.Setenv("SPINLOOP_ALIAS", "q3")

// Like the alias-argument note, this belongs on stderr so `spinloop remote
// Like the alias-argument note, this belongs on stderr so `spinloop cloud
// env` stays eval-able.
var stdout string
stderr := captureStderr(t, func() {
Expand Down Expand Up @@ -803,13 +803,13 @@ func TestEnvAlias_ReachesServe(t *testing.T) {
}
}

// TestEnvAlias_ReachesRemote checks the case that first caught this out: a
// `remote` subcommand with no argument only consults a Spinloop when one is
// TestEnvAlias_ReachesCloud checks the case that first caught this out: a
// `cloud` subcommand with no argument only consults a Spinloop when one is
// there to consult, and SPINLOOP_ALIAS names one as surely as a ./Spinloop does
// — here for its ENV instructions, which name the control plane. Without this
// the command fell through to the per-user default config and reported the
// endpoint as unconfigured.
func TestEnvAlias_ReachesRemote(t *testing.T) {
func TestEnvAlias_ReachesCloud(t *testing.T) {
isolateConfig(t)
stubAWSEnv(t)

Expand All @@ -821,8 +821,8 @@ func TestEnvAlias_ReachesRemote(t *testing.T) {
// behind.
dir := t.TempDir()
mustWrite(t, filepath.Join(dir, spinloop.DefaultFile),
"PROVIDER openai-compatible\nALIAS q3\nENV SPINLOOP_REMOTE_START_URL="+server.URL+"\nENV SPINLOOP_REMOTE_STOP_URL="+server.URL+"\nENV SPINLOOP_REMOTE_REGION=eu-west-1\n")
unsetEnvOnCleanup(t, "SPINLOOP_REMOTE_START_URL", "SPINLOOP_REMOTE_STOP_URL", "SPINLOOP_REMOTE_REGION")
"PROVIDER openai-compatible\nALIAS q3\nENV SPINLOOP_CLOUD_START_URL="+server.URL+"\nENV SPINLOOP_CLOUD_STOP_URL="+server.URL+"\nENV SPINLOOP_CLOUD_REGION=eu-west-1\n")
unsetEnvOnCleanup(t, "SPINLOOP_CLOUD_START_URL", "SPINLOOP_CLOUD_STOP_URL", "SPINLOOP_CLOUD_REGION")
captureStdout(t, func() {
if err := cmdAlias([]string{dir}); err != nil {
t.Fatalf("cmdAlias: %v", err)
Expand All @@ -832,8 +832,8 @@ func TestEnvAlias_ReachesRemote(t *testing.T) {
t.Chdir(t.TempDir()) // no ./Spinloop, so only the variable can find it
t.Setenv("SPINLOOP_ALIAS", "q3")

if err := cmdRemoteStop([]string{"--env", "default"}); err != nil {
t.Fatalf("cmdRemoteStop with SPINLOOP_ALIAS: %v", err)
if err := cmdCloudStop([]string{"--env", "default"}); err != nil {
t.Fatalf("cmdCloudStop with SPINLOOP_ALIAS: %v", err)
}
select {
case name := <-hit:
Expand Down Expand Up @@ -879,14 +879,14 @@ func TestEnvAlias_RemoteFailsRatherThanFallingBack(t *testing.T) {
t.Chdir(t.TempDir())
t.Setenv("SPINLOOP_ALIAS", "nope")

err := cmdRemoteStop([]string{"--env", "default"})
err := cmdCloudStop([]string{"--env", "default"})
if err == nil {
t.Fatal("expected an error for an unregistered SPINLOOP_ALIAS")
}
if !strings.Contains(err.Error(), "SPINLOOP_ALIAS") {
t.Errorf("error %q does not name the variable", err)
}
if strings.Contains(err.Error(), "remote is not configured") {
if strings.Contains(err.Error(), "cloud is not configured") {
t.Errorf("the variable was passed over for the default config: %v", err)
}
}
Expand All @@ -903,11 +903,11 @@ func TestEnvAlias_RemoteFallsBackWithoutREMOTE(t *testing.T) {
t.Chdir(t.TempDir())
t.Setenv("SPINLOOP_ALIAS", "q3")

err := cmdRemoteStop([]string{"--env", "default"})
err := cmdCloudStop([]string{"--env", "default"})
if err == nil {
t.Fatal("expected an error: there is no default endpoint config either")
}
if !strings.Contains(err.Error(), "remote is not configured") {
if !strings.Contains(err.Error(), "cloud is not configured") {
t.Errorf("error = %q, want the default-config failure (the fallback still applies)", err)
}
}
Expand Down
Loading
Loading