diff --git a/.github/workflows/remote.yml b/.github/workflows/remote.yml index 9558f19b..8b84fb94 100644 --- a/.github/workflows/remote.yml +++ b/.github/workflows/remote.yml @@ -1,6 +1,6 @@ name: Remote deployment -# The CDK project under remote/ that `spinloop remote` drives. It is TypeScript +# The CDK project under remote/ that `spinloop cloud` drives. It is TypeScript # rather than Go, so it has its own workflow; the paths filter keeps it off # pull requests that do not touch it. on: diff --git a/AGENTS.md b/AGENTS.md index 2822acb1..a4c5fdd6 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -18,8 +18,8 @@ This file provides guidance to coding agents, such as Claude Code, when working - **Harness configuration** — deep-merges provider settings into a coding agent's config. Three harnesses are supported: [opencode](https://opencode.ai) (config under `${XDG_CONFIG_HOME:-$HOME/.config}/opencode`), [Pi](https://github.com/earendil-works/pi) (`~/.pi/agent/models.json`), and [lucinate](https://github.com/lucinate-ai/lucinate) (`~/.lucinate/connections.json`). The harness is chosen at runtime — never baked into a Spinloop file — so the same selection applies to any of them. - **Local hosting** — `spinloop serve` runs an inference engine (llama.cpp, oMLX) directly; `spinloop daemon` supervises one instead, exposing an HTTP control API for starting/stopping it and reading its status, metrics, and logs. -- **Cloud deployment** — `spinloop remote` drives a scale-to-zero GPU instance through its own AWS control plane (`remote/`): starts it on demand, deploys a model to it, and stops it when idle. -- **Fleet monitoring and control** — `spinloop fleet` observes and drives every engine you run — local daemons and remote environments alike — across machines from one place, including an interactive dashboard, and can route a harness launch to whichever node already has (or can load) the wanted model. +- **Cloud deployment** — `spinloop cloud` drives a scale-to-zero GPU instance through its own AWS control plane (`remote/`): starts it on demand, deploys a model to it, and stops it when idle. +- **Fleet monitoring and control** — `spinloop fleet` observes and drives every engine you run — local daemons and cloud environments alike — across machines from one place, including an interactive dashboard, and can route a harness launch to whichever node already has (or can load) the wanted model. - **Fleet gateway** — `spinloop gateway` serves a fleet under one OpenAI-compatible endpoint: each request is answered by the fleet's own selector, and a stopped node is started — or the request refused — the way the fleet file's wake policy says. - **Fleet orchestration** — `spinloop orchestrator` works a backlog of work items against the fleet, through the gateway, at the pace the fleet file declares: each admitted item runs as a one-shot agent of the active harness, its inference going through the gateway. @@ -65,18 +65,18 @@ The binary lives under `cmd/`; domain logic is split into `internal/` packages s - `internal/pi` — Pi's `models.json` IO: deep-merge of one managed provider, preserving siblings and unknown fields. (`pi-integration`) - `internal/lucinate` — lucinate's `connections.json` IO: one managed connection, no secret ever written to disk. (`lucinate-integration`) - `cmd/spinloop/fleet.go`, `metrics_render.go`, `status_render.go`, `fleet_dashboard.go`, `dashboard_*.go` — the `fleet` command group and its Bubble Tea dashboard (the CLI's only TUI). (`fleet-client`, `fleet-config`) -- `internal/fleet` — the fleet client: the `Node` interface (`daemonNode`/`remoteNode`), concurrent fan-out (each reading stamped with the time its call returned), the `StartPhase` a start reports and the one function that renders it, and routing/waking a node for a launch. (`fleet-client`, `fleet-config`, `fleet-routing`, `remote-node`) +- `internal/fleet` — the fleet client: the `Node` interface (`daemonNode`/`cloudNode`), concurrent fan-out (each reading stamped with the time its call returned), the `StartPhase` a start reports and the one function that renders it, and routing/waking a node for a launch. (`fleet-client`, `fleet-config`, `fleet-routing`, `remote-node`) - `cmd/spinloop/gateway.go` + `internal/gateway` — the `gateway` command and the OpenAI-compatible front of a fleet: caller authentication, the cached fan-out behind the model list and routing, the wake behind a request where the file allows it, and the `/v1/fleet` topology the orchestrator reads. (`fleet-gateway`) - `cmd/spinloop/orchestrator.go` + `internal/orchestrator` — the `orchestrator` command and the loop that works an items backlog against the fleet: the items file, matching and admission against the gateway's topology under the file's declared concurrency, the state kept beside the items file, and the dispatch that runs each admitted item as a one-shot harness agent through the gateway. (`fleet-orchestrator`, `fleet-config`) - `examples/fleet-docker/` — a runnable multi-node fleet that doubles as the fleet integration test, run per PR by CI. (`fleet-docker-example`) -- `cmd/spinloop/remote.go` + `internal/remote` — the `remote` command group and the scale-to-zero cloud GPU control plane (SigV4-signed Lambda Function URL calls — the repo's only AWS/network dependency). (`remote-environments`, `endpoint-lifecycle`, `endpoint-provisioning`, `remote-endpoint`, `remote-seed`, `weight-seeding`, `remote-keep`, `remote-start-probe`) +- `cmd/spinloop/cloud.go` + `internal/cloud` — the `cloud` command group and the scale-to-zero cloud GPU control plane (SigV4-signed Lambda Function URL calls — the repo's only AWS/network dependency). (`remote-environments`, `endpoint-lifecycle`, `endpoint-provisioning`, `remote-endpoint`, `remote-seed`, `weight-seeding`, `remote-keep`, `remote-start-probe`) - `internal/daemon` — the engine supervisor and the HTTP control API. `Routes()` in `api.go` is checked against `docs/openapi.yaml` by `openapi_test.go` — keep them in sync when adding an endpoint. Depends on no cloud package: what to serve comes from `internal/inference`. (`daemon-api`, `daemon-api-contract`, `engine-activity`, `engine-metrics`, `api-logging`, `serve-daemon`) - `internal/inference` — `DeployConfig`, the runner-neutral description of what an engine should serve, shared by every node kind that runs one. A leaf: standard library only, so no node kind has to depend on how another is reached. - `internal/contextsize` — parses human-friendly sizes (`128k`, `1.5m`) for `CONTEXT`/`OUTPUT`. - `internal/preset` — parses llama.cpp-style preset `.ini` files, dialect-aware (LlamaCpp vs. OMLX). (`inference-runners`) - `internal/catalog/providers.yaml` — externalised provider plumbing (URLs, key env vars, npm packages) — no model ids. Add providers here, not in Go. (`provider-catalog`) - `examples/` — runnable guides, each a directory with a README and a `Spinloop`. -- `remote/` — the TypeScript CDK project `spinloop remote` drives (Lambdas, EC2 Image Builder, S3 weights), built and tested by pnpm with its own CI job. **Public repo:** nothing identifying a deployment (account ids, ARNs, hosts, bucket names) may be committed — enforced by `scripts/check-no-cloud-identifiers.sh`. +- `remote/` — the TypeScript CDK project `spinloop cloud` drives (Lambdas, EC2 Image Builder, S3 weights), built and tested by pnpm with its own CI job. **Public repo:** nothing identifying a deployment (account ids, ARNs, hosts, bucket names) may be committed — enforced by `scripts/check-no-cloud-identifiers.sh`. ## Invariants to keep diff --git a/README.md b/README.md index 5fbeef99..93934371 100644 --- a/README.md +++ b/README.md @@ -86,20 +86,20 @@ Start the whole fleet with `spinloop up` or just one with `spinloop fleet start ### 3. On a cloud GPU, for as long as you need one -Nothing on your desk with a big enough card? `spinloop remote` drives a +Nothing on your desk with a big enough card? `spinloop cloud` drives a scale-to-zero instance in your own AWS account: it boots when you ask, loads the model, and stops itself once you stop using it. ```sh -spinloop remote start # boot, wait for the model to load, print the endpoint -spinloop remote keep 4h # hold it against the idle sweep while you work -spinloop remote stop # terminate now, rather than waiting for the idle timer +spinloop cloud start # boot, wait for the model to load, print the endpoint +spinloop cloud keep 4h # hold it against the idle sweep while you work +spinloop cloud stop # terminate now, rather than waiting for the idle timer ``` -A remote is also just another node: give it `kind: remote` in `fleet.yaml` and it +A cloud environment is also just another node: give it `kind: cloud` in `fleet.yaml` and it sits on the same board as the machines you own. That is `vllm-1` in the picture above — configured, no instance running, costing nothing until someone presses -`s`. → [Remote inference instance](#remote-inference-instance) +`s`. → [Cloud inference instance](#cloud-inference-instance) ### Then point your agent at it @@ -302,8 +302,8 @@ spinloop harness open [] [-H ] [--spinloop[=]] [args...] # launch the harness (a leading Spinloop or alias is # applied first) spinloop completion # tab completion (bash, zsh, powershell) -spinloop remote [path] - # control the remote GPU inference instance +spinloop cloud [path] + # control the cloud GPU inference instance # (bootstrap does the once-per-account setup; # bake bakes the runner AMI(s) it launches from; # deploy sets what it serves, from the Spinloop; @@ -531,7 +531,7 @@ nodes: engine: port: 18080 # only when the daemon cannot report the engine's address - name: qwen - kind: remote # a `spinloop remote` environment, driven as a fleet node + kind: cloud # a `spinloop cloud` environment, driven as a fleet node ``` A node's `host`/`port` are the **daemon's**, not the model server's — those are @@ -615,14 +615,14 @@ Writing a client? [`docs/openapi.yaml`](docs/openapi.yaml) is the full contract, and it ships with every release. See [`docs/http-api.md`](docs/http-api.md) for the endpoints in prose. -## Remote inference instance +## Cloud inference instance Running a model on your own cloud GPU box? [`remote/`](remote/) deploys one. -`spinloop remote` drives its scale-to-zero lifecycle: the instance only exists +`spinloop cloud` drives its scale-to-zero lifecycle: the instance only exists while you are using it, and stops itself after a period of idleness. ```sh -spinloop remote start --env dev-2 --print-env # boot the instance, wait for the +spinloop cloud start --env dev-2 --print-env # boot the instance, wait for the # model to load, then print OPENAI_BASE_URL / # OPENAI_API_KEY exports for eval spinloop status --env --env dev-2 # instance state, endpoint health, @@ -631,27 +631,27 @@ spinloop metrics --env --env dev-2 # tokens, GPU, CPU and RAM # the same last-active spinloop logs --env --env dev-2 # what the engine (or the boot) # said, even after it's gone -spinloop remote pause --env dev-2 # stop now, but keep it re-wakeable -spinloop remote restart --env dev-2 # fresh engine, same address: stop +spinloop cloud pause --env dev-2 # stop now, but keep it re-wakeable +spinloop cloud restart --env dev-2 # fresh engine, same address: stop # it, then wake it -spinloop remote keep 4h --env dev-2 # hold it against the idle sweep +spinloop cloud keep 4h --env dev-2 # hold it against the idle sweep # for 4 hours (start --keep does the same at wake time) -spinloop remote stop --env dev-2 # terminate now instead of waiting +spinloop cloud stop --env dev-2 # terminate now instead of waiting # for the idle timer ``` -Instances ship their engine and boot output to CloudWatch, so `spinloop remote +Instances ship their engine and boot output to CloudWatch, so `spinloop cloud logs` still works once the instance has terminated — including for a start that failed before the engine came up (`--source boot`). See -[docs/commands/remote.md](docs/commands/remote.md#reading-the-logs). +[docs/commands/cloud.md](docs/commands/cloud.md#reading-the-logs). -Configuration lives in a `remote.json` per **environment**, named with the +Configuration lives in a `cloud.json` per **environment**, named with the `--env` flag on every command above (`--env dev-2` selects the file at -`remotes/dev-2/remote.json` under spinloop's config directory, +`clouds/dev-2/cloud.json` under spinloop's config directory, `${SPINLOOP_CONFIG_DIR:-${XDG_CONFIG_HOME:-~/.config}/spinloop}`); with no `--env`, the `default` environment is used. The Spinloop itself says only what the environment serves — the name is a machine-local choice, so it stays out of -the file. `spinloop remote deploy --env dev-2` writes the file for you when it +the file. `spinloop cloud deploy --env dev-2` writes the file for you when it registers the environment; deploying [`remote/`](remote/) yourself prints the same values: @@ -669,21 +669,21 @@ machine? `spinloop harness open --env dev-2` on its own — no Spinloop at all configures the harness straight from what is deployed there: ```sh -spinloop remote deploy path/to/Spinloop --env dev-2 # from wherever you deployed it +spinloop cloud deploy path/to/Spinloop --env dev-2 # from wherever you deployed it spinloop harness open --env dev-2 --prompt "..." # from anywhere with dev-2 registered ``` Every URL and the region can be overridden with the matching -[`SPINLOOP_REMOTE_*`](docs/env-vars.md) environment variable. The commands +[`SPINLOOP_CLOUD_*`](docs/env-vars.md) environment variable. The commands sign with an AWS credential resolved per region: explicit environment credentials or a named profile first, then the stored control-plane credential -from [`spinloop remote auth --store`](docs/commands/remote.md#credentials), +from [`spinloop cloud auth --store`](docs/commands/cloud.md#credentials), then the standard chain (config files, SSO sessions, instance metadata). The credential needs `lambda:InvokeFunctionUrl` allowed. A cold `start` takes a few minutes while the instance boots and loads the model; `--timeout` (default 15m) caps the wait. -The AWS credentials, region and `SPINLOOP_REMOTE_*` overrides can all travel +The AWS credentials, region and `SPINLOOP_CLOUD_*` overrides can all travel with the Spinloop, in the `.env` beside it. A value already set in your shell wins over the `.env`. To pin a value in the Spinloop itself, add an `ENV` line (`ENV AWS_PROFILE=prod`) — it may repeat and overrides both the `.env` and your shell. `ENV` applies only on your @@ -700,7 +700,7 @@ Bedrock authenticates through your AWS credentials. `spinloop harness open` carries that same local environment to the agent it launches: the whole `.env` beside the active Spinloop fills gaps, and the Spinloop's `ENV` lines -override both your shell and the `.env` — the same precedence the `spinloop remote` +override both your shell and the `.env` — the same precedence the `spinloop cloud` commands use. These variables shape only the launched agent; `spinloop` never changes its own environment. diff --git a/cmd/spinloop/alias_test.go b/cmd/spinloop/alias_test.go index 05a5fbb1..d55e6f96 100644 --- a/cmd/spinloop/alias_test.go +++ b/cmd/spinloop/alias_test.go @@ -340,7 +340,7 @@ func TestApply_ByAlias(t *testing.T) { registerSpinloop(t, "PROVIDER llamacpp\nMODEL gemma\nALIAS q3\n") t.Chdir(t.TempDir()) // somewhere else entirely - // The alias line goes to stderr, so `spinloop remote env` can be eval'd. + // The alias line goes to stderr, so `spinloop cloud env` can be eval'd. var stdout string stderr := captureStderr(t, func() { stdout = captureStdout(t, func() { @@ -538,7 +538,7 @@ func TestEnvAlias_SuppliesTheSpinloop(t *testing.T) { t.Chdir(t.TempDir()) // no Spinloop here t.Setenv("SPINLOOP_ALIAS", "q3") - // Like the alias-argument note, this belongs on stderr so `spinloop remote + // Like the alias-argument note, this belongs on stderr so `spinloop cloud // env` stays eval-able. var stdout string stderr := captureStderr(t, func() { @@ -803,13 +803,13 @@ func TestEnvAlias_ReachesServe(t *testing.T) { } } -// TestEnvAlias_ReachesRemote checks the case that first caught this out: a -// `remote` subcommand with no argument only consults a Spinloop when one is +// TestEnvAlias_ReachesCloud checks the case that first caught this out: a +// `cloud` subcommand with no argument only consults a Spinloop when one is // there to consult, and SPINLOOP_ALIAS names one as surely as a ./Spinloop does // — here for its ENV instructions, which name the control plane. Without this // the command fell through to the per-user default config and reported the // endpoint as unconfigured. -func TestEnvAlias_ReachesRemote(t *testing.T) { +func TestEnvAlias_ReachesCloud(t *testing.T) { isolateConfig(t) stubAWSEnv(t) @@ -821,8 +821,8 @@ func TestEnvAlias_ReachesRemote(t *testing.T) { // behind. dir := t.TempDir() mustWrite(t, filepath.Join(dir, spinloop.DefaultFile), - "PROVIDER openai-compatible\nALIAS q3\nENV SPINLOOP_REMOTE_START_URL="+server.URL+"\nENV SPINLOOP_REMOTE_STOP_URL="+server.URL+"\nENV SPINLOOP_REMOTE_REGION=eu-west-1\n") - unsetEnvOnCleanup(t, "SPINLOOP_REMOTE_START_URL", "SPINLOOP_REMOTE_STOP_URL", "SPINLOOP_REMOTE_REGION") + "PROVIDER openai-compatible\nALIAS q3\nENV SPINLOOP_CLOUD_START_URL="+server.URL+"\nENV SPINLOOP_CLOUD_STOP_URL="+server.URL+"\nENV SPINLOOP_CLOUD_REGION=eu-west-1\n") + unsetEnvOnCleanup(t, "SPINLOOP_CLOUD_START_URL", "SPINLOOP_CLOUD_STOP_URL", "SPINLOOP_CLOUD_REGION") captureStdout(t, func() { if err := cmdAlias([]string{dir}); err != nil { t.Fatalf("cmdAlias: %v", err) @@ -832,8 +832,8 @@ func TestEnvAlias_ReachesRemote(t *testing.T) { t.Chdir(t.TempDir()) // no ./Spinloop, so only the variable can find it t.Setenv("SPINLOOP_ALIAS", "q3") - if err := cmdRemoteStop([]string{"--env", "default"}); err != nil { - t.Fatalf("cmdRemoteStop with SPINLOOP_ALIAS: %v", err) + if err := cmdCloudStop([]string{"--env", "default"}); err != nil { + t.Fatalf("cmdCloudStop with SPINLOOP_ALIAS: %v", err) } select { case name := <-hit: @@ -879,14 +879,14 @@ func TestEnvAlias_RemoteFailsRatherThanFallingBack(t *testing.T) { t.Chdir(t.TempDir()) t.Setenv("SPINLOOP_ALIAS", "nope") - err := cmdRemoteStop([]string{"--env", "default"}) + err := cmdCloudStop([]string{"--env", "default"}) if err == nil { t.Fatal("expected an error for an unregistered SPINLOOP_ALIAS") } if !strings.Contains(err.Error(), "SPINLOOP_ALIAS") { t.Errorf("error %q does not name the variable", err) } - if strings.Contains(err.Error(), "remote is not configured") { + if strings.Contains(err.Error(), "cloud is not configured") { t.Errorf("the variable was passed over for the default config: %v", err) } } @@ -903,11 +903,11 @@ func TestEnvAlias_RemoteFallsBackWithoutREMOTE(t *testing.T) { t.Chdir(t.TempDir()) t.Setenv("SPINLOOP_ALIAS", "q3") - err := cmdRemoteStop([]string{"--env", "default"}) + err := cmdCloudStop([]string{"--env", "default"}) if err == nil { t.Fatal("expected an error: there is no default endpoint config either") } - if !strings.Contains(err.Error(), "remote is not configured") { + if !strings.Contains(err.Error(), "cloud is not configured") { t.Errorf("error = %q, want the default-config failure (the fallback still applies)", err) } } diff --git a/cmd/spinloop/apply_test.go b/cmd/spinloop/apply_test.go index 0149757b..422b2455 100644 --- a/cmd/spinloop/apply_test.go +++ b/cmd/spinloop/apply_test.go @@ -462,15 +462,15 @@ func TestCmdApply_DirectoryWithoutSpinloop(t *testing.T) { } } -// TestCmdApply_BaseURLFromRemoteConfig checks that a Spinloop with no BASEURL, +// TestCmdApply_BaseURLFromCloudConfig checks that a Spinloop with no BASEURL, // applied against a registered environment, takes the endpoint address from the -// environment's remote.json base_url — the deployment writes that file, so the +// environment's cloud.json base_url — the deployment writes that file, so the // Spinloop does not have to carry it. -func TestCmdApply_BaseURLFromRemoteConfig(t *testing.T) { +func TestCmdApply_BaseURLFromCloudConfig(t *testing.T) { dir := t.TempDir() t.Setenv("XDG_CONFIG_HOME", dir) - envConfig := filepath.Join(dir, "spinloop", "remotes", "dev-1", "remote.json") + envConfig := filepath.Join(dir, "spinloop", "clouds", "dev-1", "cloud.json") if err := os.MkdirAll(filepath.Dir(envConfig), 0o700); err != nil { t.Fatal(err) } @@ -486,23 +486,23 @@ func TestCmdApply_BaseURLFromRemoteConfig(t *testing.T) { } }) if !strings.Contains(out, "http://198.51.100.7:8000/v1") { - t.Errorf("expected the remote base URL to be reported:\n%s", out) + t.Errorf("expected the cloud base URL to be reported:\n%s", out) } m := readConfigMap(t, filepath.Join(dir, "opencode", "opencode.json")) dev1 := m["provider"].(map[string]any)["dev-1"].(map[string]any) if got := dev1["options"].(map[string]any)["baseURL"]; got != "http://198.51.100.7:8000/v1" { - t.Errorf("baseURL = %v, want the remote config's base_url", got) + t.Errorf("baseURL = %v, want the cloud config's base_url", got) } } -// TestCmdApply_SpinloopBaseURLBeatsRemoteConfig checks the precedence: a BASEURL -// the user wrote in the Spinloop wins over the environment's remote.json. -func TestCmdApply_SpinloopBaseURLBeatsRemoteConfig(t *testing.T) { +// TestCmdApply_SpinloopBaseURLBeatsCloudConfig checks the precedence: a BASEURL +// the user wrote in the Spinloop wins over the environment's cloud.json. +func TestCmdApply_SpinloopBaseURLBeatsCloudConfig(t *testing.T) { dir := t.TempDir() t.Setenv("XDG_CONFIG_HOME", dir) - envConfig := filepath.Join(dir, "spinloop", "remotes", "dev-1", "remote.json") + envConfig := filepath.Join(dir, "spinloop", "clouds", "dev-1", "cloud.json") if err := os.MkdirAll(filepath.Dir(envConfig), 0o700); err != nil { t.Fatal(err) } @@ -541,19 +541,19 @@ func TestCmdApply_UnregisteredEnvironmentFails(t *testing.T) { t.Fatal("expected an error for an unregistered environment") } if !strings.Contains(err.Error(), "not registered") || - !strings.Contains(err.Error(), "`spinloop remote deploy --env \"dev-1\"`") { + !strings.Contains(err.Error(), "`spinloop cloud deploy --env \"dev-1\"`") { t.Errorf("error = %v, want the not-registered error naming the deploy command", err) } } // TestCmdApply_RemoteConfigMalformed checks that applying against a registered -// environment whose remote.json is malformed fails loudly, rather than silently +// environment whose cloud.json is malformed fails loudly, rather than silently // applying under the wrong base URL. func TestCmdApply_RemoteConfigMalformed(t *testing.T) { dir := t.TempDir() t.Setenv("XDG_CONFIG_HOME", dir) - envConfig := filepath.Join(dir, "spinloop", "remotes", "dev-1", "remote.json") + envConfig := filepath.Join(dir, "spinloop", "clouds", "dev-1", "cloud.json") if err := os.MkdirAll(filepath.Dir(envConfig), 0o700); err != nil { t.Fatal(err) } @@ -564,22 +564,22 @@ func TestCmdApply_RemoteConfigMalformed(t *testing.T) { err := cmdApply([]string{"--env", "dev-1", spinloopFile}) if err == nil { - t.Fatal("expected an error for a malformed remote config") + t.Fatal("expected an error for a malformed cloud config") } if !strings.Contains(err.Error(), "parsing") { - t.Errorf("error = %v, want it to mention parsing the remote config", err) + t.Errorf("error = %v, want it to mention parsing the cloud config", err) } } // TestCmdApply_RemoteNameIsProviderName checks that --env keys the harness // provider on the environment name — configured from the PROVIDER's catalogue // entry — with the default model reading as / and the base URL taken -// from that environment's remote.json. +// from that environment's cloud.json. func TestCmdApply_RemoteNameIsProviderName(t *testing.T) { dir := t.TempDir() t.Setenv("XDG_CONFIG_HOME", dir) - envConfig := filepath.Join(dir, "spinloop", "remotes", "dev-1", "remote.json") + envConfig := filepath.Join(dir, "spinloop", "clouds", "dev-1", "cloud.json") if err := os.MkdirAll(filepath.Dir(envConfig), 0o700); err != nil { t.Fatal(err) } @@ -616,14 +616,14 @@ func TestCmdApply_RemoteNameIsProviderName(t *testing.T) { } } -// TestCmdApply_RemoteProviderLabelledPerEnvironment checks that a remote provider +// TestCmdApply_RemoteProviderLabelledPerEnvironment checks that a cloud provider // gets a display name qualified by its environment, so it reads distinctly from a // local engine of the same kind, which keeps the bare engine name. func TestCmdApply_RemoteProviderLabelledPerEnvironment(t *testing.T) { dir := t.TempDir() t.Setenv("XDG_CONFIG_HOME", dir) - envConfig := filepath.Join(dir, "spinloop", "remotes", "dev-2", "remote.json") + envConfig := filepath.Join(dir, "spinloop", "clouds", "dev-2", "cloud.json") if err := os.MkdirAll(filepath.Dir(envConfig), 0o700); err != nil { t.Fatal(err) } @@ -631,14 +631,14 @@ func TestCmdApply_RemoteProviderLabelledPerEnvironment(t *testing.T) { `{"start_url":"https://start.example/","stop_url":"https://stop.example/","region":"us-east-1","base_url":"http://198.51.100.7:8000/v1","environment":"dev-2"}`) spinloopDir := t.TempDir() - remoteSpinloop := filepath.Join(spinloopDir, "Spinloop") - mustWrite(t, remoteSpinloop, "PROVIDER llamacpp\nALIAS qwen\n") + cloudSpinloop := filepath.Join(spinloopDir, "Spinloop") + mustWrite(t, cloudSpinloop, "PROVIDER llamacpp\nALIAS qwen\n") localSpinloop := filepath.Join(spinloopDir, "Local") mustWrite(t, localSpinloop, "PROVIDER llamacpp\nALIAS qwen\nBASEURL http://127.0.0.1:8080/v1\n") captureStdout(t, func() { - if err := cmdApply([]string{"--env", "dev-2", remoteSpinloop}); err != nil { - t.Fatalf("cmdApply remote: %v", err) + if err := cmdApply([]string{"--env", "dev-2", cloudSpinloop}); err != nil { + t.Fatalf("cmdApply cloud: %v", err) } if err := cmdApply([]string{localSpinloop}); err != nil { t.Fatalf("cmdApply local: %v", err) @@ -651,7 +651,7 @@ func TestCmdApply_RemoteProviderLabelledPerEnvironment(t *testing.T) { t.Fatalf("expected a provider keyed %q, got %v", "dev-2", prov) } if got := dev2["name"]; got != "llama.cpp (dev-2)" { - t.Errorf("remote display name = %v, want %q", got, "llama.cpp (dev-2)") + t.Errorf("cloud display name = %v, want %q", got, "llama.cpp (dev-2)") } local, ok := prov["llamacpp"].(map[string]any) if !ok { @@ -661,7 +661,7 @@ func TestCmdApply_RemoteProviderLabelledPerEnvironment(t *testing.T) { t.Errorf("local display name = %v, want the bare engine name %q", got, "llama.cpp") } if dev2["name"] == local["name"] { - t.Errorf("remote and local providers share a display name %v; they must be distinct", dev2["name"]) + t.Errorf("cloud and local providers share a display name %v; they must be distinct", dev2["name"]) } } @@ -674,7 +674,7 @@ func TestCmdApply_RemoteReapplyRefreshesLabel(t *testing.T) { dir := t.TempDir() t.Setenv("XDG_CONFIG_HOME", dir) - envConfig := filepath.Join(dir, "spinloop", "remotes", "dev-2", "remote.json") + envConfig := filepath.Join(dir, "spinloop", "clouds", "dev-2", "cloud.json") if err := os.MkdirAll(filepath.Dir(envConfig), 0o700); err != nil { t.Fatal(err) } @@ -709,13 +709,13 @@ func TestCmdApply_RemoteReapplyRefreshesLabel(t *testing.T) { } // TestCmdUnapply_RemoveEnvironmentNamedProvider checks apply/unapply symmetry for -// a remote Spinloop: unapply removes the environment-named provider that apply +// a cloud Spinloop: unapply removes the environment-named provider that apply // wrote, not the PROVIDER-named one (which was never written). func TestCmdUnapply_RemoveEnvironmentNamedProvider(t *testing.T) { dir := t.TempDir() t.Setenv("XDG_CONFIG_HOME", dir) - envConfig := filepath.Join(dir, "spinloop", "remotes", "dev-1", "remote.json") + envConfig := filepath.Join(dir, "spinloop", "clouds", "dev-1", "cloud.json") if err := os.MkdirAll(filepath.Dir(envConfig), 0o700); err != nil { t.Fatal(err) } diff --git a/cmd/spinloop/cli_viper.go b/cmd/spinloop/cli_viper.go index 98a20d8b..0611c127 100644 --- a/cmd/spinloop/cli_viper.go +++ b/cmd/spinloop/cli_viper.go @@ -1,11 +1,11 @@ // The CLI layer's single Viper instance: the one place that reads the SPINLOOP_* -// environment values the CLI owns (SPINLOOP_ALIAS and the SPINLOOP_REMOTE_* +// environment values the CLI owns (SPINLOOP_ALIAS and the SPINLOOP_CLOUD_* // control-plane settings). Every SPINLOOP_* variable whose precedence is an // internal contract is deliberately NOT read here — SPINLOOP_PROVIDERS // (catalog), SPINLOOP_BASE_URL (the injected resolve closures), SPINLOOP_LOG_LEVEL // (daemon.ParseLevel), SPINLOOP_HARNESS (harness.Resolve, which also reports the // source of the choice), and the domain-owned SPINLOOP_API_TOKEN / -// SPINLOOP_CONFIG_DIR / SPINLOOP_REMOTE_* of other packages. Reading one of those +// SPINLOOP_CONFIG_DIR / SPINLOOP_CLOUD_* of other packages. Reading one of those // through Viper would create a second reader of the same variable — the exact // silent drift this migration removes. See the ownership table in // openspec/changes/migrate-cli-to-cobra-viper/design.md (D3). @@ -21,7 +21,7 @@ import ( // cliViper is built once per process. AutomaticEnv plus the SPINLOOP prefix means // cliViper.GetString("alias") resolves SPINLOOP_ALIAS and -// cliViper.GetString("remote_start_url") resolves SPINLOOP_REMOTE_START_URL, and +// cliViper.GetString("cloud_start_url") resolves SPINLOOP_CLOUD_START_URL, and // a dash in a key is read from the underscored variable. var cliViper = newCLIViper() @@ -34,8 +34,8 @@ func newCLIViper() *viper.Viper { } // viperGetenv returns a Viper-backed os.Getenv for the SPINLOOP_* variables the -// CLI owns. internal/remote keeps its func(string) string injection point for -// the SPINLOOP_REMOTE_* config keys; the CLI hands it this closure so the lookup +// CLI owns. internal/cloud keeps its func(string) string injection point for +// the SPINLOOP_CLOUD_* config keys; the CLI hands it this closure so the lookup // runs through cliViper instead of a raw os.Getenv. A name without the SPINLOOP_ // prefix is not a CLI-owned variable, so it falls through to the process // environment unchanged, which keeps the closure safe as a general getenv. @@ -64,7 +64,7 @@ func viperGetenv() func(string) string { // at call time rather than capturing it, so its closures follow the swap. // // In the current surface no CLI-owned variable has a flag spelling — -// SPINLOOP_ALIAS and the SPINLOOP_REMOTE_* settings have none — so every binding +// SPINLOOP_ALIAS and the SPINLOOP_CLOUD_* settings have none — so every binding // today is a no-op that exists so the precedence stays one mechanism if a flag // ever names one of them. The flags that have env counterparts today // (--providers, --base-url, --log-level, --harness) are resolved by the diff --git a/cmd/spinloop/remote.go b/cmd/spinloop/cloud.go similarity index 85% rename from cmd/spinloop/remote.go rename to cmd/spinloop/cloud.go index 3ae93588..03cf656c 100644 --- a/cmd/spinloop/remote.go +++ b/cmd/spinloop/cloud.go @@ -17,17 +17,17 @@ import ( "time" "github.com/spf13/cobra" + "github.com/spinloop-ai/spinloop/internal/cloud" "github.com/spinloop-ai/spinloop/internal/contextsize" "github.com/spinloop-ai/spinloop/internal/fleet" "github.com/spinloop-ai/spinloop/internal/inference" "github.com/spinloop-ai/spinloop/internal/opencode" "github.com/spinloop-ai/spinloop/internal/preset" - "github.com/spinloop-ai/spinloop/internal/remote" "github.com/spinloop-ai/spinloop/internal/spinloop" "github.com/spinloop-ai/spinloop/internal/spinloopsrc" ) -// cmdRemote dispatches the remote subcommands, which control the +// cmdCloud dispatches the cloud subcommands, which control the // scale-to-zero GPU inference instance defined in this repo's remote/: // start boots it and prints the endpoint exports, pause stops it without // terminating it (the sweep terminates stopped instances after their @@ -38,19 +38,19 @@ import ( // early, and deploy sets what the instance will serve from the Spinloop itself. // Each subcommand acts on one registered environment — the --env flag names // it, and without the flag the `default` environment is used; see -// resolveRemoteConfig. Each also takes an optional Spinloop path, read for +// resolveCloudConfig. Each also takes an optional Spinloop path, read for // its ENV instructions and adjacent .env only. The group itself does nothing // — see groupFallback for what a bare or mistyped invocation gets. -// cmdRemote runs the remote subcommands through the tree — the seam the suite +// cmdCloud runs the cloud subcommands through the tree — the seam the suite // calls directly. -func cmdRemote(args []string) error { - return execCmd(remoteCmd(), args) +func cmdCloud(args []string) error { + return execCmd(cloudCmd(), args) } -// applySpinloopEnv makes the remote commands respect the Spinloop's local +// applySpinloopEnv makes the cloud commands respect the Spinloop's local // environment. The AWS SDK's credential chain reads the process environment -// directly, and the SPINLOOP_REMOTE_*/AWS_REGION lookups (Viper's AutomaticEnv +// directly, and the SPINLOOP_CLOUD_*/AWS_REGION lookups (Viper's AutomaticEnv // included) ultimately do the same, so the values have to be present in the // environment itself — a lookup closure would not reach the SDK. It therefore // mutates this process's environment, in two passes that give the precedence @@ -89,50 +89,50 @@ func applySpinloopEnv(sel spinloop.Selection, spinloopPath string) error { return nil } -// envFlagUsage is the --env flag's help text on every remote subcommand that +// envFlagUsage is the --env flag's help text on every cloud subcommand that // acts on one environment. const envFlagUsage = "the registered environment to act on (required)" -// resolveRemoteConfig loads the remote config a subcommand acts on. The --env +// resolveCloudConfig loads the cloud config a subcommand acts on. The --env // flag names an environment in the per-user registry, read from -// remotes//remote.json; a name with no registered configuration fails +// clouds//cloud.json; a name with no registered configuration fails // saying the environment is not registered and how to create it. Without the // flag, the per-user config is the fallback — the `default` environment, then -// the legacy single file, then the SPINLOOP_REMOTE_* overrides alone — so -// `spinloop remote` still works outside any project. +// the legacy single file, then the SPINLOOP_CLOUD_* overrides alone — so +// `spinloop cloud` still works outside any project. // // A Spinloop given as an argument (or the ./Spinloop in the working directory, // when no argument was given) is read for its ENV instructions and adjacent // .env only, applied before any control-plane work: it never selects an // environment. -func resolveRemoteConfig(envName, spinloopArg string) (remote.Config, error) { +func resolveCloudConfig(envName, spinloopArg string) (cloud.Config, error) { if spinloopArg != "" { - sel, spinloopPath, err := readSpinloop("remote", spinloopArg) + sel, spinloopPath, err := readSpinloop("cloud", spinloopArg) if err != nil { - return remote.Config{}, err + return cloud.Config{}, err } if err := applySpinloopEnv(sel, spinloopPath); err != nil { - return remote.Config{}, err + return cloud.Config{}, err } } else if defaultSpinloopNamed() { - sel, spinloopPath, err := readSpinloop("remote", "") + sel, spinloopPath, err := readSpinloop("cloud", "") if err != nil { - return remote.Config{}, err + return cloud.Config{}, err } if err := applySpinloopEnv(sel, spinloopPath); err != nil { - return remote.Config{}, err + return cloud.Config{}, err } } if envName == "" { - return remote.Config{}, errNoEnvironment() + return cloud.Config{}, errNoEnvironment() } - if !remote.IsEnvName(envName) { - return remote.Config{}, fmt.Errorf("%q is not an environment name: an environment name is a plain identifier, with no path", envName) + if !cloud.IsEnvName(envName) { + return cloud.Config{}, fmt.Errorf("%q is not an environment name: an environment name is a plain identifier, with no path", envName) } - return remote.LoadEnvironment(envName, viperGetenv()) + return cloud.LoadEnvironment(envName, viperGetenv()) } -// errNoEnvironment is what a remote subcommand fails with when it names no +// errNoEnvironment is what a cloud subcommand fails with when it names no // environment. There is no environment to fall back to: several of these // subcommands change the state of a cloud instance, and one nobody named is // not one to start, stop or terminate. The registered names are listed, so the @@ -140,9 +140,9 @@ func resolveRemoteConfig(envName, spinloopArg string) (remote.Config, error) { func errNoEnvironment() error { var b strings.Builder b.WriteString("no environment named: pass --env ") - envs, err := remote.ListEnvironments() + envs, err := cloud.ListEnvironments() if err != nil || len(envs) == 0 { - b.WriteString(" (none registered yet — `spinloop remote deploy --env ` creates one)") + b.WriteString(" (none registered yet — `spinloop cloud deploy --env ` creates one)") return errors.New(b.String()) } names := make([]string, len(envs)) @@ -184,7 +184,7 @@ func defaultSpinloopExists() bool { } // spinloopArg returns the optional positional Spinloop path after the flags. -// spinloopArg is the first positional argument of a remote subcommand — the +// spinloopArg is the first positional argument of a cloud subcommand — the // Spinloop path — or "" when none was given. func spinloopArg(args []string) string { if len(args) > 0 { @@ -199,10 +199,10 @@ func spinloopArg(args []string) string { var heartbeatEvery = 30 * time.Second // startProgress reports what a slow start is doing, as a renderer over the -// phases fleet.StartPhases builds from remote.Start's callbacks — the same +// phases fleet.StartPhases builds from cloud.Start's callbacks — the same // phases the dashboard tile draws, so the two surfaces cannot word one // situation differently. Everything it writes goes to stderr, so `spinloop -// remote start | grep '^export '` still yields just the exports while the user +// cloud start | grep '^export '` still yields just the exports while the user // watching the terminal still sees progress. type startProgress struct { mu sync.Mutex @@ -230,7 +230,7 @@ func newStartProgress(every time.Duration) *startProgress { } // report prints one phase as the start enters it, and keeps it for the -// heartbeat to redraw. Called from remote.Start's own goroutine. +// heartbeat to redraw. Called from cloud.Start's own goroutine. func (p *startProgress) report(phase fleet.StartPhase) { p.mu.Lock() p.phase = phase @@ -238,7 +238,7 @@ func (p *startProgress) report(phase fleet.StartPhase) { p.line(fleet.RenderPhase(phase, time.Now())) } -// callbacks are the pair to hand remote.Start: the phases they build are what +// callbacks are the pair to hand cloud.Start: the phases they build are what // report renders. func (p *startProgress) callbacks() (progress func(string), onState func(string)) { return fleet.StartPhases(p.report) @@ -272,24 +272,24 @@ func (p *startProgress) close() { p.stop.Do(func() { close(p.done) }) } -// The variables a remote endpoint is addressed by. The instance is started +// The variables a cloud endpoint is addressed by. The instance is started // with its API key as --api-key on an OpenAI-compatible server, so these are // the names every consumer uses: the export lines below, and the environment // `spinloop harness open` hands the agent it launches. const ( - remoteBaseURLEnv = "OPENAI_BASE_URL" - remoteAPIKeyEnv = "OPENAI_API_KEY" + cloudBaseURLEnv = "OPENAI_BASE_URL" + cloudAPIKeyEnv = "OPENAI_API_KEY" ) -// printRemoteEnv prints the remote endpoint's environment variables as shell +// printCloudEnv prints the cloud endpoint's environment variables as shell // export lines to stdout, suitable for eval. Nothing else on this path may // write to stdout, or the eval fails on the stray line. -func printRemoteEnv(resp *remote.Response) { - fmt.Printf("export %s=%s\n", remoteBaseURLEnv, resp.BaseURL) - fmt.Printf("export %s=%s\n", remoteAPIKeyEnv, resp.APIKey) +func printCloudEnv(resp *cloud.Response) { + fmt.Printf("export %s=%s\n", cloudBaseURLEnv, resp.BaseURL) + fmt.Printf("export %s=%s\n", cloudAPIKeyEnv, resp.APIKey) } -func remoteEnvCmd() *cobra.Command { +func cloudEnvCmd() *cobra.Command { var envName string c := &cobra.Command{ Use: "env", @@ -302,7 +302,7 @@ starting it.`, ValidArgsFunction: aliasSlot, RunE: func(c *cobra.Command, args []string) error { resolve(c) - return runRemoteEnv(envName, args) + return runCloudEnv(envName, args) }, } c.Flags().StringVar(&envName, "env", "", envFlagUsage) @@ -310,21 +310,21 @@ starting it.`, return c } -// runRemoteEnv is the body of `spinloop remote env`. -func runRemoteEnv(envName string, args []string) error { - cfg, err := resolveRemoteConfig(envName, spinloopArg(args)) +// runCloudEnv is the body of `spinloop cloud env`. +func runCloudEnv(envName string, args []string) error { + cfg, err := resolveCloudConfig(envName, spinloopArg(args)) if err != nil { return err } - resp, err := remote.Env(context.Background(), cfg) + resp, err := cloud.Env(context.Background(), cfg) if err != nil { return err } - printRemoteEnv(resp) + printCloudEnv(resp) return nil } -func remoteStartCmd() *cobra.Command { +func cloudStartCmd() *cobra.Command { var timeout time.Duration const timeoutUsage = "overall time to wait for the endpoint" var envName string @@ -341,7 +341,7 @@ agent needs.`, ValidArgsFunction: aliasSlot, RunE: func(c *cobra.Command, args []string) error { resolve(c) - return runRemoteStart(envName, args, timeout, printEnv, keepD) + return runCloudStart(envName, args, timeout, printEnv, keepD) }, } fs := c.Flags() @@ -353,9 +353,9 @@ agent needs.`, return c } -// runRemoteStart is the body of `spinloop remote start`. -func runRemoteStart(envName string, args []string, timeout time.Duration, printEnv bool, keepD string) error { - cfg, err := resolveRemoteConfig(envName, spinloopArg(args)) +// runCloudStart is the body of `spinloop cloud start`. +func runCloudStart(envName string, args []string, timeout time.Duration, printEnv bool, keepD string) error { + cfg, err := resolveCloudConfig(envName, spinloopArg(args)) if err != nil { return err } @@ -381,7 +381,7 @@ func runRemoteStart(envName string, args []string, timeout time.Duration, printE ctx, cancel := context.WithTimeout(context.Background(), timeout) defer cancel() onProgress, onState := progress.callbacks() - resp, err := remote.Start(ctx, cfg, onProgress, onState, retainUntil) + resp, err := cloud.Start(ctx, cfg, onProgress, onState, retainUntil) if err != nil { return err } @@ -392,13 +392,13 @@ func runRemoteStart(envName string, args []string, timeout time.Duration, printE // reach it — the control plane (SigV4 Lambda URLs) works from anywhere, // but the inference port is guarded by a security group that admits only // one CIDR. A changed network means start succeeds but inference hangs. - if err := remote.ProbeReachability(resp.BaseURL); err != nil { + if err := cloud.ProbeReachability(resp.BaseURL); err != nil { cidr := "/32" if detected, detErr := detectPublicCIDRFn(context.Background()); detErr == nil { cidr = detected } fmt.Fprintln(os.Stderr, "the endpoint is ready but not reachable from this network — its ingress admits a different address.") - fmt.Fprintf(os.Stderr, "Re-admit this machine with:\n spinloop remote deploy --overwrite --allowed-cidr %s\n", cidr) + fmt.Fprintf(os.Stderr, "Re-admit this machine with:\n spinloop cloud deploy --overwrite --allowed-cidr %s\n", cidr) } if retainUntil != nil || resp.RetainUntil != "" { @@ -411,21 +411,21 @@ func runRemoteStart(envName string, args []string, timeout time.Duration, printE if printEnv { envCtx, envCancel := context.WithTimeout(context.Background(), 30*time.Second) - envResp, err := remote.Env(envCtx, cfg) + envResp, err := cloud.Env(envCtx, cfg) envCancel() if err != nil { return err } - printRemoteEnv(envResp) + printCloudEnv(envResp) } return nil } -// cmdRemoteList prints the registered remote environments, each with its base -// URL and region, marking any whose remote.json is missing or unreadable. It +// cmdCloudList prints the registered cloud environments, each with its base +// URL and region, marking any whose cloud.json is missing or unreadable. It // contacts no endpoint. Environments are registered under -// ~/.config/spinloop/remotes// (by `spinloop remote bootstrap`, or by hand). -func remoteListCmd() *cobra.Command { +// ~/.config/spinloop/clouds// (by `spinloop cloud bootstrap`, or by hand). +func cloudListCmd() *cobra.Command { return &cobra.Command{ Use: "ls", Short: "list the registered environments", @@ -435,24 +435,24 @@ func remoteListCmd() *cobra.Command { ValidArgsFunction: noPositionals, RunE: func(c *cobra.Command, _ []string) error { resolve(c) - return runRemoteList() + return runCloudList() }, } } -// runRemoteList is the body of `spinloop remote ls`. -func runRemoteList() error { - envs, err := remote.ListEnvironments() +// runCloudList is the body of `spinloop cloud ls`. +func runCloudList() error { + envs, err := cloud.ListEnvironments() if err != nil { return err } if len(envs) == 0 { - fmt.Println("No remote environments registered. Register one with `spinloop remote bootstrap`.") + fmt.Println("No cloud environments registered. Register one with `spinloop cloud bootstrap`.") return nil } for _, e := range envs { if !e.OK { - fmt.Printf("%s\t(missing or unreadable remote.json)\n", e.Name) + fmt.Printf("%s\t(missing or unreadable cloud.json)\n", e.Name) continue } base := e.BaseURL @@ -464,7 +464,7 @@ func runRemoteList() error { return nil } -func remoteKeepCmd() *cobra.Command { +func cloudKeepCmd() *cobra.Command { var envName string c := &cobra.Command{ Use: "keep", @@ -475,7 +475,7 @@ func remoteKeepCmd() *cobra.Command { ValidArgsFunction: keepSlot, RunE: func(c *cobra.Command, rest []string) error { resolve(c) - return runRemoteKeep(envName, rest) + return runCloudKeep(envName, rest) }, } c.Flags().StringVar(&envName, "env", "", envFlagUsage) @@ -483,10 +483,10 @@ func remoteKeepCmd() *cobra.Command { return c } -// runRemoteKeep is the body of `spinloop remote keep`. -func runRemoteKeep(envName string, rest []string) error { +// runCloudKeep is the body of `spinloop cloud keep`. +func runCloudKeep(envName string, rest []string) error { if len(rest) == 0 { - return fmt.Errorf("usage: spinloop remote keep [path]") + return fmt.Errorf("usage: spinloop cloud keep [path]") } durationStr := rest[0] d, err := time.ParseDuration(durationStr) @@ -501,12 +501,12 @@ func runRemoteKeep(envName string, rest []string) error { if len(rest) > 1 { spinloopPath = rest[1] } - cfg, err := resolveRemoteConfig(envName, spinloopPath) + cfg, err := resolveCloudConfig(envName, spinloopPath) if err != nil { return err } retainUntil := time.Now().Add(d) - _, err = remote.Keep(context.Background(), cfg, retainUntil) + _, err = cloud.Keep(context.Background(), cfg, retainUntil) if err != nil { return err } @@ -514,7 +514,7 @@ func runRemoteKeep(envName string, rest []string) error { return nil } -func remotePauseCmd() *cobra.Command { +func cloudPauseCmd() *cobra.Command { var envName string c := &cobra.Command{ Use: "pause", @@ -525,7 +525,7 @@ func remotePauseCmd() *cobra.Command { ValidArgsFunction: aliasSlot, RunE: func(c *cobra.Command, args []string) error { resolve(c) - return runRemotePause(envName, args) + return runCloudPause(envName, args) }, } c.Flags().StringVar(&envName, "env", "", envFlagUsage) @@ -533,13 +533,13 @@ func remotePauseCmd() *cobra.Command { return c } -// runRemotePause is the body of `spinloop remote pause`. -func runRemotePause(envName string, args []string) error { - cfg, err := resolveRemoteConfig(envName, spinloopArg(args)) +// runCloudPause is the body of `spinloop cloud pause`. +func runCloudPause(envName string, args []string) error { + cfg, err := resolveCloudConfig(envName, spinloopArg(args)) if err != nil { return err } - resp, err := remote.Pause(context.Background(), cfg, false) + resp, err := cloud.Pause(context.Background(), cfg, false) if err != nil { return err } @@ -547,11 +547,11 @@ func runRemotePause(envName string, args []string) error { // was); either way the instance is re-wakeable, and the control plane's // sweep terminates it once the stop retention passes. fmt.Printf("state: %s\n", resp.State) - fmt.Println("the endpoint can be re-woken with `spinloop remote start`, or terminated now with `spinloop remote stop`") + fmt.Println("the endpoint can be re-woken with `spinloop cloud start`, or terminated now with `spinloop cloud stop`") return nil } -func remoteRestartCmd() *cobra.Command { +func cloudRestartCmd() *cobra.Command { var force bool var timeout time.Duration const timeoutUsage = "overall time to wait for the endpoint" @@ -569,7 +569,7 @@ skipped: for when the engine or its daemon will not answer it.`, ValidArgsFunction: aliasSlot, RunE: func(c *cobra.Command, args []string) error { resolve(c) - return runRemoteRestart(envName, args, force, timeout) + return runCloudRestart(envName, args, force, timeout) }, } fs := c.Flags() @@ -580,12 +580,12 @@ skipped: for when the engine or its daemon will not answer it.`, return c } -// runRemoteRestart is the body of `spinloop remote restart`. It stops the +// runCloudRestart is the body of `spinloop cloud restart`. It stops the // instance in the pause manner — without terminating it, so the boot disk and // weights survive and the address does not change — and reuses the wake's own // deadline and retry handling to block until the model serves again. -func runRemoteRestart(envName string, args []string, force bool, timeout time.Duration) error { - cfg, err := resolveRemoteConfig(envName, spinloopArg(args)) +func runCloudRestart(envName string, args []string, force bool, timeout time.Duration) error { + cfg, err := resolveCloudConfig(envName, spinloopArg(args)) if err != nil { return err } @@ -597,7 +597,7 @@ func runRemoteRestart(envName string, args []string, force bool, timeout time.Du // saying where the instance is now (running, or already stopped / no // instance). It gates nothing — the stop Lambda is correct for every // state — so a failed check just means we do not know the starting point. - if status, err := remote.Status(context.Background(), cfg); err == nil { + if status, err := cloud.Status(context.Background(), cfg); err == nil { switch status.State { case "stopped", "undeployed": progress.line(fmt.Sprintf("the instance is already %s; waking it", status.State)) @@ -615,7 +615,7 @@ func runRemoteRestart(envName string, args []string, force bool, timeout time.Du ctx, cancel := context.WithTimeout(context.Background(), timeout) defer cancel() onProgress, onState := progress.callbacks() - resp, err := remote.Restart(ctx, cfg, force, onProgress, onState) + resp, err := cloud.Restart(ctx, cfg, force, onProgress, onState) if err != nil { return err } @@ -625,7 +625,7 @@ func runRemoteRestart(envName string, args []string, force bool, timeout time.Du return nil } -func remoteStopCmd() *cobra.Command { +func cloudStopCmd() *cobra.Command { var envName string c := &cobra.Command{ Use: "stop", @@ -637,7 +637,7 @@ func remoteStopCmd() *cobra.Command { ValidArgsFunction: aliasSlot, RunE: func(c *cobra.Command, args []string) error { resolve(c) - return runRemoteStop(envName, args) + return runCloudStop(envName, args) }, } c.Flags().StringVar(&envName, "env", "", envFlagUsage) @@ -645,13 +645,13 @@ func remoteStopCmd() *cobra.Command { return c } -// runRemoteStop is the body of `spinloop remote stop`. -func runRemoteStop(envName string, args []string) error { - cfg, err := resolveRemoteConfig(envName, spinloopArg(args)) +// runCloudStop is the body of `spinloop cloud stop`. +func runCloudStop(envName string, args []string) error { + cfg, err := resolveCloudConfig(envName, spinloopArg(args)) if err != nil { return err } - resp, err := remote.Stop(context.Background(), cfg) + resp, err := cloud.Stop(context.Background(), cfg) if err != nil { return err } @@ -693,12 +693,12 @@ func formatBytes(b int64) string { // from the AWS Price List API. Uses GetProducts with a filter on instance type // and operation (Linux/Windows). Returns the price per hour. func getOnDemandPrice(ctx context.Context, region, instanceType string) (float64, error) { - return remote.GetOnDemandPrice(ctx, region, instanceType) + return cloud.GetOnDemandPrice(ctx, region, instanceType) } // runnerFor maps a Spinloop's PROVIDER to the inference runner the cloud should // run. PROVIDER already names the engine — `spinloop serve` starts that engine -// locally, `spinloop remote deploy` asks the cloud for the same one — so no +// locally, `spinloop cloud deploy` asks the cloud for the same one — so no // separate keyword is needed. Providers that are not self-hosted engines have // nothing to deploy. func runnerFor(provider string) (string, error) { @@ -707,7 +707,7 @@ func runnerFor(provider string) (string, error) { return provider, nil default: return "", fmt.Errorf( - "PROVIDER %q cannot be deployed: remote deploy runs a self-hosted engine, so use llamacpp or vllm", + "PROVIDER %q cannot be deployed: cloud deploy runs a self-hosted engine, so use llamacpp or vllm", provider) } } @@ -1086,9 +1086,9 @@ func presetValue(key string, layers ...[]preset.Param) string { // Seams for the deploy flow, so tests drive it without AWS or a network. var ( - deployDiscoverFn = remote.DiscoverControlPlane - remoteDeployFn = remote.Deploy - remoteStatusFn = remote.Status + deployDiscoverFn = cloud.DiscoverControlPlane + cloudDeployFn = cloud.Deploy + cloudStatusFn = cloud.Status detectPublicCIDRFn = detectPublicCIDR ) @@ -1096,7 +1096,7 @@ var ( // Tests override this to avoid sleeping 60 seconds. var metricsWatchInterval = 60 * time.Second -func remoteDeployCmd() *cobra.Command { +func cloudDeployCmd() *cobra.Command { var ( dryRun bool overwrite bool @@ -1125,7 +1125,7 @@ from its next fresh launch.`, ValidArgsFunction: aliasSlot, RunE: func(c *cobra.Command, args []string) error { resolve(c) - return runRemoteDeploy(args, envName, dryRun, overwrite, reseed, allowedCidr, region, spinloopVersion, instanceType, apiKeyEnv) + return runCloudDeploy(args, envName, dryRun, overwrite, reseed, allowedCidr, region, spinloopVersion, instanceType, apiKeyEnv) }, } fs := c.Flags() @@ -1142,9 +1142,9 @@ from its next fresh launch.`, return c } -// runRemoteDeploy is the body of `spinloop remote deploy`. -func runRemoteDeploy(args []string, envName string, dryRun, overwrite, reseed bool, allowedCidr, region, spinloopVersion, instanceType, apiKeyEnv string) error { - _, spinloopPath, dc, err := deriveDeployTarget("spinloop remote deploy ", spinloopArg(args), envName) +// runCloudDeploy is the body of `spinloop cloud deploy`. +func runCloudDeploy(args []string, envName string, dryRun, overwrite, reseed bool, allowedCidr, region, spinloopVersion, instanceType, apiKeyEnv string) error { + _, spinloopPath, dc, err := deriveDeployTarget("spinloop cloud deploy ", spinloopArg(args), envName) if err != nil { return err } @@ -1169,13 +1169,13 @@ func runRemoteDeploy(args []string, envName string, dryRun, overwrite, reseed bo // path, or a URL — whatever readSpinloop accepts) and the environment name // the caller supplies into everything a deploy needs: the Spinloop it read, // the deploy config it derives, and the name to register under. usage is -// readSpinloop's error-message context, so a caller other than `remote deploy` +// readSpinloop's error-message context, so a caller other than `cloud deploy` // (namely `fleet deploy`, which names the environment with the node's own // name) gets a message naming itself rather than a hard-coded command line. // // The Spinloop says what the environment serves; the name arrives from -// outside it — the --env flag for `remote deploy`, the node name for `fleet -// deploy` — so a node's resolved Spinloop source and a standalone `remote +// outside it — the --env flag for `cloud deploy`, the node name for `fleet +// deploy` — so a node's resolved Spinloop source and a standalone `cloud // deploy` of the same file agree about what they deploy and each names its // own environment. func deriveDeployTarget(usage, spinloopArg, env string) (sel spinloop.Selection, spinloopPath string, dc inference.DeployConfig, err error) { @@ -1185,7 +1185,7 @@ func deriveDeployTarget(usage, spinloopArg, env string) (sel spinloop.Selection, } // Respect the Spinloop's local environment (.env beside it, then its ENV // lines) before any AWS work, so the credentials the deploy signs with, the - // region, and the SPINLOOP_REMOTE_* overrides all see it. ENV stays local — it + // region, and the SPINLOOP_CLOUD_* overrides all see it. ENV stays local — it // never enters dc, so nothing here reaches the deployed instance. if err = applySpinloopEnv(sel, spinloopPath); err != nil { return @@ -1198,7 +1198,7 @@ func deriveDeployTarget(usage, spinloopArg, env string) (sel spinloop.Selection, err = fmt.Errorf("deploy must name the environment it creates: pass --env (e.g. --env %s)", dc.ServedModelName) return } - if !remote.IsEnvName(env) { + if !cloud.IsEnvName(env) { err = fmt.Errorf("%q is not an environment name: an environment name is a plain identifier, with no path", env) return } @@ -1206,7 +1206,7 @@ func deriveDeployTarget(usage, spinloopArg, env string) (sel spinloop.Selection, } // deployOpts are the flags a deploy takes, independent of the Spinloop file -// itself — shared by `remote deploy` and `fleet deploy`, which each collect +// itself — shared by `cloud deploy` and `fleet deploy`, which each collect // them from their own flag set. type deployOpts struct { dryRun bool @@ -1228,7 +1228,7 @@ type deployOpts struct { } // deployOutcome is a successful (or dry-run) deploy's result: the plan/result -// text a standalone `remote deploy` prints verbatim, plus the values a +// text a standalone `cloud deploy` prints verbatim, plus the values a // caller driving several nodes wants without re-parsing that text. A failed // or guarded deploy is reported through the error runDeploy returns instead // — an errDeployGuarded distinguishes "needs --overwrite" from any other @@ -1280,7 +1280,7 @@ func runDeploy(spinloopPath, env string, dc inference.DeployConfig, opts deployO // that machine. Checked here, so a typo is named now rather than as a // RunInstances failure inside a deploy nobody is watching. if name := strings.TrimSpace(opts.instanceType); name != "" { - if !remote.IsInstanceType(name) { + if !cloud.IsInstanceType(name) { return deployOutcome{}, fmt.Errorf( "--instance-type must be an EC2 instance type (a family and size separated by a dot, e.g. g6e.xlarge), got %q", opts.instanceType) @@ -1356,7 +1356,7 @@ func runDeploy(spinloopPath, env string, dc inference.DeployConfig, opts deployO // The control URLs come from the control plane's stack outputs — the // environment may not exist yet, so there is nothing local to resolve. ctx := context.Background() - awsCfg, err := remote.LoadAWSConfig(ctx, resolveRegion(opts.region)) + awsCfg, err := cloud.LoadAWSConfig(ctx, resolveRegion(opts.region)) if err != nil { return deployOutcome{}, err } @@ -1369,7 +1369,7 @@ func runDeploy(spinloopPath, env string, dc inference.DeployConfig, opts deployO // Refuse to clobber silently: an environment that is already registered, or // whose instance is live, needs explicit consent to redeploy over. - envConfigPath, err := remote.EnvConfigPath(env) + envConfigPath, err := cloud.EnvConfigPath(env) if err != nil { return deployOutcome{}, err } @@ -1378,7 +1378,7 @@ func runDeploy(spinloopPath, env string, dc inference.DeployConfig, opts deployO registered = true } live := false - if status, err := remoteStatusFn(ctx, cfg); err == nil { + if status, err := cloudStatusFn(ctx, cfg); err == nil { live = status.State == "running" || status.State == "pending" || status.State == "starting" } if (registered || live) && !opts.overwrite { @@ -1401,15 +1401,15 @@ func runDeploy(spinloopPath, env string, dc inference.DeployConfig, opts deployO fmt.Fprintf(&buf, " ingress: %s (your public IP; override with --allowed-cidr)\n", allowedCidr) } - resp, err := remoteDeployFn(ctx, cfg, dc, allowedCidr, opts.reseed, apiKey) + resp, err := cloudDeployFn(ctx, cfg, dc, allowedCidr, opts.reseed, apiKey) if err != nil { return deployOutcome{}, err } - // Register the environment so REMOTE (and the other remote + // Register the environment so REMOTE (and the other cloud // subcommands) resolve to it from now on. cfg.BaseURL = resp.BaseURL - if err := remote.SaveEnvironment(env, cfg); err != nil { + if err := cloud.SaveEnvironment(env, cfg); err != nil { return deployOutcome{}, err } @@ -1425,11 +1425,11 @@ func runDeploy(spinloopPath, env string, dc inference.DeployConfig, opts deployO buf.WriteString("api key: created\n") } if resp.Seeding { - fmt.Fprintf(&buf, "seeding the weights — follow it with `spinloop remote seed status %s`.\n", resp.SeedID) - buf.WriteString("Wait for it to finish before `spinloop remote start`, or the instance will\n") + fmt.Fprintf(&buf, "seeding the weights — follow it with `spinloop cloud seed status %s`.\n", resp.SeedID) + buf.WriteString("Wait for it to finish before `spinloop cloud start`, or the instance will\n") buf.WriteString("start against an incomplete download.\n") } else { - buf.WriteString("weights already in place — `spinloop remote start` will serve this.\n") + buf.WriteString("weights already in place — `spinloop cloud start` will serve this.\n") } return deployOutcome{ @@ -1474,12 +1474,12 @@ func detectPublicCIDR(ctx context.Context) (string, error) { // Test seams: the suite calls these the way the tree does — each runs its // command through execCmd rather than parsing a private FlagSet. -func cmdRemoteBootstrap(args []string) error { return execCmd(remoteBootstrapCmd(), args) } -func cmdRemoteStart(args []string) error { return execCmd(remoteStartCmd(), args) } -func cmdRemotePause(args []string) error { return execCmd(remotePauseCmd(), args) } -func cmdRemoteRestart(args []string) error { return execCmd(remoteRestartCmd(), args) } -func cmdRemoteStop(args []string) error { return execCmd(remoteStopCmd(), args) } -func cmdRemoteDeploy(args []string) error { return execCmd(remoteDeployCmd(), args) } -func cmdRemoteEnv(args []string) error { return execCmd(remoteEnvCmd(), args) } -func cmdRemoteList(args []string) error { return execCmd(remoteListCmd(), args) } -func cmdRemoteKeep(args []string) error { return execCmd(remoteKeepCmd(), args) } +func cmdCloudBootstrap(args []string) error { return execCmd(cloudBootstrapCmd(), args) } +func cmdCloudStart(args []string) error { return execCmd(cloudStartCmd(), args) } +func cmdCloudPause(args []string) error { return execCmd(cloudPauseCmd(), args) } +func cmdCloudRestart(args []string) error { return execCmd(cloudRestartCmd(), args) } +func cmdCloudStop(args []string) error { return execCmd(cloudStopCmd(), args) } +func cmdCloudDeploy(args []string) error { return execCmd(cloudDeployCmd(), args) } +func cmdCloudEnv(args []string) error { return execCmd(cloudEnvCmd(), args) } +func cmdCloudList(args []string) error { return execCmd(cloudListCmd(), args) } +func cmdCloudKeep(args []string) error { return execCmd(cloudKeepCmd(), args) } diff --git a/cmd/spinloop/remote_auth.go b/cmd/spinloop/cloud_auth.go similarity index 77% rename from cmd/spinloop/remote_auth.go rename to cmd/spinloop/cloud_auth.go index c358b026..a4ae251b 100644 --- a/cmd/spinloop/remote_auth.go +++ b/cmd/spinloop/cloud_auth.go @@ -10,7 +10,7 @@ import ( "github.com/aws/aws-sdk-go-v2/aws" "github.com/spf13/cobra" - "github.com/spinloop-ai/spinloop/internal/remote" + "github.com/spinloop-ai/spinloop/internal/cloud" ) // Seams: package variables so tests drive the auth flow without AWS. The IAM @@ -18,17 +18,17 @@ import ( // ambient config for a first store and a stored-credential config for a // rotation, and records which credential each call resolved with. var ( - iamUserExistsFn = remote.IAMUserExists - iamUserAccessKeysFn = remote.IAMUserAccessKeyIDs - iamCreateAccessKeyFn = remote.IAMCreateAccessKey - iamDeleteAccessKeyFn = remote.IAMDeleteAccessKey - authCallerIdentityFn = remote.CallerIdentity + iamUserExistsFn = cloud.IAMUserExists + iamUserAccessKeysFn = cloud.IAMUserAccessKeyIDs + iamCreateAccessKeyFn = cloud.IAMCreateAccessKey + iamDeleteAccessKeyFn = cloud.IAMDeleteAccessKey + authCallerIdentityFn = cloud.CallerIdentity ) -// remoteAuthCmd is `spinloop remote auth`: store, report, or clear the +// cloudAuthCmd is `spinloop cloud auth`: store, report, or clear the // long-lived control-plane credential this machine signs day-to-day commands // with. With no flag it reports what is stored, from the local store only. -func remoteAuthCmd() *cobra.Command { +func cloudAuthCmd() *cobra.Command { var ( store bool clear bool @@ -39,7 +39,7 @@ func remoteAuthCmd() *cobra.Command { Short: "store, report, or clear the control-plane credential", Long: `stores a long-lived control-plane credential in this machine's OS keystore (keychain, credential manager, or secret service) so the day-to-day -remote commands sign without a fresh SSO log-in. With no flag it reports what +cloud commands sign without a fresh SSO log-in. With no flag it reports what is stored; --store stores the credential for a region, rotating it when one is already stored; --clear removes it and deletes the access key.`, Args: cobra.ArbitraryArgs, @@ -48,7 +48,7 @@ is already stored; --clear removes it and deletes the access key.`, ValidArgsFunction: noPositionals, RunE: func(c *cobra.Command, _ []string) error { resolve(c) - return runRemoteAuth(store, clear, region) + return runCloudAuth(store, clear, region) }, } fs := c.Flags() @@ -60,32 +60,32 @@ is already stored; --clear removes it and deletes the access key.`, return c } -// runRemoteAuth is the body of `spinloop remote auth`. -func runRemoteAuth(store, clear bool, regionFlag string) error { +// runCloudAuth is the body of `spinloop cloud auth`. +func runCloudAuth(store, clear bool, regionFlag string) error { region := resolveRegion(regionFlag) ctx, stop := signal.NotifyContext(context.Background(), os.Interrupt) defer stop() switch { case clear: - return runRemoteAuthClear(ctx, region) + return runCloudAuthClear(ctx, region) case store: - return runRemoteAuthStore(ctx, region) + return runCloudAuthStore(ctx, region) default: - return runRemoteAuthReport() + return runCloudAuthReport() } } -// runRemoteAuthReport lists every stored credential from the local store. It +// runCloudAuthReport lists every stored credential from the local store. It // makes no AWS call: the report is this machine's own book, and it must work // on a machine whose credentials are expired or absent — which is the point // of the stored key. -func runRemoteAuthReport() error { - creds, err := remote.ListStoredCredentials() +func runCloudAuthReport() error { + creds, err := cloud.ListStoredCredentials() if err != nil { return fmt.Errorf("reading the stored credentials: %w", err) } if len(creds) == 0 { - fmt.Println("No stored credential. Store one with `spinloop remote auth --store`.") + fmt.Println("No stored credential. Store one with `spinloop cloud auth --store`.") return nil } w := os.Stdout @@ -97,20 +97,20 @@ func runRemoteAuthReport() error { return nil } -// runRemoteAuthStore stores a credential for the region. With nothing stored +// runCloudAuthStore stores a credential for the region. With nothing stored // it is a first store: it runs on the caller's ambient credentials, which are // the administrator's, and verifies the new key against the caller's account. // With an entry already stored it is a rotation: it runs on the stored // credential alone — no administrator or other ambient credential required — // and deletes the superseded key on the AWS side once the new one is verified // and swapped in. -func runRemoteAuthStore(ctx context.Context, region string) error { - existing, rotating := remote.LookupStoredCredential(region) +func runCloudAuthStore(ctx context.Context, region string) error { + existing, rotating := cloud.LookupStoredCredential(region) var cfg aws.Config var expectedAccount string if rotating { - cfg = remote.ConfigFromStored(existing) + cfg = cloud.ConfigFromStored(existing) expectedAccount = existing.Account } else { var err error @@ -125,28 +125,28 @@ func runRemoteAuthStore(ctx context.Context, region string) error { } } - exists, err := iamUserExistsFn(ctx, cfg, remote.ControlPlaneUserName) + exists, err := iamUserExistsFn(ctx, cfg, cloud.ControlPlaneUserName) if err != nil { return fmt.Errorf("checking the control-plane user: %w", err) } if !exists { - return fmt.Errorf("the control-plane user %q does not exist in this account — re-run `spinloop remote bootstrap` to create it", - remote.ControlPlaneUserName) + return fmt.Errorf("the control-plane user %q does not exist in this account — re-run `spinloop cloud bootstrap` to create it", + cloud.ControlPlaneUserName) } if !rotating { // IAM allows two access keys per user. A first store would be the // third, so check before creating rather than after a failure. - keys, err := iamUserAccessKeysFn(ctx, cfg, remote.ControlPlaneUserName) + keys, err := iamUserAccessKeysFn(ctx, cfg, cloud.ControlPlaneUserName) if err != nil { return fmt.Errorf("listing the control-plane user's access keys: %w", err) } if len(keys) >= 2 { - return fmt.Errorf("the control-plane user already has two access keys — delete one and run `spinloop remote auth --store` again") + return fmt.Errorf("the control-plane user already has two access keys — delete one and run `spinloop cloud auth --store` again") } } - keyID, secret, err := iamCreateAccessKeyFn(ctx, cfg, remote.ControlPlaneUserName) + keyID, secret, err := iamCreateAccessKeyFn(ctx, cfg, cloud.ControlPlaneUserName) if err != nil { return fmt.Errorf("creating the access key: %w", err) } @@ -154,7 +154,7 @@ func runRemoteAuthStore(ctx context.Context, region string) error { // Verify the new key resolves to the expected account before anything is // stored. A key that cannot be verified, or that resolves elsewhere, is // deleted on the AWS side rather than left behind. - probe := remote.ConfigFromStored(remote.StoredCredential{AccessKeyID: keyID, SecretAccessKey: secret, Region: region}) + probe := cloud.ConfigFromStored(cloud.StoredCredential{AccessKeyID: keyID, SecretAccessKey: secret, Region: region}) account, err := verifyNewKey(ctx, probe) if err != nil { deleteStrayKey(ctx, cfg, keyID) @@ -166,20 +166,20 @@ func runRemoteAuthStore(ctx context.Context, region string) error { account, expectedAccount) } - cred := remote.StoredCredential{ + cred := cloud.StoredCredential{ AccessKeyID: keyID, SecretAccessKey: secret, Account: account, - User: remote.ControlPlaneUserName, + User: cloud.ControlPlaneUserName, Region: region, StoredAt: time.Now().UTC(), } - if err := remote.StoreCredential(cred); err != nil { + if err := cloud.StoreCredential(cred); err != nil { return fmt.Errorf("storing the credential: %w", err) } // Read the entry back: it carries the store kind the write recorded, and // a store that cannot hand back what it was given has not stored it. - cred, ok := remote.LookupStoredCredential(region) + cred, ok := cloud.LookupStoredCredential(region) if !ok { return fmt.Errorf("the credential was stored but cannot be read back") } @@ -188,11 +188,11 @@ func runRemoteAuthStore(ctx context.Context, region string) error { fmt.Fprintf(w, "Stored the control-plane credential for %s:\n", region) fmt.Fprintf(w, " Account: %s\n", account) fmt.Fprintf(w, " Region: %s\n", region) - fmt.Fprintf(w, " User: %s\n", remote.ControlPlaneUserName) + fmt.Fprintf(w, " User: %s\n", cloud.ControlPlaneUserName) fmt.Fprintf(w, " Key id: %s\n", keyID) fmt.Fprintf(w, " Store: %s\n", cred.Store) if rotating { - if err := iamDeleteAccessKeyFn(ctx, probe, remote.ControlPlaneUserName, existing.AccessKeyID); err != nil { + if err := iamDeleteAccessKeyFn(ctx, probe, cloud.ControlPlaneUserName, existing.AccessKeyID); err != nil { fmt.Fprintf(w, "The superseded key %s could not be deleted on the AWS side (%v) — it may still exist.\n", existing.AccessKeyID, err) } else { @@ -253,24 +253,24 @@ func invalidClientToken(err error) bool { // to verify, so a failed store leaves no live key on the AWS side. A failure // here is reported, not returned: the store has already failed. func deleteStrayKey(ctx context.Context, cfg aws.Config, keyID string) { - if err := iamDeleteAccessKeyFn(ctx, cfg, remote.ControlPlaneUserName, keyID); err != nil { + if err := iamDeleteAccessKeyFn(ctx, cfg, cloud.ControlPlaneUserName, keyID); err != nil { fmt.Fprintf(os.Stderr, "The new key %s could not be deleted on the AWS side (%v) — delete it manually.\n", keyID, err) } } -// runRemoteAuthClear removes the stored credential for the region and deletes +// runCloudAuthClear removes the stored credential for the region and deletes // its access key on the AWS side, using the stored credential, so a cleared // key does not linger in the account. If the AWS-side deletion cannot be // made, the local entry is still removed and the failure reported. -func runRemoteAuthClear(ctx context.Context, region string) error { - cred, ok := remote.LookupStoredCredential(region) +func runCloudAuthClear(ctx context.Context, region string) error { + cred, ok := cloud.LookupStoredCredential(region) if !ok { fmt.Printf("No stored credential for %s.\n", region) return nil } - cfg := remote.ConfigFromStored(cred) + cfg := cloud.ConfigFromStored(cred) deleteErr := iamDeleteAccessKeyFn(ctx, cfg, cred.User, cred.AccessKeyID) - if err := remote.DeleteStoredCredential(region); err != nil { + if err := cloud.DeleteStoredCredential(region); err != nil { return fmt.Errorf("removing the stored credential: %w", err) } w := os.Stderr diff --git a/cmd/spinloop/remote_auth_test.go b/cmd/spinloop/cloud_auth_test.go similarity index 81% rename from cmd/spinloop/remote_auth_test.go rename to cmd/spinloop/cloud_auth_test.go index 47000557..2b32cde9 100644 --- a/cmd/spinloop/remote_auth_test.go +++ b/cmd/spinloop/cloud_auth_test.go @@ -9,7 +9,7 @@ import ( "time" "github.com/aws/aws-sdk-go-v2/aws" - "github.com/spinloop-ai/spinloop/internal/remote" + "github.com/spinloop-ai/spinloop/internal/cloud" ) // authSeamState is what the stubbed IAM/STS seams answer with and record. @@ -111,7 +111,7 @@ func authStoreEnv(t *testing.T) { t.Helper() isolateConfig(t) t.Setenv("SPINLOOP_CONFIG_DIR", "") - t.Setenv("SPINLOOP_REMOTE_KEYSTORE", "file") + t.Setenv("SPINLOOP_CLOUD_KEYSTORE", "file") } // noAmbientCreds pins the process to have no ambient AWS credential at all: @@ -129,11 +129,11 @@ func noAmbientCreds(t *testing.T) { func seedStoredCred(t *testing.T, region, keyID, secret, account string) { t.Helper() - if err := remote.StoreCredential(remote.StoredCredential{ + if err := cloud.StoreCredential(cloud.StoredCredential{ AccessKeyID: keyID, SecretAccessKey: secret, Account: account, - User: remote.ControlPlaneUserName, + User: cloud.ControlPlaneUserName, Region: region, StoredAt: time.Now().UTC(), }); err != nil { @@ -141,7 +141,7 @@ func seedStoredCred(t *testing.T, region, keyID, secret, account string) { } } -func TestRemoteAuthStoreFirstStore(t *testing.T) { +func TestCloudAuthStoreFirstStore(t *testing.T) { authStoreEnv(t) stubAWSEnv(t) st := stubAuthSeams(t) @@ -153,19 +153,19 @@ func TestRemoteAuthStoreFirstStore(t *testing.T) { st.accounts["AKIANEWNEWNEWNEW01"] = "1" out := captureStderr(t, func() { - if err := runRemoteAuth(true, false, "ap-southeast-2"); err != nil { - t.Fatalf("runRemoteAuth --store: %v", err) + if err := runCloudAuth(true, false, "ap-southeast-2"); err != nil { + t.Fatalf("runCloudAuth --store: %v", err) } }) - cred, ok := remote.LookupStoredCredential("ap-southeast-2") + cred, ok := cloud.LookupStoredCredential("ap-southeast-2") if !ok { t.Fatal("the credential was not stored") } if cred.AccessKeyID != "AKIANEWNEWNEWNEW01" || cred.SecretAccessKey != "new-secret" { t.Errorf("stored key = %q/%q", cred.AccessKeyID, cred.SecretAccessKey) } - if cred.Account != "1" || cred.Region != "ap-southeast-2" || cred.User != remote.ControlPlaneUserName { + if cred.Account != "1" || cred.Region != "ap-southeast-2" || cred.User != cloud.ControlPlaneUserName { t.Errorf("stored entry = %+v", cred) } if cred.Store != "file" { @@ -174,13 +174,13 @@ func TestRemoteAuthStoreFirstStore(t *testing.T) { if st.existsKey != "AKIATESTTESTTESTTEST" { t.Errorf("the user check ran with %q, want the ambient credential", st.existsKey) } - if st.createdFor != remote.ControlPlaneUserName { + if st.createdFor != cloud.ControlPlaneUserName { t.Errorf("the key was created for %q", st.createdFor) } if len(st.deleted) != 0 { t.Errorf("a failed-free store deleted keys: %v", st.deleted) } - for _, want := range []string{"ap-southeast-2", "1", remote.ControlPlaneUserName, "AKIANEWNEWNEWNEW01", "file"} { + for _, want := range []string{"ap-southeast-2", "1", cloud.ControlPlaneUserName, "AKIANEWNEWNEWNEW01", "file"} { if !strings.Contains(out, want) { t.Errorf("confirmation missing %q:\n%s", want, out) } @@ -190,22 +190,22 @@ func TestRemoteAuthStoreFirstStore(t *testing.T) { } } -func TestRemoteAuthStoreMissingUser(t *testing.T) { +func TestCloudAuthStoreMissingUser(t *testing.T) { authStoreEnv(t) stubAWSEnv(t) st := stubAuthSeams(t) st.userExists = false - err := runRemoteAuth(true, false, "ap-southeast-2") - if err == nil || !strings.Contains(err.Error(), "spinloop remote bootstrap") { - t.Fatalf("error = %v, want it naming spinloop remote bootstrap", err) + err := runCloudAuth(true, false, "ap-southeast-2") + if err == nil || !strings.Contains(err.Error(), "spinloop cloud bootstrap") { + t.Fatalf("error = %v, want it naming spinloop cloud bootstrap", err) } - if _, ok := remote.LookupStoredCredential("ap-southeast-2"); ok { + if _, ok := cloud.LookupStoredCredential("ap-southeast-2"); ok { t.Error("a credential was stored despite the missing user") } } -func TestRemoteAuthStoreAccountMismatch(t *testing.T) { +func TestCloudAuthStoreAccountMismatch(t *testing.T) { authStoreEnv(t) stubAWSEnv(t) st := stubAuthSeams(t) @@ -215,11 +215,11 @@ func TestRemoteAuthStoreAccountMismatch(t *testing.T) { st.account = "1" st.accounts["AKIANEWNEWNEWNEW01"] = "2" - err := runRemoteAuth(true, false, "ap-southeast-2") + err := runCloudAuth(true, false, "ap-southeast-2") if err == nil || !strings.Contains(err.Error(), "resolves to account 2") { t.Fatalf("error = %v, want the account mismatch", err) } - if _, ok := remote.LookupStoredCredential("ap-southeast-2"); ok { + if _, ok := cloud.LookupStoredCredential("ap-southeast-2"); ok { t.Error("a key that resolves elsewhere was stored") } if len(st.deleted) != 1 || st.deleted[0] != "AKIANEWNEWNEWNEW01" { @@ -227,7 +227,7 @@ func TestRemoteAuthStoreAccountMismatch(t *testing.T) { } } -func TestRemoteAuthStoreRetriesUnpropagatedKey(t *testing.T) { +func TestCloudAuthStoreRetriesUnpropagatedKey(t *testing.T) { authStoreEnv(t) stubAWSEnv(t) st := stubAuthSeams(t) @@ -240,7 +240,7 @@ func TestRemoteAuthStoreRetriesUnpropagatedKey(t *testing.T) { st.identityErrs = []error{invalidClientTokenErr, invalidClientTokenErr} out := captureStderr(t, func() { - if err := runRemoteAuth(true, false, "ap-southeast-2"); err != nil { + if err := runCloudAuth(true, false, "ap-southeast-2"); err != nil { t.Fatalf("a key that propagates late must still be stored: %v", err) } }) @@ -250,7 +250,7 @@ func TestRemoteAuthStoreRetriesUnpropagatedKey(t *testing.T) { if len(st.deleted) != 0 { t.Errorf("a late-propagating key was deleted on the AWS side: %v", st.deleted) } - if _, ok := remote.LookupStoredCredential("ap-southeast-2"); !ok { + if _, ok := cloud.LookupStoredCredential("ap-southeast-2"); !ok { t.Error("the credential was not stored") } if !strings.Contains(out, "not resolvable yet") { @@ -258,7 +258,7 @@ func TestRemoteAuthStoreRetriesUnpropagatedKey(t *testing.T) { } } -func TestRemoteAuthStoreVerifyExhausted(t *testing.T) { +func TestCloudAuthStoreVerifyExhausted(t *testing.T) { authStoreEnv(t) stubAWSEnv(t) st := stubAuthSeams(t) @@ -272,7 +272,7 @@ func TestRemoteAuthStoreVerifyExhausted(t *testing.T) { st.identityErrs[i] = invalidClientTokenErr } - err := runRemoteAuth(true, false, "ap-southeast-2") + err := runCloudAuth(true, false, "ap-southeast-2") if err == nil || !strings.Contains(err.Error(), "verifying the new access key") { t.Fatalf("error = %v, want the exhausted verification", err) } @@ -282,12 +282,12 @@ func TestRemoteAuthStoreVerifyExhausted(t *testing.T) { if len(st.deleted) != 1 || st.deleted[0] != "AKIANEWNEWNEWNEW01" { t.Errorf("the unverifiable key was not deleted: %v", st.deleted) } - if _, ok := remote.LookupStoredCredential("ap-southeast-2"); ok { + if _, ok := cloud.LookupStoredCredential("ap-southeast-2"); ok { t.Error("an unverifiable key was stored") } } -func TestRemoteAuthStoreVerifyOtherErrorNoRetry(t *testing.T) { +func TestCloudAuthStoreVerifyOtherErrorNoRetry(t *testing.T) { authStoreEnv(t) stubAWSEnv(t) st := stubAuthSeams(t) @@ -298,7 +298,7 @@ func TestRemoteAuthStoreVerifyOtherErrorNoRetry(t *testing.T) { st.identityErrFor = "AKIANEWNEWNEWNEW01" st.identityErrs = []error{errors.New("operation error STS: GetCallerIdentity, api error AccessDenied: not authorised")} - err := runRemoteAuth(true, false, "ap-southeast-2") + err := runCloudAuth(true, false, "ap-southeast-2") if err == nil || !strings.Contains(err.Error(), "verifying the new access key") { t.Fatalf("error = %v, want the verification failure", err) } @@ -310,23 +310,23 @@ func TestRemoteAuthStoreVerifyOtherErrorNoRetry(t *testing.T) { } } -func TestRemoteAuthStoreTwoKeys(t *testing.T) { +func TestCloudAuthStoreTwoKeys(t *testing.T) { authStoreEnv(t) stubAWSEnv(t) st := stubAuthSeams(t) st.userExists = true st.accessKeys = []string{"AKIAFIRSTKEYKEYKEY01", "AKIASECONDKEYKEY01"} - err := runRemoteAuth(true, false, "ap-southeast-2") + err := runCloudAuth(true, false, "ap-southeast-2") if err == nil || !strings.Contains(err.Error(), "two access keys") { t.Fatalf("error = %v, want the two-key cap named", err) } - if _, ok := remote.LookupStoredCredential("ap-southeast-2"); ok { + if _, ok := cloud.LookupStoredCredential("ap-southeast-2"); ok { t.Error("a credential was stored at the two-key cap") } } -func TestRemoteAuthStoreRotationNoAmbient(t *testing.T) { +func TestCloudAuthStoreRotationNoAmbient(t *testing.T) { authStoreEnv(t) noAmbientCreds(t) seedStoredCred(t, "ap-southeast-2", "AKIAOLDOLDOLDOLD01", "old-secret", "1") @@ -337,7 +337,7 @@ func TestRemoteAuthStoreRotationNoAmbient(t *testing.T) { st.accounts["AKIANEWNEWNEWNEW01"] = "1" out := captureStderr(t, func() { - if err := runRemoteAuth(true, false, "ap-southeast-2"); err != nil { + if err := runCloudAuth(true, false, "ap-southeast-2"); err != nil { t.Fatalf("rotation without ambient credentials: %v", err) } }) @@ -348,7 +348,7 @@ func TestRemoteAuthStoreRotationNoAmbient(t *testing.T) { if len(st.identity) != 1 || st.identity[0] != "AKIANEWNEWNEWNEW01" { t.Errorf("identity calls resolved %v, want the new key only", st.identity) } - cred, ok := remote.LookupStoredCredential("ap-southeast-2") + cred, ok := cloud.LookupStoredCredential("ap-southeast-2") if !ok { t.Fatal("the rotated credential is missing") } @@ -363,18 +363,18 @@ func TestRemoteAuthStoreRotationNoAmbient(t *testing.T) { } } -func TestRemoteAuthClear(t *testing.T) { +func TestCloudAuthClear(t *testing.T) { authStoreEnv(t) noAmbientCreds(t) seedStoredCred(t, "ap-southeast-2", "AKIACLEARMEKEYKEY01", "secret", "1") st := stubAuthSeams(t) out := captureStderr(t, func() { - if err := runRemoteAuth(false, true, "ap-southeast-2"); err != nil { - t.Fatalf("runRemoteAuth --clear: %v", err) + if err := runCloudAuth(false, true, "ap-southeast-2"); err != nil { + t.Fatalf("runCloudAuth --clear: %v", err) } }) - if _, ok := remote.LookupStoredCredential("ap-southeast-2"); ok { + if _, ok := cloud.LookupStoredCredential("ap-southeast-2"); ok { t.Error("the entry was not removed") } if len(st.deleted) != 1 || st.deleted[0] != "AKIACLEARMEKEYKEY01" { @@ -385,7 +385,7 @@ func TestRemoteAuthClear(t *testing.T) { } } -func TestRemoteAuthClearAWSDeploymentFails(t *testing.T) { +func TestCloudAuthClearAWSDeploymentFails(t *testing.T) { authStoreEnv(t) noAmbientCreds(t) seedStoredCred(t, "ap-southeast-2", "AKIACLEARMEKEYKEY01", "secret", "1") @@ -393,11 +393,11 @@ func TestRemoteAuthClearAWSDeploymentFails(t *testing.T) { st.deleteErr = errors.New("boom") out := captureStderr(t, func() { - if err := runRemoteAuth(false, true, "ap-southeast-2"); err != nil { + if err := runCloudAuth(false, true, "ap-southeast-2"); err != nil { t.Fatalf("a failed AWS-side deletion must not fail the clear: %v", err) } }) - if _, ok := remote.LookupStoredCredential("ap-southeast-2"); ok { + if _, ok := cloud.LookupStoredCredential("ap-southeast-2"); ok { t.Error("the entry was not removed after a failed AWS-side deletion") } if !strings.Contains(out, "boom") || !strings.Contains(out, "may still exist") { @@ -405,12 +405,12 @@ func TestRemoteAuthClearAWSDeploymentFails(t *testing.T) { } } -func TestRemoteAuthClearNone(t *testing.T) { +func TestCloudAuthClearNone(t *testing.T) { authStoreEnv(t) noAmbientCreds(t) out := captureStdout(t, func() { - if err := runRemoteAuth(false, true, "ap-southeast-2"); err != nil { + if err := runCloudAuth(false, true, "ap-southeast-2"); err != nil { t.Fatalf("clearing nothing is not an error: %v", err) } }) @@ -419,12 +419,12 @@ func TestRemoteAuthClearNone(t *testing.T) { } } -func TestRemoteAuthReportEmpty(t *testing.T) { +func TestCloudAuthReportEmpty(t *testing.T) { authStoreEnv(t) out := captureStdout(t, func() { - if err := runRemoteAuth(false, false, "ap-southeast-2"); err != nil { - t.Fatalf("runRemoteAuth: %v", err) + if err := runCloudAuth(false, false, "ap-southeast-2"); err != nil { + t.Fatalf("runCloudAuth: %v", err) } }) if !strings.Contains(out, "No stored credential") || !strings.Contains(out, "--store") { @@ -432,7 +432,7 @@ func TestRemoteAuthReportEmpty(t *testing.T) { } } -func TestRemoteAuthReport(t *testing.T) { +func TestCloudAuthReport(t *testing.T) { authStoreEnv(t) noAmbientCreds(t) seedStoredCred(t, "ap-southeast-2", "AKIASECONDKEYKEY01", "second-secret", "1") @@ -440,8 +440,8 @@ func TestRemoteAuthReport(t *testing.T) { st := stubAuthSeams(t) out := captureStdout(t, func() { - if err := runRemoteAuth(false, false, ""); err != nil { - t.Fatalf("runRemoteAuth: %v", err) + if err := runCloudAuth(false, false, ""); err != nil { + t.Fatalf("runCloudAuth: %v", err) } }) if len(st.identity) != 0 || len(st.deleted) != 0 { @@ -455,8 +455,8 @@ func TestRemoteAuthReport(t *testing.T) { t.Errorf("unexpected header:\n%s", lines[0]) } for i, want := range []string{ - "ap-southeast-2\t1\t" + remote.ControlPlaneUserName + "\tAKIASECONDKEYKEY01", - "eu-west-1\t2\t" + remote.ControlPlaneUserName + "\tAKIAFIRSTKEYKEYKEY01", + "ap-southeast-2\t1\t" + cloud.ControlPlaneUserName + "\tAKIASECONDKEYKEY01", + "eu-west-1\t2\t" + cloud.ControlPlaneUserName + "\tAKIAFIRSTKEYKEYKEY01", } { if !strings.HasPrefix(lines[i+1], want) { t.Errorf("entry %d = %q, want it starting %q", i+1, lines[i+1], want) diff --git a/cmd/spinloop/remote_bake.go b/cmd/spinloop/cloud_bake.go similarity index 86% rename from cmd/spinloop/remote_bake.go rename to cmd/spinloop/cloud_bake.go index 9156edfe..9cc9dc3a 100644 --- a/cmd/spinloop/remote_bake.go +++ b/cmd/spinloop/cloud_bake.go @@ -10,7 +10,7 @@ import ( "github.com/spf13/cobra" ) -// defaultBakeRunners is what `spinloop remote bake` bakes when no runners are +// defaultBakeRunners is what `spinloop cloud bake` bakes when no runners are // named, and the full accepted set — the two engines an environment can run. var defaultBakeRunners = []string{"llamacpp", "vllm"} @@ -40,12 +40,12 @@ func bakeRunnerSlot(_ *cobra.Command, args []string, _ string) ([]string, cobra. return out, cobra.ShellCompDirectiveNoFileComp } -// remoteBakeCmd starts an AMI bake for each named runner. It drives the same +// cloudBakeCmd starts an AMI bake for each named runner. It drives the same // CDK project bootstrap deploys but deploys nothing — the control plane (and // its Image Builder pipelines) must already exist, so a missing one fails fast // naming bootstrap rather than deploying implicitly. The wait is the default: -// the step after a bake is `spinloop remote deploy`, which needs the AMI. -func remoteBakeCmd() *cobra.Command { +// the step after a bake is `spinloop cloud deploy`, which needs the AMI. +func cloudBakeCmd() *cobra.Command { var ( noWait bool ref string @@ -67,7 +67,7 @@ queued, reporting how to check on them.`, ValidArgsFunction: bakeRunnerSlot, RunE: func(c *cobra.Command, args []string) error { resolve(c) - return runRemoteBake(args, noWait, ref, dir, region, pkgMgr) + return runCloudBake(args, noWait, ref, dir, region, pkgMgr) }, } fs := c.Flags() @@ -80,8 +80,8 @@ queued, reporting how to check on them.`, return c } -// runRemoteBake is the body of `spinloop remote bake`. -func runRemoteBake(args []string, noWait bool, ref, dir, region string, pkgMgr string) error { +// runCloudBake is the body of `spinloop cloud bake`. +func runCloudBake(args []string, noWait bool, ref, dir, region string, pkgMgr string) error { runners := defaultBakeRunners if len(args) > 0 { runners = args @@ -123,7 +123,7 @@ func runRemoteBake(args []string, noWait bool, ref, dir, region string, pkgMgr s } if !deployed { return fmt.Errorf( - "the control plane is not deployed in %s — run `spinloop remote bootstrap` first", resolvedRegion) + "the control plane is not deployed in %s — run `spinloop cloud bootstrap` first", resolvedRegion) } if err := downloadFn(ctx, loc.ref, loc.dir); err != nil { @@ -158,4 +158,4 @@ func runRemoteBake(args []string, noWait bool, ref, dir, region string, pkgMgr s return pruneSourceCaches(loc) } -func cmdRemoteBake(args []string) error { return execCmd(remoteBakeCmd(), args) } +func cmdCloudBake(args []string) error { return execCmd(cloudBakeCmd(), args) } diff --git a/cmd/spinloop/remote_bake_test.go b/cmd/spinloop/cloud_bake_test.go similarity index 86% rename from cmd/spinloop/remote_bake_test.go rename to cmd/spinloop/cloud_bake_test.go index 07a8e141..eac79ac8 100644 --- a/cmd/spinloop/remote_bake_test.go +++ b/cmd/spinloop/cloud_bake_test.go @@ -9,7 +9,7 @@ import ( "time" "github.com/aws/aws-sdk-go-v2/aws" - "github.com/spinloop-ai/spinloop/internal/remote" + "github.com/spinloop-ai/spinloop/internal/cloud" ) // stubBakeSeams wires the shared seams to hermetic fakes for the bake flow: @@ -50,7 +50,7 @@ func TestBake_DefaultRunners(t *testing.T) { isolateConfig(t) stubAWSEnv(t) steps, bakedCalls := stubBakeSeams(t, true) - if err := cmdRemoteBake([]string{"--region", "us-east-1", "--no-wait"}); err != nil { + if err := cmdCloudBake([]string{"--region", "us-east-1", "--no-wait"}); err != nil { t.Fatal(err) } var got []string @@ -70,7 +70,7 @@ func TestBake_SingleRunner(t *testing.T) { isolateConfig(t) stubAWSEnv(t) steps, _ := stubBakeSeams(t, true) - if err := cmdRemoteBake([]string{"llamacpp", "--region", "us-east-1", "--no-wait"}); err != nil { + if err := cmdCloudBake([]string{"llamacpp", "--region", "us-east-1", "--no-wait"}); err != nil { t.Fatal(err) } var got []string @@ -87,7 +87,7 @@ func TestBake_UnknownRunnerRejected(t *testing.T) { isolateConfig(t) stubAWSEnv(t) steps, _ := stubBakeSeams(t, true) - err := cmdRemoteBake([]string{"bogus", "--region", "us-east-1"}) + err := cmdCloudBake([]string{"bogus", "--region", "us-east-1"}) if err == nil { t.Fatal("expected an error for a bad runner") } @@ -103,11 +103,11 @@ func TestBake_NoControlPlane(t *testing.T) { isolateConfig(t) stubAWSEnv(t) steps, _ := stubBakeSeams(t, false) - err := cmdRemoteBake([]string{"--region", "us-east-1"}) + err := cmdCloudBake([]string{"--region", "us-east-1"}) if err == nil { t.Fatal("expected an error when the control plane is missing") } - if !strings.Contains(err.Error(), "spinloop remote bootstrap") { + if !strings.Contains(err.Error(), "spinloop cloud bootstrap") { t.Errorf("error should point at bootstrap, got %v", err) } if len(*steps) != 0 { @@ -135,7 +135,7 @@ func TestBake_WaitsByDefault(t *testing.T) { } t.Cleanup(func() { bakedFn = origBaked }) - if err := cmdRemoteBake([]string{"--region", "us-east-1"}); err != nil { + if err := cmdCloudBake([]string{"--region", "us-east-1"}); err != nil { t.Fatal(err) } if calls < 2 { @@ -150,11 +150,11 @@ func TestBake_SkipsInstallWhenPresent(t *testing.T) { isolateConfig(t) stubAWSEnv(t) steps, _ := stubBakeSeams(t, true) - cdkDir := must1(remote.SourceDir(remote.ResolveRef(version, ""))) + cdkDir := must1(cloud.SourceDir(cloud.ResolveRef(version, ""))) if err := os.MkdirAll(filepath.Join(cdkDir, "node_modules"), 0o755); err != nil { t.Fatal(err) } - if err := cmdRemoteBake([]string{"--region", "us-east-1", "--no-wait"}); err != nil { + if err := cmdCloudBake([]string{"--region", "us-east-1", "--no-wait"}); err != nil { t.Fatal(err) } for _, s := range *steps { @@ -183,12 +183,12 @@ func TestBake_PrunesOtherRefsByDefault(t *testing.T) { isolateConfig(t) stubAWSEnv(t) stubBakeSeams(t, true) - root := must1(remote.SourceRoot()) + root := must1(cloud.SourceRoot()) stale := filepath.Join(root, "v9.9.9") if err := os.MkdirAll(stale, 0o755); err != nil { t.Fatal(err) } - if err := cmdRemoteBake([]string{"--region", "us-east-1", "--no-wait"}); err != nil { + if err := cmdCloudBake([]string{"--region", "us-east-1", "--no-wait"}); err != nil { t.Fatal(err) } if _, err := os.Stat(stale); !os.IsNotExist(err) { @@ -200,12 +200,12 @@ func TestBake_ExplicitDirNotPruned(t *testing.T) { isolateConfig(t) stubAWSEnv(t) stubBakeSeams(t, true) - root := must1(remote.SourceRoot()) + root := must1(cloud.SourceRoot()) stale := filepath.Join(root, "v9.9.9") if err := os.MkdirAll(stale, 0o755); err != nil { t.Fatal(err) } - if err := cmdRemoteBake([]string{"--region", "us-east-1", "--no-wait", "--dir", t.TempDir()}); err != nil { + if err := cmdCloudBake([]string{"--region", "us-east-1", "--no-wait", "--dir", t.TempDir()}); err != nil { t.Fatal(err) } if _, err := os.Stat(stale); err != nil { diff --git a/cmd/spinloop/remote_bootstrap.go b/cmd/spinloop/cloud_bootstrap.go similarity index 89% rename from cmd/spinloop/remote_bootstrap.go rename to cmd/spinloop/cloud_bootstrap.go index 4e236f13..ed4b4b36 100644 --- a/cmd/spinloop/remote_bootstrap.go +++ b/cmd/spinloop/cloud_bootstrap.go @@ -15,14 +15,14 @@ import ( "github.com/aws/aws-sdk-go-v2/aws" "github.com/spf13/cobra" - "github.com/spinloop-ai/spinloop/internal/remote" + "github.com/spinloop-ai/spinloop/internal/cloud" ) // controlPlaneStackName is the CloudFormation stack holding the account-level -// control plane that bootstrap deploys and `spinloop remote deploy` discovers. It +// control plane that bootstrap deploys and `spinloop cloud deploy` discovers. It // is also the namespace the control plane's log groups sit under, so the name -// has one home in internal/remote rather than one per caller. -const controlPlaneStackName = remote.ControlPlaneStackName +// has one home in internal/cloud rather than one per caller. +const controlPlaneStackName = cloud.ControlPlaneStackName // Seams: package variables so tests drive the flow without AWS, a network, or // spawning a package manager/cdk. Shared by bootstrap and bake. @@ -30,10 +30,10 @@ type stepRunner func(ctx context.Context, name string, argv []string, workDir st var ( runStep stepRunner = execStep - downloadFn = remote.DownloadRemote - accountFn = remote.CallerIdentity - stackDeployedFn = remote.ControlPlaneStackDeployed - bakedFn = remote.BakedRunners + downloadFn = cloud.DownloadRemote + accountFn = cloud.CallerIdentity + stackDeployedFn = cloud.ControlPlaneStackDeployed + bakedFn = cloud.BakedRunners preflightFn = checkNodeAndPackageManager ) @@ -43,7 +43,7 @@ const controlPlaneVersionEnv = "SPINLOOP_CONTROL_PLANE_VERSION" // packageManagerEnv pins the Node package manager bootstrap drives the CDK // project with, when the --package-manager flag is not given. -const packageManagerEnv = "SPINLOOP_REMOTE_PACKAGE_MANAGER" +const packageManagerEnv = "SPINLOOP_CLOUD_PACKAGE_MANAGER" // packageManager shapes the argv for the two things bootstrap asks of a Node // package manager: installing dependencies and running a package.json script. @@ -117,7 +117,7 @@ func validatePackageManagerName(name string) error { } // resolvePackageManagerName applies the override precedence — the flag, then the -// SPINLOOP_REMOTE_PACKAGE_MANAGER env var — returning the pinned name (or "" to +// SPINLOOP_CLOUD_PACKAGE_MANAGER env var — returning the pinned name (or "" to // auto-detect) and whether the choice was pinned. An unrecognised value errors. func resolvePackageManagerName(flagVal string) (name string, pinned bool, err error) { if flagVal != "" { @@ -128,7 +128,7 @@ func resolvePackageManagerName(flagVal string) (name string, pinned bool, err er } // The env leg of the precedence runs through the CLI's one Viper, like // every other SPINLOOP_ read the CLI owns; the flag keeps its say first. - if env := cliViper.GetString("remote_package_manager"); env != "" { + if env := cliViper.GetString("cloud_package_manager"); env != "" { if err := validatePackageManagerName(env); err != nil { return "", false, fmt.Errorf("%s: %w", packageManagerEnv, err) } @@ -137,12 +137,12 @@ func resolvePackageManagerName(flagVal string) (name string, pinned bool, err er return "", false, nil } -// remoteBootstrapCmd deploys the account-level control plane once — +// cloudBootstrapCmd deploys the account-level control plane once — // analogous to `cdk bootstrap` — by downloading the remote/ CDK project and // driving its control-plane deploy. It creates no EIP, instance, or environment; -// those come from `spinloop remote deploy`. It starts no AMI bake either — -// `spinloop remote bake` is the separate step after it. -func remoteBootstrapCmd() *cobra.Command { +// those come from `spinloop cloud deploy`. It starts no AMI bake either — +// `spinloop cloud bake` is the separate step after it. +func cloudBootstrapCmd() *cobra.Command { var ( hfToken string ref string @@ -157,14 +157,14 @@ func remoteBootstrapCmd() *cobra.Command { Short: "set up the once-per-account control plane", Long: `does the once-per-account control-plane setup (Image Builder, the lifecycle Lambdas, shared bucket/roles/VPC) with a consent gate. It bakes no -AMIs — spinloop remote bake is the next step after it.`, +AMIs — spinloop cloud bake is the next step after it.`, Args: cobra.ArbitraryArgs, SilenceErrors: true, SilenceUsage: true, ValidArgsFunction: noPositionals, RunE: func(c *cobra.Command, _ []string) error { resolve(c) - return runRemoteBootstrap(hfToken, ref, dir, region, dryRun, assumeYes, pkgMgr) + return runCloudBootstrap(hfToken, ref, dir, region, dryRun, assumeYes, pkgMgr) }, } fs := c.Flags() @@ -180,8 +180,8 @@ AMIs — spinloop remote bake is the next step after it.`, return c } -// runRemoteBootstrap is the body of `spinloop remote bootstrap`. -func runRemoteBootstrap(hfToken, ref, dir, region string, dryRun, assumeYes bool, pkgMgr string) error { +// runCloudBootstrap is the body of `spinloop cloud bootstrap`. +func runCloudBootstrap(hfToken, ref, dir, region string, dryRun, assumeYes bool, pkgMgr string) error { pmName, pmPinned, err := resolvePackageManagerName(pkgMgr) if err != nil { return err @@ -250,8 +250,8 @@ func runRemoteBootstrap(hfToken, ref, dir, region string, dryRun, assumeYes bool } fmt.Println("\nThe account is bootstrapped. Before an environment can start, its AMI needs baking:") - fmt.Println(" spinloop remote bake # bakes the AMI(s) an environment runs from; waits until available") - fmt.Println(" spinloop remote deploy # names an environment; discovers this control plane") + fmt.Println(" spinloop cloud bake # bakes the AMI(s) an environment runs from; waits until available") + fmt.Println(" spinloop cloud deploy # names an environment; discovers this control plane") return pruneSourceCaches(loc) } @@ -315,7 +315,7 @@ func resolveRegion(flagVal string) string { // stored control-plane key is deliberately not consulted here — a day-to-day // key must not stand in for the admin while the account is being (re)built. func loadCreds(ctx context.Context, region string) (aws.Config, error) { - cfg, err := remote.LoadAmbientAWSConfig(ctx, region) + cfg, err := cloud.LoadAmbientAWSConfig(ctx, region) if err != nil { return aws.Config{}, err } @@ -396,7 +396,7 @@ func upsertEnvVar(path, key, value string) error { func renderBootstrapPlan(account, region string, ref, cdkDir string, alreadyBootstrapped bool, pm packageManager) { w := os.Stderr - fmt.Fprintln(w, "spinloop remote bootstrap — account-level control-plane setup (once per account)") + fmt.Fprintln(w, "spinloop cloud bootstrap — account-level control-plane setup (once per account)") fmt.Fprintln(w) fmt.Fprintf(w, "AWS account: %s\n", account) fmt.Fprintf(w, "Region: %s\n", region) @@ -406,7 +406,7 @@ func renderBootstrapPlan(account, region string, ref, cdkDir string, alreadyBoot } fmt.Fprintln(w) fmt.Fprintln(w, "This deploys the control plane every environment reuses:") - fmt.Fprintln(w, " • EC2 Image Builder pipelines (the AMI bakes are a separate step: spinloop remote bake)") + fmt.Fprintln(w, " • EC2 Image Builder pipelines (the AMI bakes are a separate step: spinloop cloud bake)") fmt.Fprintln(w, " • the lifecycle Lambdas (start/stop/monitor/deploy) and their IAM") fmt.Fprintln(w, " • the shared S3 weights bucket, IAM roles, and VPC") fmt.Fprintln(w) diff --git a/cmd/spinloop/remote_bootstrap_test.go b/cmd/spinloop/cloud_bootstrap_test.go similarity index 94% rename from cmd/spinloop/remote_bootstrap_test.go rename to cmd/spinloop/cloud_bootstrap_test.go index 73e2f5d3..88f34163 100644 --- a/cmd/spinloop/remote_bootstrap_test.go +++ b/cmd/spinloop/cloud_bootstrap_test.go @@ -8,7 +8,7 @@ import ( "testing" "github.com/aws/aws-sdk-go-v2/aws" - "github.com/spinloop-ai/spinloop/internal/remote" + "github.com/spinloop-ai/spinloop/internal/cloud" ) type recordedStep struct { @@ -86,7 +86,7 @@ func TestBootstrap_PlanOutput(t *testing.T) { out := captureStderr(t, func() { renderBootstrapPlan("1", "us-east-1", "v1.10.0", "/tmp/cdk/v1.10.0", false, pnpmManager) }) - for _, want := range []string{"AWS account: 1\n", "us-east-1", "Image Builder", "spinloop remote bake", "Cost:", "pnpm run deploy\n"} { + for _, want := range []string{"AWS account: 1\n", "us-east-1", "Image Builder", "spinloop cloud bake", "Cost:", "pnpm run deploy\n"} { if !strings.Contains(out, want) { t.Errorf("plan missing %q:\n%s", want, out) } @@ -103,7 +103,7 @@ func TestBootstrap_DryRunRunsNothing(t *testing.T) { isolateConfig(t) stubAWSEnv(t) steps := stubBootstrapSeams(t, false) - if err := cmdRemoteBootstrap([]string{"--region", "us-east-1", "--dry-run"}); err != nil { + if err := cmdCloudBootstrap([]string{"--region", "us-east-1", "--dry-run"}); err != nil { t.Fatal(err) } if len(*steps) != 0 { @@ -117,7 +117,7 @@ func TestBootstrap_ConfirmGate(t *testing.T) { stubAWSEnv(t) steps := stubBootstrapSeams(t, false) withStdin(t, "n\n") - if err := cmdRemoteBootstrap([]string{"--region", "us-east-1"}); err != nil { + if err := cmdCloudBootstrap([]string{"--region", "us-east-1"}); err != nil { t.Fatal(err) } if len(*steps) != 0 { @@ -129,11 +129,11 @@ func TestBootstrap_ConfirmGate(t *testing.T) { isolateConfig(t) stubAWSEnv(t) steps := stubBootstrapSeams(t, false) - if err := cmdRemoteBootstrap([]string{"--region", "us-east-1", "--yes"}); err != nil { + if err := cmdCloudBootstrap([]string{"--region", "us-east-1", "--yes"}); err != nil { t.Fatal(err) } var got []string - cdkDir := must1(remote.SourceDir(remote.ResolveRef(version, ""))) + cdkDir := must1(cloud.SourceDir(cloud.ResolveRef(version, ""))) for _, s := range *steps { got = append(got, strings.Join(s.argv, " ")) if s.dir != cdkDir { @@ -154,12 +154,12 @@ func TestBootstrap_SignpostsTheBake(t *testing.T) { stubAWSEnv(t) stubBootstrapSeams(t, false) out := captureStdout(t, func() { - if err := cmdRemoteBootstrap([]string{"--region", "us-east-1", "--yes"}); err != nil { + if err := cmdCloudBootstrap([]string{"--region", "us-east-1", "--yes"}); err != nil { t.Fatal(err) } }) - bakeIdx := strings.Index(out, "spinloop remote bake") - deployIdx := strings.Index(out, "spinloop remote deploy") + bakeIdx := strings.Index(out, "spinloop cloud bake") + deployIdx := strings.Index(out, "spinloop cloud deploy") if bakeIdx < 0 || deployIdx < 0 { t.Fatalf("success output should signpost bake then deploy:\n%s", out) } @@ -172,7 +172,7 @@ func TestBootstrap_SignpostsTheBake(t *testing.T) { // preflight resolves the ambient chain only: with no ambient credentials // there is nothing to fall back to — the stored control-plane key must not // stand in for the administrator. The loader itself is pinned to that -// behaviour by TestLoadAmbientAWSConfigIgnoresStoredKey in internal/remote. +// behaviour by TestLoadAmbientAWSConfigIgnoresStoredKey in internal/cloud. func TestLoadCredsRequiresAmbientCredentials(t *testing.T) { isolateConfig(t) t.Setenv("AWS_ACCESS_KEY_ID", "") @@ -367,7 +367,7 @@ func TestBootstrap_NpmOverrideDrivesNpmCommands(t *testing.T) { isolateConfig(t) stubAWSEnv(t) steps := stubBootstrapSeams(t, false) - if err := cmdRemoteBootstrap([]string{"--region", "us-east-1", "--yes", "--package-manager", "npm"}); err != nil { + if err := cmdCloudBootstrap([]string{"--region", "us-east-1", "--yes", "--package-manager", "npm"}); err != nil { t.Fatal(err) } var got []string diff --git a/cmd/spinloop/remote_deploy_test.go b/cmd/spinloop/cloud_deploy_test.go similarity index 87% rename from cmd/spinloop/remote_deploy_test.go rename to cmd/spinloop/cloud_deploy_test.go index c2fd55cb..d7ee0345 100644 --- a/cmd/spinloop/remote_deploy_test.go +++ b/cmd/spinloop/cloud_deploy_test.go @@ -16,8 +16,8 @@ import ( "time" "github.com/aws/aws-sdk-go-v2/aws" + "github.com/spinloop-ai/spinloop/internal/cloud" "github.com/spinloop-ai/spinloop/internal/inference" - "github.com/spinloop-ai/spinloop/internal/remote" ) // writeDeploySpinloop writes a Spinloop (and optionally a preset) into a temp dir @@ -312,17 +312,17 @@ func TestSplitModelQuant(t *testing.T) { // and the public-IP probe is fixed. Restores on cleanup. func stubDeploySeams(t *testing.T, serverURL, statusState string) { t.Helper() - origDiscover, origStatus, origDetect := deployDiscoverFn, remoteStatusFn, detectPublicCIDRFn + origDiscover, origStatus, origDetect := deployDiscoverFn, cloudStatusFn, detectPublicCIDRFn t.Cleanup(func() { - deployDiscoverFn, remoteStatusFn, detectPublicCIDRFn = origDiscover, origStatus, origDetect + deployDiscoverFn, cloudStatusFn, detectPublicCIDRFn = origDiscover, origStatus, origDetect }) - deployDiscoverFn = func(context.Context, aws.Config, string) (remote.ControlPlane, error) { - return remote.ControlPlane{Config: remote.Config{ + deployDiscoverFn = func(context.Context, aws.Config, string) (cloud.ControlPlane, error) { + return cloud.ControlPlane{Config: cloud.Config{ StartURL: serverURL, StopURL: serverURL, DeployURL: serverURL, Region: "us-east-1", }}, nil } - remoteStatusFn = func(context.Context, remote.Config) (*remote.Response, error) { - return &remote.Response{StatusCode: 200, State: statusState}, nil + cloudStatusFn = func(context.Context, cloud.Config) (*cloud.Response, error) { + return &cloud.Response{StatusCode: 200, State: statusState}, nil } detectPublicCIDRFn = func(context.Context) (string, error) { return "203.0.113.7/32", nil } } @@ -342,7 +342,7 @@ func writeDeploySpinloopCwd(t *testing.T) { } } -func TestRemoteDeploy_PostsTheConfigAndRegisters(t *testing.T) { +func TestCloudDeploy_PostsTheConfigAndRegisters(t *testing.T) { isolateConfig(t) stubAWSEnv(t) @@ -365,8 +365,8 @@ func TestRemoteDeploy_PostsTheConfigAndRegisters(t *testing.T) { writeDeploySpinloopCwd(t) out := captureStdout(t, func() { - if err := cmdRemoteDeploy([]string{"--env", "testenv"}); err != nil { - t.Errorf("cmdRemoteDeploy: %v", err) + if err := cmdCloudDeploy([]string{"--env", "testenv"}); err != nil { + t.Errorf("cmdCloudDeploy: %v", err) } }) @@ -390,21 +390,21 @@ func TestRemoteDeploy_PostsTheConfigAndRegisters(t *testing.T) { } // The seed is named by its stable id, and the output states the command // that follows it rather than quoting an estimate to wait out. - if !strings.Contains(out, "spinloop remote seed status llamacpp--unsloth-Qwen3.6-27B-MTP-GGUF--UD-Q6_K_XL") { + if !strings.Contains(out, "spinloop cloud seed status llamacpp--unsloth-Qwen3.6-27B-MTP-GGUF--UD-Q6_K_XL") { t.Errorf("seeding should name the follow-up command, got:\n%s", out) } // The environment is registered, owner-only, carrying the base URL and id. - path := must1(remote.EnvConfigPath("testenv")) + path := must1(cloud.EnvConfigPath("testenv")) fi, err := os.Stat(path) if err != nil { t.Fatalf("environment not registered: %v", err) } if fi.Mode().Perm() != 0o600 { - t.Errorf("registered remote.json mode = %v, want 0600", fi.Mode().Perm()) + t.Errorf("registered cloud.json mode = %v, want 0600", fi.Mode().Perm()) } data, _ := os.ReadFile(path) - var saved remote.Config + var saved cloud.Config if err := json.Unmarshal(data, &saved); err != nil { t.Fatal(err) } @@ -413,25 +413,25 @@ func TestRemoteDeploy_PostsTheConfigAndRegisters(t *testing.T) { } } -func TestRemoteDeploy_NotBootstrapped(t *testing.T) { +func TestCloudDeploy_NotBootstrapped(t *testing.T) { isolateConfig(t) stubAWSEnv(t) stubDeploySeams(t, "https://unused", "undeployed") - deployDiscoverFn = func(context.Context, aws.Config, string) (remote.ControlPlane, error) { - return remote.ControlPlane{}, fmt.Errorf("the control plane (stack %q) is not deployed in this account and region — run `spinloop remote bootstrap` first", "cloud-vm-llm") + deployDiscoverFn = func(context.Context, aws.Config, string) (cloud.ControlPlane, error) { + return cloud.ControlPlane{}, fmt.Errorf("the control plane (stack %q) is not deployed in this account and region — run `spinloop cloud bootstrap` first", "cloud-vm-llm") } writeDeploySpinloopCwd(t) - err := cmdRemoteDeploy([]string{"--env", "testenv"}) - if err == nil || !strings.Contains(err.Error(), "spinloop remote bootstrap") { + err := cmdCloudDeploy([]string{"--env", "testenv"}) + if err == nil || !strings.Contains(err.Error(), "spinloop cloud bootstrap") { t.Errorf("want a bootstrap-first error, got %v", err) } - if _, statErr := os.Stat(must1(remote.EnvConfigPath("testenv"))); !os.IsNotExist(statErr) { + if _, statErr := os.Stat(must1(cloud.EnvConfigPath("testenv"))); !os.IsNotExist(statErr) { t.Error("nothing should be registered when the account is not bootstrapped") } } -func TestRemoteDeploy_RequiresEnvName(t *testing.T) { +func TestCloudDeploy_RequiresEnvName(t *testing.T) { isolateConfig(t) t.Chdir(t.TempDir()) if err := os.WriteFile("Spinloop", []byte( @@ -440,29 +440,29 @@ func TestRemoteDeploy_RequiresEnvName(t *testing.T) { t.Fatal(err) } // No --env: deploy cannot tell which environment to create. - err := cmdRemoteDeploy(nil) + err := cmdCloudDeploy(nil) if err == nil || !strings.Contains(err.Error(), "deploy must name the environment it creates") { t.Errorf("want an error asking for --env, got %v", err) } // A path-shaped value is not an environment name. - err = cmdRemoteDeploy([]string{"--env", "./remote.json"}) + err = cmdCloudDeploy([]string{"--env", "./cloud.json"}) if err == nil || !strings.Contains(err.Error(), "is not an environment name") { t.Errorf("want an environment-name error, got %v", err) } } -func TestRemoteDeploy_OverwriteGuard(t *testing.T) { +func TestCloudDeploy_OverwriteGuard(t *testing.T) { deployBody := `{"deployed":true,"environment":"testenv","base_url":"http://198.51.100.9:8000/v1"}` t.Run("registered environment needs --overwrite", func(t *testing.T) { isolateConfig(t) stubAWSEnv(t) stubDeploySeams(t, "https://unused", "undeployed") - if err := remote.SaveEnvironment("testenv", remote.Config{StartURL: "https://s", StopURL: "https://x", Region: "us-east-1"}); err != nil { + if err := cloud.SaveEnvironment("testenv", cloud.Config{StartURL: "https://s", StopURL: "https://x", Region: "us-east-1"}); err != nil { t.Fatal(err) } writeDeploySpinloopCwd(t) - err := cmdRemoteDeploy([]string{"--env", "testenv"}) + err := cmdCloudDeploy([]string{"--env", "testenv"}) if err == nil || !strings.Contains(err.Error(), "--overwrite") { t.Errorf("want an overwrite refusal, got %v", err) } @@ -473,7 +473,7 @@ func TestRemoteDeploy_OverwriteGuard(t *testing.T) { stubAWSEnv(t) stubDeploySeams(t, "https://unused", "running") writeDeploySpinloopCwd(t) - err := cmdRemoteDeploy([]string{"--env", "testenv"}) + err := cmdCloudDeploy([]string{"--env", "testenv"}) if err == nil || !strings.Contains(err.Error(), "--overwrite") { t.Errorf("want an overwrite refusal for a live instance, got %v", err) } @@ -490,8 +490,8 @@ func TestRemoteDeploy_OverwriteGuard(t *testing.T) { stubDeploySeams(t, server.URL, "running") writeDeploySpinloopCwd(t) out := captureStdout(t, func() { - if err := cmdRemoteDeploy([]string{"--env", "testenv", "--overwrite"}); err != nil { - t.Errorf("cmdRemoteDeploy --overwrite: %v", err) + if err := cmdCloudDeploy([]string{"--env", "testenv", "--overwrite"}); err != nil { + t.Errorf("cmdCloudDeploy --overwrite: %v", err) } }) if !strings.Contains(out, "deployed: environment testenv") { @@ -500,30 +500,30 @@ func TestRemoteDeploy_OverwriteGuard(t *testing.T) { }) } -func TestRemoteDeploy_RejectsBadCIDR(t *testing.T) { +func TestCloudDeploy_RejectsBadCIDR(t *testing.T) { isolateConfig(t) stubDeploySeams(t, "https://unused", "undeployed") writeDeploySpinloopCwd(t) - err := cmdRemoteDeploy([]string{"--env", "testenv", "--allowed-cidr", "bogus"}) + err := cmdCloudDeploy([]string{"--env", "testenv", "--allowed-cidr", "bogus"}) if err == nil || !strings.Contains(err.Error(), "IPv4 CIDR") { t.Errorf("want a CIDR validation error, got %v", err) } } -func TestRemoteDeploy_DryRunSendsNothing(t *testing.T) { +func TestCloudDeploy_DryRunSendsNothing(t *testing.T) { isolateConfig(t) called := false origDiscover := deployDiscoverFn t.Cleanup(func() { deployDiscoverFn = origDiscover }) - deployDiscoverFn = func(context.Context, aws.Config, string) (remote.ControlPlane, error) { + deployDiscoverFn = func(context.Context, aws.Config, string) (cloud.ControlPlane, error) { called = true - return remote.ControlPlane{}, fmt.Errorf("must not be called") + return cloud.ControlPlane{}, fmt.Errorf("must not be called") } writeDeploySpinloopCwd(t) out := captureStdout(t, func() { - if err := cmdRemoteDeploy([]string{"--env", "testenv", "--dry-run"}); err != nil { - t.Errorf("cmdRemoteDeploy --dry-run: %v", err) + if err := cmdCloudDeploy([]string{"--env", "testenv", "--dry-run"}); err != nil { + t.Errorf("cmdCloudDeploy --dry-run: %v", err) } }) if called { @@ -539,7 +539,7 @@ func TestRemoteDeploy_DryRunSendsNothing(t *testing.T) { } } -func TestRemoteDeploy_SpinloopVersion(t *testing.T) { +func TestCloudDeploy_SpinloopVersion(t *testing.T) { t.Run("the pin reaches the deploy body and the plan", func(t *testing.T) { isolateConfig(t) stubAWSEnv(t) @@ -557,8 +557,8 @@ func TestRemoteDeploy_SpinloopVersion(t *testing.T) { writeDeploySpinloopCwd(t) out := captureStdout(t, func() { - if err := cmdRemoteDeploy([]string{"--env", "testenv", "--spinloop-version", "v1.26.1"}); err != nil { - t.Errorf("cmdRemoteDeploy: %v", err) + if err := cmdCloudDeploy([]string{"--env", "testenv", "--spinloop-version", "v1.26.1"}); err != nil { + t.Errorf("cmdCloudDeploy: %v", err) } }) @@ -589,8 +589,8 @@ func TestRemoteDeploy_SpinloopVersion(t *testing.T) { stubDeploySeams(t, server.URL, "undeployed") writeDeploySpinloopCwd(t) - if err := cmdRemoteDeploy([]string{"--env", "testenv"}); err != nil { - t.Errorf("cmdRemoteDeploy: %v", err) + if err := cmdCloudDeploy([]string{"--env", "testenv"}); err != nil { + t.Errorf("cmdCloudDeploy: %v", err) } // Absent, not null and not "latest": an unpinned deploy sends exactly // what a control plane predating the pin expects. @@ -604,8 +604,8 @@ func TestRemoteDeploy_SpinloopVersion(t *testing.T) { for _, flag := range []string{"latest", "v"} { writeDeploySpinloopCwd(t) out := captureStdout(t, func() { - if err := cmdRemoteDeploy([]string{"--env", "testenv", "--spinloop-version", flag, "--dry-run"}); err != nil { - t.Errorf("cmdRemoteDeploy --spinloop-version %s: %v", flag, err) + if err := cmdCloudDeploy([]string{"--env", "testenv", "--spinloop-version", flag, "--dry-run"}); err != nil { + t.Errorf("cmdCloudDeploy --spinloop-version %s: %v", flag, err) } }) if !strings.Contains(out, "spinloop: latest") { @@ -618,7 +618,7 @@ func TestRemoteDeploy_SpinloopVersion(t *testing.T) { isolateConfig(t) stubDeploySeams(t, "https://unused", "undeployed") writeDeploySpinloopCwd(t) - err := cmdRemoteDeploy([]string{"--env", "testenv", "--spinloop-version", "1.2.6 beta"}) + err := cmdCloudDeploy([]string{"--env", "testenv", "--spinloop-version", "1.2.6 beta"}) if err == nil || !strings.Contains(err.Error(), "--spinloop-version") { t.Errorf("want a --spinloop-version validation error, got %v", err) } @@ -628,7 +628,7 @@ func TestRemoteDeploy_SpinloopVersion(t *testing.T) { // An instance type is recorded on the environment: it reaches the signed body // so the deploy Lambda persists it, and the plan names the machine the // environment will launch as — a promise when named, the default otherwise. -func TestRemoteDeploy_InstanceType(t *testing.T) { +func TestCloudDeploy_InstanceType(t *testing.T) { t.Run("the type reaches the deploy body and the plan", func(t *testing.T) { isolateConfig(t) stubAWSEnv(t) @@ -646,8 +646,8 @@ func TestRemoteDeploy_InstanceType(t *testing.T) { writeDeploySpinloopCwd(t) out := captureStdout(t, func() { - if err := cmdRemoteDeploy([]string{"--env", "testenv", "--instance-type", "g6e.2xlarge"}); err != nil { - t.Errorf("cmdRemoteDeploy: %v", err) + if err := cmdCloudDeploy([]string{"--env", "testenv", "--instance-type", "g6e.2xlarge"}); err != nil { + t.Errorf("cmdCloudDeploy: %v", err) } }) @@ -676,8 +676,8 @@ func TestRemoteDeploy_InstanceType(t *testing.T) { stubDeploySeams(t, server.URL, "undeployed") writeDeploySpinloopCwd(t) - if err := cmdRemoteDeploy([]string{"--env", "testenv"}); err != nil { - t.Errorf("cmdRemoteDeploy: %v", err) + if err := cmdCloudDeploy([]string{"--env", "testenv"}); err != nil { + t.Errorf("cmdCloudDeploy: %v", err) } // Absent, not null and not empty: an untyped deploy sends exactly what // a control plane predating the field expects. @@ -690,8 +690,8 @@ func TestRemoteDeploy_InstanceType(t *testing.T) { isolateConfig(t) writeDeploySpinloopCwd(t) out := captureStdout(t, func() { - if err := cmdRemoteDeploy([]string{"--env", "testenv", "--dry-run"}); err != nil { - t.Errorf("cmdRemoteDeploy --dry-run: %v", err) + if err := cmdCloudDeploy([]string{"--env", "testenv", "--dry-run"}); err != nil { + t.Errorf("cmdCloudDeploy --dry-run: %v", err) } }) if !strings.Contains(out, "instance: the control plane's default") { @@ -703,7 +703,7 @@ func TestRemoteDeploy_InstanceType(t *testing.T) { isolateConfig(t) stubDeploySeams(t, "https://unused", "undeployed") writeDeploySpinloopCwd(t) - err := cmdRemoteDeploy([]string{"--env", "testenv", "--instance-type", "g6exlarge"}) + err := cmdCloudDeploy([]string{"--env", "testenv", "--instance-type", "g6exlarge"}) if err == nil || !strings.Contains(err.Error(), "--instance-type") { t.Errorf("want a --instance-type validation error, got %v", err) } @@ -713,7 +713,7 @@ func TestRemoteDeploy_InstanceType(t *testing.T) { // A supplied key is resolved from the environment the Spinloop's local // environment populated, reaches the signed body as a request-scoped field, // and the report says what happened to it — the action, never the value. -func TestRemoteDeploy_APIKeyEnvReachesTheRequest(t *testing.T) { +func TestCloudDeploy_APIKeyEnvReachesTheRequest(t *testing.T) { isolateConfig(t) stubAWSEnv(t) @@ -734,8 +734,8 @@ func TestRemoteDeploy_APIKeyEnvReachesTheRequest(t *testing.T) { t.Setenv("SHARED_KEY", "sk-supplied") out := captureStdout(t, func() { - if err := cmdRemoteDeploy([]string{"--env", "testenv", "--api-key-env", "SHARED_KEY"}); err != nil { - t.Fatalf("cmdRemoteDeploy: %v", err) + if err := cmdCloudDeploy([]string{"--env", "testenv", "--api-key-env", "SHARED_KEY"}); err != nil { + t.Fatalf("cmdCloudDeploy: %v", err) } }) @@ -752,7 +752,7 @@ func TestRemoteDeploy_APIKeyEnvReachesTheRequest(t *testing.T) { } // A named variable that is set nowhere fails before anything is sent. -func TestRemoteDeploy_APIKeyEnvUnsetFailsBeforeSending(t *testing.T) { +func TestCloudDeploy_APIKeyEnvUnsetFailsBeforeSending(t *testing.T) { isolateConfig(t) stubAWSEnv(t) @@ -766,7 +766,7 @@ func TestRemoteDeploy_APIKeyEnvUnsetFailsBeforeSending(t *testing.T) { writeDeploySpinloopCwd(t) os.Unsetenv("NOWHERE_DEPLOY_KEY") - err := cmdRemoteDeploy([]string{"--env", "testenv", "--api-key-env", "NOWHERE_DEPLOY_KEY"}) + err := cmdCloudDeploy([]string{"--env", "testenv", "--api-key-env", "NOWHERE_DEPLOY_KEY"}) if err == nil || !strings.Contains(err.Error(), "NOWHERE_DEPLOY_KEY") { t.Errorf("want an error naming the variable, got %v", err) } @@ -777,21 +777,21 @@ func TestRemoteDeploy_APIKeyEnvUnsetFailsBeforeSending(t *testing.T) { // A dry run with a key names the variable it will store from — never the // value — and still touches nothing. -func TestRemoteDeploy_APIKeyEnvDryRunNamesTheVariable(t *testing.T) { +func TestCloudDeploy_APIKeyEnvDryRunNamesTheVariable(t *testing.T) { isolateConfig(t) called := false origDiscover := deployDiscoverFn t.Cleanup(func() { deployDiscoverFn = origDiscover }) - deployDiscoverFn = func(context.Context, aws.Config, string) (remote.ControlPlane, error) { + deployDiscoverFn = func(context.Context, aws.Config, string) (cloud.ControlPlane, error) { called = true - return remote.ControlPlane{}, fmt.Errorf("must not be called") + return cloud.ControlPlane{}, fmt.Errorf("must not be called") } writeDeploySpinloopCwd(t) t.Setenv("SHARED_KEY", "sk-supplied") out := captureStdout(t, func() { - if err := cmdRemoteDeploy([]string{"--env", "testenv", "--dry-run", "--api-key-env", "SHARED_KEY"}); err != nil { - t.Fatalf("cmdRemoteDeploy --dry-run: %v", err) + if err := cmdCloudDeploy([]string{"--env", "testenv", "--dry-run", "--api-key-env", "SHARED_KEY"}); err != nil { + t.Fatalf("cmdCloudDeploy --dry-run: %v", err) } }) if called { @@ -840,7 +840,7 @@ func TestProviderForRunner(t *testing.T) { // A cold start blocks in one request for minutes, so the command must say what // it is doing rather than sit silent — and must say it on stderr, so piping the // exports still works. -func TestRemoteStart_ReportsProgressWhileWaiting(t *testing.T) { +func TestCloudStart_ReportsProgressWhileWaiting(t *testing.T) { isolateConfig(t) stubAWSEnv(t) @@ -850,11 +850,11 @@ func TestRemoteStart_ReportsProgressWhileWaiting(t *testing.T) { w.Write([]byte(`{"state":"ready","base_url":"http://198.51.100.1:8000/v1","api_key":"sk-test"}`)) })) defer server.Close() - writeRemoteConfig(t, server.URL) + writeCloudConfig(t, server.URL) stderr := captureStderr(t, func() { - if err := cmdRemoteStart([]string{"--env", "default"}); err != nil { - t.Errorf("cmdRemoteStart: %v", err) + if err := cmdCloudStart([]string{"--env", "default"}); err != nil { + t.Errorf("cmdCloudStart: %v", err) } }) @@ -888,7 +888,7 @@ func TestStartProgress_HeartbeatsAndStops(t *testing.T) { // coming up. It renders the phase the start is in, so the wording is the tile's // wording and the numbers in it move as the heartbeat repeats. func TestStartProgress_HeartbeatReflectsPhase(t *testing.T) { - // driveStart feeds one start's callbacks the way remote.Start would. + // driveStart feeds one start's callbacks the way cloud.Start would. driveStart := func(p *startProgress, do func(progress, onState func(string))) { progress, onState := p.callbacks() do(progress, onState) @@ -940,7 +940,7 @@ func TestStartProgress_HeartbeatReflectsPhase(t *testing.T) { if got := p.heartbeat(); !strings.Contains(got, "waiting for capacity") { t.Errorf("after no-capacity, heartbeat = %q, want a capacity wait", got) } - onState(remote.StateInFlight) + onState(cloud.StateInFlight) onState("starting") if got := p.heartbeat(); !strings.Contains(got, "booting") { t.Errorf("after booting, heartbeat = %q, want a booting line", got) @@ -958,7 +958,7 @@ func TestStartProgress_HeartbeatReflectsPhase(t *testing.T) { if got := p.heartbeat(); !strings.Contains(got, "waiting for capacity") { t.Errorf("after no-capacity, heartbeat = %q, want a capacity wait", got) } - onState(remote.StateInFlight) + onState(cloud.StateInFlight) got := p.heartbeat() if !strings.Contains(got, "waking the instance") { t.Errorf("in-flight after a capacity wait, heartbeat = %q, want the attempt", got) @@ -975,7 +975,7 @@ func TestStartProgress_HeartbeatReflectsPhase(t *testing.T) { // the instance boots. The output must say it is waiting for capacity during the // wait and stop saying so once the booting attempt is in flight — the refused // attempt's report must not outlive the attempt it described. -func TestRemoteStart_HeartbeatTracksTheCapacityWaitEnding(t *testing.T) { +func TestCloudStart_HeartbeatTracksTheCapacityWaitEnding(t *testing.T) { isolateConfig(t) stubAWSEnv(t) @@ -983,9 +983,9 @@ func TestRemoteStart_HeartbeatTracksTheCapacityWaitEnding(t *testing.T) { heartbeatEvery = 20 * time.Millisecond t.Cleanup(func() { heartbeatEvery = origEvery }) - origProbe := remote.ProbeTimeout - remote.ProbeTimeout = 100 * time.Millisecond - t.Cleanup(func() { remote.ProbeTimeout = origProbe }) + origProbe := cloud.ProbeTimeout + cloud.ProbeTimeout = 100 * time.Millisecond + t.Cleanup(func() { cloud.ProbeTimeout = origProbe }) // A reachable endpoint, so the post-ready probe adds no noise. l, err := net.Listen("tcp", "127.0.0.1:0") @@ -1019,11 +1019,11 @@ func TestRemoteStart_HeartbeatTracksTheCapacityWaitEnding(t *testing.T) { w.Write([]byte(fmt.Sprintf(`{"state":"ready","base_url":"%s","api_key":"sk-test"}`, baseURL))) })) defer server.Close() - writeRemoteConfig(t, server.URL) + writeCloudConfig(t, server.URL) stderr := captureStderr(t, func() { - if err := cmdRemoteStart([]string{"--env", "default"}); err != nil { - t.Fatalf("cmdRemoteStart: %v", err) + if err := cmdCloudStart([]string{"--env", "default"}); err != nil { + t.Fatalf("cmdCloudStart: %v", err) } }) @@ -1050,7 +1050,7 @@ func TestRemoteStart_HeartbeatTracksTheCapacityWaitEnding(t *testing.T) { // -t is a shorthand for --timeout: a very short one makes a never-ready start // give up promptly rather than block for the 15m default. -func TestRemoteStart_TimeoutShorthand(t *testing.T) { +func TestCloudStart_TimeoutShorthand(t *testing.T) { isolateConfig(t) stubAWSEnv(t) // Never returns ready, so only the timeout ends the wait. @@ -1060,7 +1060,7 @@ func TestRemoteStart_TimeoutShorthand(t *testing.T) { w.Write([]byte(`{"state":"starting","retry_after_seconds":0}`)) })) defer server.Close() - writeRemoteConfig(t, server.URL) + writeCloudConfig(t, server.URL) // Silence the progress line this writes to stderr; the wait is what matters. oldStderr := os.Stderr @@ -1069,7 +1069,7 @@ func TestRemoteStart_TimeoutShorthand(t *testing.T) { defer func() { os.Stderr = oldStderr; w.Close() }() done := make(chan error, 1) - go func() { done <- cmdRemoteStart([]string{"--env", "default", "-t", "80ms"}) }() + go func() { done <- cmdCloudStart([]string{"--env", "default", "-t", "80ms"}) }() select { case err := <-done: if err == nil { @@ -1080,9 +1080,9 @@ func TestRemoteStart_TimeoutShorthand(t *testing.T) { } } -// Progress must not land on stdout, or `spinloop remote start | grep '^export '` +// Progress must not land on stdout, or `spinloop cloud start | grep '^export '` // would pick it up. -func TestRemoteStart_StdoutCarriesOnlyTheResult(t *testing.T) { +func TestCloudStart_StdoutCarriesOnlyTheResult(t *testing.T) { isolateConfig(t) stubAWSEnv(t) server := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) { @@ -1090,11 +1090,11 @@ func TestRemoteStart_StdoutCarriesOnlyTheResult(t *testing.T) { w.Write([]byte(`{"state":"ready","base_url":"http://198.51.100.1:8000/v1","api_key":"sk-test"}`)) })) defer server.Close() - writeRemoteConfig(t, server.URL) + writeCloudConfig(t, server.URL) out := captureStdout(t, func() { - if err := cmdRemoteStart([]string{"--env", "default", "--print-env"}); err != nil { - t.Errorf("cmdRemoteStart: %v", err) + if err := cmdCloudStart([]string{"--env", "default", "--print-env"}); err != nil { + t.Errorf("cmdCloudStart: %v", err) } }) lines := strings.Split(strings.TrimSpace(out), "\n") diff --git a/cmd/spinloop/remote_env_test.go b/cmd/spinloop/cloud_env_test.go similarity index 79% rename from cmd/spinloop/remote_env_test.go rename to cmd/spinloop/cloud_env_test.go index f2f70a7a..3075cd57 100644 --- a/cmd/spinloop/remote_env_test.go +++ b/cmd/spinloop/cloud_env_test.go @@ -9,7 +9,7 @@ import ( "strings" "testing" - "github.com/spinloop-ai/spinloop/internal/remote" + "github.com/spinloop-ai/spinloop/internal/cloud" "github.com/spinloop-ai/spinloop/internal/spinloop" ) @@ -98,14 +98,14 @@ func hitRecorder(name string, hit chan<- string) http.HandlerFunc { } } -// A .env beside the Spinloop reaches a control command: its SPINLOOP_REMOTE_STOP_URL -// override wins over the remote.json value, so the .env server is the one hit. -func TestRemoteStop_RespectsDotEnv(t *testing.T) { +// A .env beside the Spinloop reaches a control command: its SPINLOOP_CLOUD_STOP_URL +// override wins over the cloud.json value, so the .env server is the one hit. +func TestCloudStop_RespectsDotEnv(t *testing.T) { isolateConfig(t) stubAWSEnv(t) hit := make(chan string, 1) - fromConfig := httptest.NewServer(hitRecorder("remote.json", hit)) + fromConfig := httptest.NewServer(hitRecorder("cloud.json", hit)) defer fromConfig.Close() fromDotEnv := httptest.NewServer(hitRecorder("dotenv", hit)) defer fromDotEnv.Close() @@ -114,17 +114,17 @@ func TestRemoteStop_RespectsDotEnv(t *testing.T) { if err := os.WriteFile("Spinloop", []byte("PROVIDER openai-compatible\n"), 0o600); err != nil { t.Fatal(err) } - if err := remote.SaveEnvironment("testenv", remote.Config{StartURL: fromConfig.URL, StopURL: fromConfig.URL, Region: "eu-west-1"}); err != nil { + if err := cloud.SaveEnvironment("testenv", cloud.Config{StartURL: fromConfig.URL, StopURL: fromConfig.URL, Region: "eu-west-1"}); err != nil { t.Fatal(err) } - if err := os.WriteFile(".env", []byte("SPINLOOP_REMOTE_STOP_URL="+fromDotEnv.URL+"\n"), 0o600); err != nil { + if err := os.WriteFile(".env", []byte("SPINLOOP_CLOUD_STOP_URL="+fromDotEnv.URL+"\n"), 0o600); err != nil { t.Fatal(err) } - unsetEnvOnCleanup(t, "SPINLOOP_REMOTE_STOP_URL") + unsetEnvOnCleanup(t, "SPINLOOP_CLOUD_STOP_URL") - // An explicit Spinloop path, exercising resolveRemoteConfig's explicit-arg branch. - if err := cmdRemoteStop([]string{"Spinloop", "--env", "testenv"}); err != nil { - t.Fatalf("cmdRemoteStop: %v", err) + // An explicit Spinloop path, exercising resolveCloudConfig's explicit-arg branch. + if err := cmdCloudStop([]string{"Spinloop", "--env", "testenv"}); err != nil { + t.Fatalf("cmdCloudStop: %v", err) } select { case name := <-hit: @@ -137,8 +137,8 @@ func TestRemoteStop_RespectsDotEnv(t *testing.T) { } // An ENV instruction overrides the .env end-to-end: with both setting -// SPINLOOP_REMOTE_STOP_URL, the ENV server is the one hit. -func TestRemoteStop_EnvKeywordOverridesDotEnv(t *testing.T) { +// SPINLOOP_CLOUD_STOP_URL, the ENV server is the one hit. +func TestCloudStop_EnvKeywordOverridesDotEnv(t *testing.T) { isolateConfig(t) stubAWSEnv(t) @@ -150,21 +150,21 @@ func TestRemoteStop_EnvKeywordOverridesDotEnv(t *testing.T) { t.Chdir(t.TempDir()) spinloopBody := "PROVIDER openai-compatible\n" + - "ENV SPINLOOP_REMOTE_STOP_URL=" + fromEnvKw.URL + "\n" + "ENV SPINLOOP_CLOUD_STOP_URL=" + fromEnvKw.URL + "\n" if err := os.WriteFile("Spinloop", []byte(spinloopBody), 0o600); err != nil { t.Fatal(err) } // The environment still supplies the required StartURL; StopURL is overridden. - if err := remote.SaveEnvironment("testenv", remote.Config{StartURL: fromDotEnv.URL, StopURL: fromDotEnv.URL, Region: "eu-west-1"}); err != nil { + if err := cloud.SaveEnvironment("testenv", cloud.Config{StartURL: fromDotEnv.URL, StopURL: fromDotEnv.URL, Region: "eu-west-1"}); err != nil { t.Fatal(err) } - if err := os.WriteFile(".env", []byte("SPINLOOP_REMOTE_STOP_URL="+fromDotEnv.URL+"\n"), 0o600); err != nil { + if err := os.WriteFile(".env", []byte("SPINLOOP_CLOUD_STOP_URL="+fromDotEnv.URL+"\n"), 0o600); err != nil { t.Fatal(err) } - unsetEnvOnCleanup(t, "SPINLOOP_REMOTE_STOP_URL") + unsetEnvOnCleanup(t, "SPINLOOP_CLOUD_STOP_URL") - if err := cmdRemoteStop([]string{"--env", "testenv"}); err != nil { - t.Fatalf("cmdRemoteStop: %v", err) + if err := cmdCloudStop([]string{"--env", "testenv"}); err != nil { + t.Fatalf("cmdCloudStop: %v", err) } select { case name := <-hit: @@ -178,7 +178,7 @@ func TestRemoteStop_EnvKeywordOverridesDotEnv(t *testing.T) { // The local-only guarantee: ENV and .env values shape the deploying process's // own environment but never enter the payload sent to the instance. -func TestRemoteDeploy_DoesNotForwardEnvToInstance(t *testing.T) { +func TestCloudDeploy_DoesNotForwardEnvToInstance(t *testing.T) { isolateConfig(t) stubAWSEnv(t) @@ -205,8 +205,8 @@ func TestRemoteDeploy_DoesNotForwardEnvToInstance(t *testing.T) { } unsetEnvOnCleanup(t, "SPINLOOP_SECRET_TOKEN", "SPINLOOP_DOTENV_SECRET") - if err := cmdRemoteDeploy([]string{"--env", "testenv"}); err != nil { - t.Fatalf("cmdRemoteDeploy: %v", err) + if err := cmdCloudDeploy([]string{"--env", "testenv"}); err != nil { + t.Fatalf("cmdCloudDeploy: %v", err) } for _, leak := range []string{"SPINLOOP_SECRET_TOKEN", "do-not-leak", "SPINLOOP_DOTENV_SECRET", "also-no"} { if strings.Contains(string(body), leak) { diff --git a/cmd/spinloop/remote_environments_test.go b/cmd/spinloop/cloud_environments_test.go similarity index 64% rename from cmd/spinloop/remote_environments_test.go rename to cmd/spinloop/cloud_environments_test.go index 14d266b1..66728410 100644 --- a/cmd/spinloop/remote_environments_test.go +++ b/cmd/spinloop/cloud_environments_test.go @@ -9,21 +9,21 @@ import ( "strings" "testing" + "github.com/spinloop-ai/spinloop/internal/cloud" "github.com/spinloop-ai/spinloop/internal/config" - "github.com/spinloop-ai/spinloop/internal/remote" ) -// registerEnv writes an environment's remote.json into the registry. -func registerEnv(t *testing.T, name string, cfg remote.Config) { +// registerEnv writes an environment's cloud.json into the registry. +func registerEnv(t *testing.T, name string, cfg cloud.Config) { t.Helper() - if err := os.MkdirAll(must1(remote.EnvDir(name)), 0o700); err != nil { + if err := os.MkdirAll(must1(cloud.EnvDir(name)), 0o700); err != nil { t.Fatal(err) } data, err := json.Marshal(cfg) if err != nil { t.Fatal(err) } - if err := os.WriteFile(must1(remote.EnvConfigPath(name)), data, 0o600); err != nil { + if err := os.WriteFile(must1(cloud.EnvConfigPath(name)), data, 0o600); err != nil { t.Fatal(err) } } @@ -37,12 +37,12 @@ func stateServer(t *testing.T) *httptest.Server { } // A bare environment name on --env resolves through the registry. -func TestRemote_EnvNameResolves(t *testing.T) { +func TestCloud_EnvNameResolves(t *testing.T) { isolateConfig(t) stubAWSEnv(t) server := stateServer(t) defer server.Close() - registerEnv(t, "prodenv", remote.Config{StartURL: server.URL, StopURL: server.URL, Region: "eu-west-1"}) + registerEnv(t, "prodenv", cloud.Config{StartURL: server.URL, StopURL: server.URL, Region: "eu-west-1"}) t.Chdir(t.TempDir()) if err := os.WriteFile("Spinloop", []byte("PROVIDER openai-compatible\n"), 0o600); err != nil { @@ -59,12 +59,12 @@ func TestRemote_EnvNameResolves(t *testing.T) { } // With no Spinloop in play, the default environment is used. -func TestRemote_DefaultEnvironment(t *testing.T) { +func TestCloud_DefaultEnvironment(t *testing.T) { isolateConfig(t) stubAWSEnv(t) server := stateServer(t) defer server.Close() - registerEnv(t, "default", remote.Config{StartURL: server.URL, StopURL: server.URL, Region: "eu-west-1"}) + registerEnv(t, "default", cloud.Config{StartURL: server.URL, StopURL: server.URL, Region: "eu-west-1"}) t.Chdir(t.TempDir()) // no ./Spinloop here out := captureStdout(t, func() { @@ -79,18 +79,18 @@ func TestRemote_DefaultEnvironment(t *testing.T) { // A file at the superseded path configures nothing: no path outside the // registry is read for any name. -func TestRemote_SupersededFileIsNotRead(t *testing.T) { +func TestCloud_SupersededFileIsNotRead(t *testing.T) { isolateConfig(t) stubAWSEnv(t) server := stateServer(t) defer server.Close() - data, _ := json.Marshal(remote.Config{StartURL: server.URL, StopURL: server.URL, Region: "eu-west-1"}) + data, _ := json.Marshal(cloud.Config{StartURL: server.URL, StopURL: server.URL, Region: "eu-west-1"}) home := must1(config.Dir()) if err := os.MkdirAll(home, 0o700); err != nil { t.Fatal(err) } - if err := os.WriteFile(filepath.Join(home, "remote.json"), data, 0o600); err != nil { + if err := os.WriteFile(filepath.Join(home, "cloud.json"), data, 0o600); err != nil { t.Fatal(err) } t.Chdir(t.TempDir()) @@ -98,34 +98,34 @@ func TestRemote_SupersededFileIsNotRead(t *testing.T) { if err == nil { t.Fatal("the superseded file must not configure an environment") } - if !strings.Contains(err.Error(), "remotes/default/remote.json") { + if !strings.Contains(err.Error(), "clouds/default/cloud.json") { t.Errorf("the failure should name the registry path, got %v", err) } } -func TestRemoteList(t *testing.T) { +func TestCloudList(t *testing.T) { isolateConfig(t) // Empty registry says so. out := captureStdout(t, func() { - if err := cmdRemoteList(nil); err != nil { + if err := cmdCloudList(nil); err != nil { t.Errorf("ls empty: %v", err) } }) - if !strings.Contains(out, "No remote environments") { + if !strings.Contains(out, "No cloud environments") { t.Errorf("empty ls should say so, got:\n%s", out) } - registerEnv(t, "prod", remote.Config{StartURL: "https://s", StopURL: "https://x", Region: "eu-west-1", BaseURL: "http://1.2.3.4:8000/v1"}) - if err := os.MkdirAll(must1(remote.EnvDir("broken")), 0o700); err != nil { + registerEnv(t, "prod", cloud.Config{StartURL: "https://s", StopURL: "https://x", Region: "eu-west-1", BaseURL: "http://1.2.3.4:8000/v1"}) + if err := os.MkdirAll(must1(cloud.EnvDir("broken")), 0o700); err != nil { t.Fatal(err) } - if err := os.WriteFile(must1(remote.EnvConfigPath("broken")), []byte("nope"), 0o600); err != nil { + if err := os.WriteFile(must1(cloud.EnvConfigPath("broken")), []byte("nope"), 0o600); err != nil { t.Fatal(err) } out = captureStdout(t, func() { - if err := cmdRemoteList(nil); err != nil { + if err := cmdCloudList(nil); err != nil { t.Errorf("ls: %v", err) } }) diff --git a/cmd/spinloop/remote_logs_test.go b/cmd/spinloop/cloud_logs_test.go similarity index 76% rename from cmd/spinloop/remote_logs_test.go rename to cmd/spinloop/cloud_logs_test.go index 043785de..f4f55f3e 100644 --- a/cmd/spinloop/remote_logs_test.go +++ b/cmd/spinloop/cloud_logs_test.go @@ -5,28 +5,28 @@ import ( "testing" "time" + "github.com/spinloop-ai/spinloop/internal/cloud" "github.com/spinloop-ai/spinloop/internal/fleet" - "github.com/spinloop-ai/spinloop/internal/remote" ) // stubTopLevelLogsFetch substitutes the log store read for the duration of a // test, handing the stub each query so it can assert on what the command // asked for. The read goes through fleet.FetchLogsFn rather than an HTTP // server: CloudWatch is reached through the AWS SDK, not a Function URL. -func stubTopLevelLogsFetch(t *testing.T, fn func(q remote.LogQuery) (remote.LogResult, error)) { +func stubTopLevelLogsFetch(t *testing.T, fn func(q cloud.LogQuery) (cloud.LogResult, error)) { t.Helper() prev := fleet.FetchLogsFn - fleet.FetchLogsFn = func(_ context.Context, _ remote.Config, q remote.LogQuery) (remote.LogResult, error) { + fleet.FetchLogsFn = func(_ context.Context, _ cloud.Config, q cloud.LogQuery) (cloud.LogResult, error) { return fn(q) } t.Cleanup(func() { fleet.FetchLogsFn = prev }) } func TestRunnerForAcceptsExactlyTheRunnersWithLogGroups(t *testing.T) { - for _, runner := range remote.Runners { + for _, runner := range cloud.Runners { if _, err := runnerFor(runner); err != nil { t.Errorf("deploy rejects runner %q, but its engine log group %q is read: %v", - runner, remote.EngineLogGroup(runner), err) + runner, cloud.EngineLogGroup(runner), err) } } // The reverse direction: a runner deploy accepts must have a group read for @@ -37,13 +37,13 @@ func TestRunnerForAcceptsExactlyTheRunnersWithLogGroups(t *testing.T) { t.Fatalf("runnerFor(%q): %v", provider, err) } found := false - for _, known := range remote.Runners { + for _, known := range cloud.Runners { if known == runner { found = true } } if !found { - t.Errorf("runner %q can be deployed but has no engine log group in remote.Runners", runner) + t.Errorf("runner %q can be deployed but has no engine log group in cloud.Runners", runner) } } } @@ -53,22 +53,22 @@ func TestRunnerForAcceptsExactlyTheRunnersWithLogGroups(t *testing.T) { // log store query. func TestTopLevelLogsDefaultsToTheEngineSourceAndAnHourWindow(t *testing.T) { isolateConfig(t) - registerEnv(t, "prod", remote.Config{ + registerEnv(t, "prod", cloud.Config{ StartURL: "https://start.lambda-url.eu-west-1.on.aws/", StopURL: "https://stop.lambda-url.eu-west-1.on.aws/", Region: "eu-west-1", }) - var got remote.LogQuery - stubTopLevelLogsFetch(t, func(q remote.LogQuery) (remote.LogResult, error) { + var got cloud.LogQuery + stubTopLevelLogsFetch(t, func(q cloud.LogQuery) (cloud.LogResult, error) { got = q - return remote.LogResult{}, nil + return cloud.LogResult{}, nil }) if err := cmdLogs([]string{"--env", "prod"}); err != nil { t.Fatal(err) } - if got.Source != remote.LogSourceEngine { + if got.Source != cloud.LogSourceEngine { t.Errorf("source = %q, want engine by default", got.Source) } if got.Limit != 200*bytesPerLineGuess { @@ -81,23 +81,23 @@ func TestTopLevelLogsDefaultsToTheEngineSourceAndAnHourWindow(t *testing.T) { func TestTopLevelLogsPassesTheFlagsThrough(t *testing.T) { isolateConfig(t) - registerEnv(t, "prod", remote.Config{ + registerEnv(t, "prod", cloud.Config{ StartURL: "https://start.lambda-url.eu-west-1.on.aws/", StopURL: "https://stop.lambda-url.eu-west-1.on.aws/", Region: "eu-west-1", }) - var got remote.LogQuery - stubTopLevelLogsFetch(t, func(q remote.LogQuery) (remote.LogResult, error) { + var got cloud.LogQuery + stubTopLevelLogsFetch(t, func(q cloud.LogQuery) (cloud.LogResult, error) { got = q - return remote.LogResult{}, nil + return cloud.LogResult{}, nil }) if err := cmdLogs([]string{"--env", "prod", "--source", "boot", "--since", "15m", "--limit", "5", "--instance", "i-42"}); err != nil { t.Fatal(err) } - if got.Source != remote.LogSourceBoot { + if got.Source != cloud.LogSourceBoot { t.Errorf("source = %q, want boot", got.Source) } if got.Instance != "i-42" { diff --git a/cmd/spinloop/cloud_rename_test.go b/cmd/spinloop/cloud_rename_test.go new file mode 100644 index 00000000..e28b3924 --- /dev/null +++ b/cmd/spinloop/cloud_rename_test.go @@ -0,0 +1,35 @@ +package main + +import ( + "slices" + "testing" +) + +// The old command group is gone: `remote` fails like any unknown command and +// completion does not offer it. +func TestRemoteCommandGroupIsRemoved(t *testing.T) { + isolateConfig(t) + if err := run([]string{"remote", "status"}); err == nil { + t.Fatal("spinloop remote status should fail: the group was renamed to cloud") + } + got, _ := complete(t, "") + if slices.Contains(got, "remote") { + t.Errorf("completion still offers remote: %v", got) + } + if !slices.Contains(got, "cloud") { + t.Errorf("completion does not offer cloud: %v", got) + } +} + +// Only SPINLOOP_CLOUD_* is read; the old SPINLOOP_REMOTE_* spelling is ignored. +func TestCloudEnvVarsIgnoreTheOldPrefix(t *testing.T) { + t.Setenv("SPINLOOP_CLOUD_REGION", "") + t.Setenv("SPINLOOP_REMOTE_REGION", "eu-west-2") + if got := newCLIViper().GetString("cloud_region"); got != "" { + t.Errorf("region = %q from SPINLOOP_REMOTE_REGION, want it ignored", got) + } + t.Setenv("SPINLOOP_CLOUD_REGION", "us-east-1") + if got := newCLIViper().GetString("cloud_region"); got != "us-east-1" { + t.Errorf("region = %q, want us-east-1 from SPINLOOP_CLOUD_REGION", got) + } +} diff --git a/cmd/spinloop/remote_seed.go b/cmd/spinloop/cloud_seed.go similarity index 76% rename from cmd/spinloop/remote_seed.go rename to cmd/spinloop/cloud_seed.go index a163fc88..1f500e6a 100644 --- a/cmd/spinloop/remote_seed.go +++ b/cmd/spinloop/cloud_seed.go @@ -7,16 +7,16 @@ import ( "strings" "github.com/spf13/cobra" - "github.com/spinloop-ai/spinloop/internal/remote" + "github.com/spinloop-ai/spinloop/internal/cloud" ) -// remoteSeedCmd is the seed subcommand parent. Seeds are account-wide — one +// cloudSeedCmd is the seed subcommand parent. Seeds are account-wide — one // model seeded once serves every environment that names it — so unlike the -// other remote subcommands these do not act on an environment. What to seed -// still comes from a Spinloop, resolved exactly as `spinloop remote deploy` +// other cloud subcommands these do not act on an environment. What to seed +// still comes from a Spinloop, resolved exactly as `spinloop cloud deploy` // resolves it, so seeding and deploying in the same directory always speak // about the same model. -func remoteSeedCmd() *cobra.Command { +func cloudSeedCmd() *cobra.Command { seed := &cobra.Command{ Use: "seed", Short: "start, watch and stop model weight seeds", @@ -25,10 +25,10 @@ func remoteSeedCmd() *cobra.Command { RunE: groupFallback, } seed.AddCommand( - remoteSeedStartCmd(), - remoteSeedStatusCmd(), - remoteSeedListCmd(), - remoteSeedStopCmd(), + cloudSeedStartCmd(), + cloudSeedStatusCmd(), + cloudSeedListCmd(), + cloudSeedStopCmd(), ) return seed } @@ -36,19 +36,19 @@ func remoteSeedCmd() *cobra.Command { // seedControlConfig finds the control plane. The seed endpoints are shared // across environments and the stack outputs carry them, so this is the same // discovery `deploy` performs. -func seedControlConfig(ctx context.Context, region string) (remote.Config, error) { - awsCfg, err := remote.LoadAWSConfig(ctx, resolveRegion(region)) +func seedControlConfig(ctx context.Context, region string) (cloud.Config, error) { + awsCfg, err := cloud.LoadAWSConfig(ctx, resolveRegion(region)) if err != nil { - return remote.Config{}, err + return cloud.Config{}, err } layer, err := deployDiscoverFn(ctx, awsCfg, controlPlaneStackName) if err != nil { - return remote.Config{}, err + return cloud.Config{}, err } return layer.Config, nil } -func remoteSeedStartCmd() *cobra.Command { +func cloudSeedStartCmd() *cobra.Command { var ( force bool revision string @@ -63,7 +63,7 @@ func remoteSeedStartCmd() *cobra.Command { ValidArgsFunction: aliasSlot, RunE: func(c *cobra.Command, args []string) error { resolve(c) - return runRemoteSeedStart(args, force, revision, region) + return runCloudSeedStart(args, force, revision, region) }, } fs := c.Flags() @@ -73,11 +73,11 @@ func remoteSeedStartCmd() *cobra.Command { return c } -// runRemoteSeedStart is the body of `spinloop remote seed start`. -func runRemoteSeedStart(args []string, force bool, revision, region string) error { +// runCloudSeedStart is the body of `spinloop cloud seed start`. +func runCloudSeedStart(args []string, force bool, revision, region string) error { // Like deploy, this reads the Spinloop for what to seed, so it always needs // one — there is nothing else that says which model. - sel, spinloopPath, err := readSpinloop("spinloop remote seed start ", spinloopArg(args)) + sel, spinloopPath, err := readSpinloop("spinloop cloud seed start ", spinloopArg(args)) if err != nil { return err } @@ -94,7 +94,7 @@ func runRemoteSeedStart(args []string, force bool, revision, region string) erro if err != nil { return err } - started, err := remote.SeedStart(ctx, cfg, remote.SeedRequest{ + started, err := cloud.SeedStart(ctx, cfg, cloud.SeedRequest{ Runner: dc.Runner, ModelID: dc.ModelID, Quant: dc.Quant, @@ -118,12 +118,12 @@ func runRemoteSeedStart(args []string, force bool, revision, region string) erro fmt.Printf(" weights: %s\n", started.WeightsPrefix) } if started.SeedID != "" && !started.AlreadySeeded { - fmt.Printf("\nFollow it:\n spinloop remote seed status %s\n", started.SeedID) + fmt.Printf("\nFollow it:\n spinloop cloud seed status %s\n", started.SeedID) } return nil } -func remoteSeedStatusCmd() *cobra.Command { +func cloudSeedStatusCmd() *cobra.Command { var region string c := &cobra.Command{ Use: "status ", @@ -133,18 +133,18 @@ func remoteSeedStatusCmd() *cobra.Command { SilenceUsage: true, RunE: func(c *cobra.Command, args []string) error { resolve(c) - return runRemoteSeedStatus(args, region) + return runCloudSeedStatus(args, region) }, } c.Flags().StringVar(®ion, "region", "", "AWS region of the control plane (default: AWS_REGION or us-east-1)") return c } -// runRemoteSeedStatus is the body of `spinloop remote seed status`. -func runRemoteSeedStatus(args []string, region string) error { +// runCloudSeedStatus is the body of `spinloop cloud seed status`. +func runCloudSeedStatus(args []string, region string) error { seedID := spinloopArg(args) if seedID == "" { - return fmt.Errorf("usage: spinloop remote seed status (list them with `spinloop remote seed ls`)") + return fmt.Errorf("usage: spinloop cloud seed status (list them with `spinloop cloud seed ls`)") } ctx := context.Background() @@ -152,7 +152,7 @@ func runRemoteSeedStatus(args []string, region string) error { if err != nil { return err } - status, err := remote.SeedGet(ctx, cfg, seedID) + status, err := cloud.SeedGet(ctx, cfg, seedID) if err != nil { return err } @@ -194,7 +194,7 @@ func isTerminalSeedState(state string) bool { return state == "succeeded" || state == "failed" || state == "stopped" } -func remoteSeedListCmd() *cobra.Command { +func cloudSeedListCmd() *cobra.Command { var region string c := &cobra.Command{ Use: "ls", @@ -205,21 +205,21 @@ func remoteSeedListCmd() *cobra.Command { SilenceUsage: true, RunE: func(c *cobra.Command, args []string) error { resolve(c) - return runRemoteSeedList(region) + return runCloudSeedList(region) }, } c.Flags().StringVar(®ion, "region", "", "AWS region of the control plane (default: AWS_REGION or us-east-1)") return c } -// runRemoteSeedList is the body of `spinloop remote seed ls`. -func runRemoteSeedList(region string) error { +// runCloudSeedList is the body of `spinloop cloud seed ls`. +func runCloudSeedList(region string) error { ctx := context.Background() cfg, err := seedControlConfig(ctx, region) if err != nil { return err } - seeds, err := remote.SeedList(ctx, cfg) + seeds, err := cloud.SeedList(ctx, cfg) if err != nil { return err } @@ -236,7 +236,7 @@ func runRemoteSeedList(region string) error { return nil } -func remoteSeedStopCmd() *cobra.Command { +func cloudSeedStopCmd() *cobra.Command { var region string c := &cobra.Command{ Use: "stop ", @@ -246,18 +246,18 @@ func remoteSeedStopCmd() *cobra.Command { SilenceUsage: true, RunE: func(c *cobra.Command, args []string) error { resolve(c) - return runRemoteSeedStop(args, region) + return runCloudSeedStop(args, region) }, } c.Flags().StringVar(®ion, "region", "", "AWS region of the control plane (default: AWS_REGION or us-east-1)") return c } -// runRemoteSeedStop is the body of `spinloop remote seed stop`. -func runRemoteSeedStop(args []string, region string) error { +// runCloudSeedStop is the body of `spinloop cloud seed stop`. +func runCloudSeedStop(args []string, region string) error { seedID := spinloopArg(args) if seedID == "" { - return fmt.Errorf("usage: spinloop remote seed stop ") + return fmt.Errorf("usage: spinloop cloud seed stop ") } ctx := context.Background() @@ -265,7 +265,7 @@ func runRemoteSeedStop(args []string, region string) error { if err != nil { return err } - stopped, err := remote.SeedStop(ctx, cfg, seedID) + stopped, err := cloud.SeedStop(ctx, cfg, seedID) if err != nil { return err } @@ -306,4 +306,4 @@ func humanAge(seconds int) string { } } -func cmdRemoteSeed(args []string) error { return execCmd(remoteSeedCmd(), args) } +func cmdCloudSeed(args []string) error { return execCmd(cloudSeedCmd(), args) } diff --git a/cmd/spinloop/remote_seed_test.go b/cmd/spinloop/cloud_seed_test.go similarity index 79% rename from cmd/spinloop/remote_seed_test.go rename to cmd/spinloop/cloud_seed_test.go index bd45a3c3..9cf5f32f 100644 --- a/cmd/spinloop/remote_seed_test.go +++ b/cmd/spinloop/cloud_seed_test.go @@ -11,7 +11,7 @@ import ( "github.com/aws/aws-sdk-go-v2/aws" - "github.com/spinloop-ai/spinloop/internal/remote" + "github.com/spinloop-ai/spinloop/internal/cloud" ) // stubSeedSeams points the control-plane discovery at a test server, with the @@ -27,8 +27,8 @@ func stubSeedSeams(t *testing.T, serverURL, seedURL string) { stubAWSEnv(t) orig := deployDiscoverFn t.Cleanup(func() { deployDiscoverFn = orig }) - deployDiscoverFn = func(context.Context, aws.Config, string) (remote.ControlPlane, error) { - return remote.ControlPlane{Config: remote.Config{ + deployDiscoverFn = func(context.Context, aws.Config, string) (cloud.ControlPlane, error) { + return cloud.ControlPlane{Config: cloud.Config{ StartURL: serverURL, StopURL: serverURL, DeployURL: serverURL, SeedURL: seedURL, Region: "us-east-1", }}, nil @@ -63,14 +63,14 @@ func seedServer(t *testing.T, status int, body string) (*httptest.Server, *struc return server, got } -func TestRemoteSeed_StartPostsWhatTheSpinloopNames(t *testing.T) { +func TestCloudSeed_StartPostsWhatTheSpinloopNames(t *testing.T) { server, got := seedServer(t, 200, `{"seedId":"llamacpp--m","instanceId":"i-1","started":true,"weightsPrefix":"models/llamacpp/m/"}`) stubSeedSeams(t, server.URL, server.URL) writeDeploySpinloopCwd(t) out := captureStdout(t, func() { - if err := cmdRemoteSeed([]string{"start"}); err != nil { + if err := cmdCloudSeed([]string{"start"}); err != nil { t.Fatalf("seed start: %v", err) } }) @@ -86,18 +86,18 @@ func TestRemoteSeed_StartPostsWhatTheSpinloopNames(t *testing.T) { if _, ok := got.Body["weightsPrefix"]; ok { t.Error("the request must not carry a weights prefix") } - if !strings.Contains(out, "llamacpp--m") || !strings.Contains(out, "spinloop remote seed status") { + if !strings.Contains(out, "llamacpp--m") || !strings.Contains(out, "spinloop cloud seed status") { t.Errorf("start should name the seed and how to follow it, got:\n%s", out) } } -func TestRemoteSeed_StartSaysWhenItJoinedRatherThanStarted(t *testing.T) { +func TestCloudSeed_StartSaysWhenItJoinedRatherThanStarted(t *testing.T) { server, _ := seedServer(t, 200, `{"seedId":"llamacpp--m","started":false,"joined":true}`) stubSeedSeams(t, server.URL, server.URL) writeDeploySpinloopCwd(t) out := captureStdout(t, func() { - if err := cmdRemoteSeed([]string{"start"}); err != nil { + if err := cmdCloudSeed([]string{"start"}); err != nil { t.Fatalf("seed start: %v", err) } }) @@ -107,13 +107,13 @@ func TestRemoteSeed_StartSaysWhenItJoinedRatherThanStarted(t *testing.T) { } } -func TestRemoteSeed_StartReportsAlreadySeeded(t *testing.T) { +func TestCloudSeed_StartReportsAlreadySeeded(t *testing.T) { server, _ := seedServer(t, 200, `{"seedId":"llamacpp--m","started":false,"alreadySeeded":true}`) stubSeedSeams(t, server.URL, server.URL) writeDeploySpinloopCwd(t) out := captureStdout(t, func() { - if err := cmdRemoteSeed([]string{"start"}); err != nil { + if err := cmdCloudSeed([]string{"start"}); err != nil { t.Fatalf("seed start: %v", err) } }) @@ -122,13 +122,13 @@ func TestRemoteSeed_StartReportsAlreadySeeded(t *testing.T) { } } -func TestRemoteSeed_ForceAndRevisionTravel(t *testing.T) { +func TestCloudSeed_ForceAndRevisionTravel(t *testing.T) { server, got := seedServer(t, 200, `{"seedId":"llamacpp--m","started":true}`) stubSeedSeams(t, server.URL, server.URL) writeDeploySpinloopCwd(t) captureStdout(t, func() { - if err := cmdRemoteSeed([]string{"start", "--force", "--revision", "abc123"}); err != nil { + if err := cmdCloudSeed([]string{"start", "--force", "--revision", "abc123"}); err != nil { t.Fatalf("seed start: %v", err) } }) @@ -137,13 +137,13 @@ func TestRemoteSeed_ForceAndRevisionTravel(t *testing.T) { } } -func TestRemoteSeed_StatusReportsAFailedSeed(t *testing.T) { +func TestCloudSeed_StatusReportsAFailedSeed(t *testing.T) { server, got := seedServer(t, 200, `{"seedId":"vllm--m","state":"failed","modelId":"org/m","error":"checksum mismatch","progressPercent":41}`) stubSeedSeams(t, server.URL, server.URL) out := captureStdout(t, func() { - if err := cmdRemoteSeed([]string{"status", "vllm--m"}); err != nil { + if err := cmdCloudSeed([]string{"status", "vllm--m"}); err != nil { t.Fatalf("seed status: %v", err) } }) @@ -155,14 +155,14 @@ func TestRemoteSeed_StatusReportsAFailedSeed(t *testing.T) { } } -func TestRemoteSeed_StatusShowsTheFileOnlyWhileItIsStillWorking(t *testing.T) { +func TestCloudSeed_StatusShowsTheFileOnlyWhileItIsStillWorking(t *testing.T) { // While transferring, the current file is what you want to see. server, _ := seedServer(t, 200, `{"seedId":"vllm--m","state":"transferring","currentFile":"model-00009.safetensors","filesTotal":17,"filesDone":8,"bytesTotal":2048,"bytesDone":1024,"progressPercent":50}`) stubSeedSeams(t, server.URL, server.URL) out := captureStdout(t, func() { - if err := cmdRemoteSeed([]string{"status", "vllm--m"}); err != nil { + if err := cmdCloudSeed([]string{"status", "vllm--m"}); err != nil { t.Fatalf("seed status: %v", err) } }) @@ -174,14 +174,14 @@ func TestRemoteSeed_StatusShowsTheFileOnlyWhileItIsStillWorking(t *testing.T) { } } -func TestRemoteSeed_StatusHidesTheFileOnceFinished(t *testing.T) { +func TestCloudSeed_StatusHidesTheFileOnceFinished(t *testing.T) { // Once terminal, the last file it happened to touch is noise. server, _ := seedServer(t, 200, `{"seedId":"vllm--m","state":"succeeded","currentFile":"model-00017.safetensors","durationSeconds":312}`) stubSeedSeams(t, server.URL, server.URL) out := captureStdout(t, func() { - if err := cmdRemoteSeed([]string{"status", "vllm--m"}); err != nil { + if err := cmdCloudSeed([]string{"status", "vllm--m"}); err != nil { t.Fatalf("seed status: %v", err) } }) @@ -193,22 +193,22 @@ func TestRemoteSeed_StatusHidesTheFileOnceFinished(t *testing.T) { } } -func TestRemoteSeed_StatusRejectsAnUnknownSeed(t *testing.T) { +func TestCloudSeed_StatusRejectsAnUnknownSeed(t *testing.T) { server, _ := seedServer(t, 404, `{"seedId":"nope","state":"unknown","error":"no seed known"}`) stubSeedSeams(t, server.URL, server.URL) - err := cmdRemoteSeed([]string{"status", "nope"}) + err := cmdCloudSeed([]string{"status", "nope"}) if err == nil || !strings.Contains(err.Error(), `no seed "nope" is known`) { t.Errorf("want an unknown-seed error, got %v", err) } } -func TestRemoteSeed_ListStatesPlainlyWhenEmpty(t *testing.T) { +func TestCloudSeed_ListStatesPlainlyWhenEmpty(t *testing.T) { server, _ := seedServer(t, 200, `{"seeds":[],"count":0}`) stubSeedSeams(t, server.URL, server.URL) out := captureStdout(t, func() { - if err := cmdRemoteSeed([]string{"ls"}); err != nil { + if err := cmdCloudSeed([]string{"ls"}); err != nil { t.Fatalf("seed ls: %v", err) } }) @@ -218,13 +218,13 @@ func TestRemoteSeed_ListStatesPlainlyWhenEmpty(t *testing.T) { } } -func TestRemoteSeed_ListShowsEachSeed(t *testing.T) { +func TestCloudSeed_ListShowsEachSeed(t *testing.T) { server, _ := seedServer(t, 200, `{"seeds":[{"seedId":"vllm--m","modelId":"org/m","state":"transferring","ageSeconds":125}],"count":1}`) stubSeedSeams(t, server.URL, server.URL) out := captureStdout(t, func() { - if err := cmdRemoteSeed([]string{"ls"}); err != nil { + if err := cmdCloudSeed([]string{"ls"}); err != nil { t.Fatalf("seed ls: %v", err) } }) @@ -235,13 +235,13 @@ func TestRemoteSeed_ListShowsEachSeed(t *testing.T) { } } -func TestRemoteSeed_StopIsSafeWhenNothingIsRunning(t *testing.T) { +func TestCloudSeed_StopIsSafeWhenNothingIsRunning(t *testing.T) { server, got := seedServer(t, 200, `{"seedId":"vllm--m","stopped":false,"message":"not running"}`) stubSeedSeams(t, server.URL, server.URL) out := captureStdout(t, func() { // Stopping twice must not be an error. - if err := cmdRemoteSeed([]string{"stop", "vllm--m"}); err != nil { + if err := cmdCloudSeed([]string{"stop", "vllm--m"}); err != nil { t.Fatalf("seed stop: %v", err) } }) @@ -253,12 +253,12 @@ func TestRemoteSeed_StopIsSafeWhenNothingIsRunning(t *testing.T) { } } -func TestRemoteSeed_StopReportsWhatItStopped(t *testing.T) { +func TestCloudSeed_StopReportsWhatItStopped(t *testing.T) { server, _ := seedServer(t, 200, `{"seedId":"vllm--m","stopped":true,"instanceIds":["i-1"]}`) stubSeedSeams(t, server.URL, server.URL) out := captureStdout(t, func() { - if err := cmdRemoteSeed([]string{"stop", "vllm--m"}); err != nil { + if err := cmdCloudSeed([]string{"stop", "vllm--m"}); err != nil { t.Fatalf("seed stop: %v", err) } }) @@ -267,45 +267,45 @@ func TestRemoteSeed_StopReportsWhatItStopped(t *testing.T) { } } -func TestRemoteSeed_CapReachedIsNamed(t *testing.T) { +func TestCloudSeed_CapReachedIsNamed(t *testing.T) { server, _ := seedServer(t, 429, `{"error":"3 seeds are already running (cap 3) — wait for one to finish"}`) stubSeedSeams(t, server.URL, server.URL) writeDeploySpinloopCwd(t) - err := cmdRemoteSeed([]string{"start"}) + err := cmdCloudSeed([]string{"start"}) if err == nil || !strings.Contains(err.Error(), "cap 3") { t.Errorf("a refusal at the cap should say so, got %v", err) } } -func TestRemoteSeed_NoSeedURLNamesTheValueToAdd(t *testing.T) { +func TestCloudSeed_NoSeedURLNamesTheValueToAdd(t *testing.T) { server, _ := seedServer(t, 200, `{}`) - // A remote config written before the seed Lambda existed. + // A cloud config written before the seed Lambda existed. stubSeedSeams(t, server.URL, "") - err := cmdRemoteSeed([]string{"ls"}) + err := cmdCloudSeed([]string{"ls"}) if err == nil { t.Fatal("want an error when no seed endpoint is configured") } - for _, want := range []string{"seed_url", "SeedUrl", "SPINLOOP_REMOTE_SEED_URL"} { + for _, want := range []string{"seed_url", "SeedUrl", "SPINLOOP_CLOUD_SEED_URL"} { if !strings.Contains(err.Error(), want) { t.Errorf("the error should name %q, got: %v", want, err) } } } -func TestRemoteSeed_UnknownSubcommand(t *testing.T) { - err := cmdRemoteSeed([]string{"frobnicate"}) +func TestCloudSeed_UnknownSubcommand(t *testing.T) { + err := cmdCloudSeed([]string{"frobnicate"}) if err == nil || !strings.Contains(err.Error(), "unknown command") { t.Errorf("want an unknown-subcommand error, got %v", err) } } -func TestRemoteSeed_NoSubcommandShowsHelp(t *testing.T) { +func TestCloudSeed_NoSubcommandShowsHelp(t *testing.T) { // Bare: no error — cobra shows the group's own help, with the // subcommand list generated from the tree. out := captureStdout(t, func() { - if err := cmdRemoteSeed([]string{}); err != nil { + if err := cmdCloudSeed([]string{}); err != nil { t.Fatalf("bare seed should show its help, got %v", err) } }) @@ -314,8 +314,8 @@ func TestRemoteSeed_NoSubcommandShowsHelp(t *testing.T) { } } -func TestRemoteSeed_StatusNeedsASeedID(t *testing.T) { - err := cmdRemoteSeed([]string{"status"}) +func TestCloudSeed_StatusNeedsASeedID(t *testing.T) { + err := cmdCloudSeed([]string{"status"}) if err == nil || !strings.Contains(err.Error(), "seed ls") { t.Errorf("want usage pointing at ls, got %v", err) } diff --git a/cmd/spinloop/remote_sources.go b/cmd/spinloop/cloud_sources.go similarity index 83% rename from cmd/spinloop/remote_sources.go rename to cmd/spinloop/cloud_sources.go index 201a2b85..98cc2879 100644 --- a/cmd/spinloop/remote_sources.go +++ b/cmd/spinloop/cloud_sources.go @@ -1,7 +1,7 @@ package main import ( - "github.com/spinloop-ai/spinloop/internal/remote" + "github.com/spinloop-ai/spinloop/internal/cloud" ) // sourceLocation is where the CDK project's sources live for a run: the @@ -18,8 +18,8 @@ type sourceLocation struct { // ref-keyed default directory, and pruning of other refs — except that an // explicit --dir is the user's own location, neither keyed by ref nor pruned. func resolveSourceLocation(ref, dir string) (sourceLocation, error) { - resolvedRef := remote.ResolveRef(version, ref) - locDir, err := remote.SourceDir(resolvedRef) + resolvedRef := cloud.ResolveRef(version, ref) + locDir, err := cloud.SourceDir(resolvedRef) if err != nil { return sourceLocation{}, err } @@ -38,9 +38,9 @@ func pruneSourceCaches(loc sourceLocation) error { if !loc.prune { return nil } - sourceRoot, err := remote.SourceRoot() + sourceRoot, err := cloud.SourceRoot() if err != nil { return err } - return remote.PruneSources(sourceRoot, loc.ref) + return cloud.PruneSources(sourceRoot, loc.ref) } diff --git a/cmd/spinloop/remote_test.go b/cmd/spinloop/cloud_test.go similarity index 81% rename from cmd/spinloop/remote_test.go rename to cmd/spinloop/cloud_test.go index cde88c89..42c19d7f 100644 --- a/cmd/spinloop/remote_test.go +++ b/cmd/spinloop/cloud_test.go @@ -13,7 +13,7 @@ import ( "testing" "time" - "github.com/spinloop-ai/spinloop/internal/remote" + "github.com/spinloop-ai/spinloop/internal/cloud" ) // stubAWSEnv pins the default credential chain to static env credentials so @@ -29,15 +29,15 @@ func stubAWSEnv(t *testing.T) { t.Setenv("AWS_EC2_METADATA_DISABLED", "true") } -// writeRemoteConfig registers the `default` environment pointing at the test +// writeCloudConfig registers the `default` environment pointing at the test // server — the registry is the only place a configuration is read from. -func writeRemoteConfig(t *testing.T, serverURL string) { +func writeCloudConfig(t *testing.T, serverURL string) { t.Helper() - path := must1(remote.EnvConfigPath("default")) + path := must1(cloud.EnvConfigPath("default")) if err := os.MkdirAll(filepath.Dir(path), 0o700); err != nil { t.Fatal(err) } - data, err := json.Marshal(remote.Config{ + data, err := json.Marshal(cloud.Config{ StartURL: serverURL, StopURL: serverURL, EnvURL: serverURL, @@ -51,23 +51,23 @@ func writeRemoteConfig(t *testing.T, serverURL string) { } } -func TestRemoteDispatch(t *testing.T) { - // Bare remote shows the group's own help — generated from the tree, so +func TestCloudDispatch(t *testing.T) { + // Bare cloud shows the group's own help — generated from the tree, so // its subcommand list cannot drift — rather than an error. out := captureStdout(t, func() { - if err := run([]string{"remote"}); err != nil { - t.Fatalf("bare remote should show its help, got %v", err) + if err := run([]string{"cloud"}); err != nil { + t.Fatalf("bare cloud should show its help, got %v", err) } }) if !strings.Contains(out, "bootstrap") || !strings.Contains(out, "bake") { - t.Errorf("bare remote help should name its subcommands, got:\n%s", out) + t.Errorf("bare cloud help should name its subcommands, got:\n%s", out) } - if err := run([]string{"remote", "bogus"}); err == nil || !strings.Contains(err.Error(), "bogus") { + if err := run([]string{"cloud", "bogus"}); err == nil || !strings.Contains(err.Error(), "bogus") { t.Errorf("unknown subcommand should error, got %v", err) } } -func TestRemote_Unconfigured(t *testing.T) { +func TestCloud_Unconfigured(t *testing.T) { isolateConfig(t) // deploy needs a Spinloop, covered separately. metrics, logs and status // moved to the top level — cmd/spinloop/commands_test.go covers their @@ -76,21 +76,21 @@ func TestRemote_Unconfigured(t *testing.T) { // Naming no environment is its own failure: these commands act on one // instance, and an instance nobody named is not one to act on. for _, sub := range subs { - err := run([]string{"remote", sub}) + err := run([]string{"cloud", sub}) if err == nil || !strings.Contains(err.Error(), "pass --env") { - t.Errorf("remote %s with no environment should name the flag, got %v", sub, err) + t.Errorf("cloud %s with no environment should name the flag, got %v", sub, err) } } // Naming one that has no configuration explains the setup. for _, sub := range subs { - err := run([]string{"remote", sub, "--env", "default"}) + err := run([]string{"cloud", sub, "--env", "default"}) if err == nil || !strings.Contains(err.Error(), "not configured") { - t.Errorf("remote %s without config should explain setup, got %v", sub, err) + t.Errorf("cloud %s without config should explain setup, got %v", sub, err) } } } -func TestRemoteStart_PrintsExports(t *testing.T) { +func TestCloudStart_PrintsExports(t *testing.T) { isolateConfig(t) stubAWSEnv(t) server := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) { @@ -98,11 +98,11 @@ func TestRemoteStart_PrintsExports(t *testing.T) { w.Write([]byte(`{"state":"ready","base_url":"http://198.51.100.1:8000/v1","api_key":"sk-test"}`)) })) defer server.Close() - writeRemoteConfig(t, server.URL) + writeCloudConfig(t, server.URL) out := captureStdout(t, func() { - if err := cmdRemoteStart([]string{"--env", "default", "--print-env"}); err != nil { - t.Errorf("cmdRemoteStart: %v", err) + if err := cmdCloudStart([]string{"--env", "default", "--print-env"}); err != nil { + t.Errorf("cmdCloudStart: %v", err) } }) if !strings.Contains(out, "export OPENAI_BASE_URL=http://198.51.100.1:8000/v1") || @@ -112,7 +112,7 @@ func TestRemoteStart_PrintsExports(t *testing.T) { } // Start without --print-env prints nothing to stdout (progress goes to stderr). -func TestRemoteStart_NoExportsWithoutFlag(t *testing.T) { +func TestCloudStart_NoExportsWithoutFlag(t *testing.T) { isolateConfig(t) stubAWSEnv(t) server := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) { @@ -120,11 +120,11 @@ func TestRemoteStart_NoExportsWithoutFlag(t *testing.T) { w.Write([]byte(`{"state":"ready","base_url":"http://198.51.100.1:8000/v1","api_key":"sk-test"}`)) })) defer server.Close() - writeRemoteConfig(t, server.URL) + writeCloudConfig(t, server.URL) out := captureStdout(t, func() { - if err := cmdRemoteStart([]string{"--env", "default"}); err != nil { - t.Errorf("cmdRemoteStart: %v", err) + if err := cmdCloudStart([]string{"--env", "default"}); err != nil { + t.Errorf("cmdCloudStart: %v", err) } }) if strings.Contains(out, "export OPENAI_") { @@ -134,9 +134,9 @@ func TestRemoteStart_NoExportsWithoutFlag(t *testing.T) { // Start with --print-env after a positional argument still parses the flag. // Regression test: Go's flag package stops at the first non-flag argument, -// so `spinloop remote start path --print-env` would silently ignore +// so `spinloop cloud start path --print-env` would silently ignore // --print-env without sortFlagsBeforeArgs. -func TestRemoteStart_FlagAfterPositional(t *testing.T) { +func TestCloudStart_FlagAfterPositional(t *testing.T) { isolateConfig(t) stubAWSEnv(t) server := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) { @@ -144,7 +144,7 @@ func TestRemoteStart_FlagAfterPositional(t *testing.T) { w.Write([]byte(`{"state":"ready","base_url":"http://198.51.100.1:8000/v1","api_key":"sk-test"}`)) })) defer server.Close() - registerEnv(t, "testenv", remote.Config{StartURL: server.URL, StopURL: server.URL, EnvURL: server.URL, Region: "eu-west-1"}) + registerEnv(t, "testenv", cloud.Config{StartURL: server.URL, StopURL: server.URL, EnvURL: server.URL, Region: "eu-west-1"}) dir := t.TempDir() spinloopFile := "PROVIDER openai-compatible\n" @@ -153,8 +153,8 @@ func TestRemoteStart_FlagAfterPositional(t *testing.T) { } out := captureStdout(t, func() { - if err := cmdRemoteStart([]string{dir, "--env", "testenv", "--print-env"}); err != nil { - t.Errorf("cmdRemoteStart: %v", err) + if err := cmdCloudStart([]string{dir, "--env", "testenv", "--print-env"}); err != nil { + t.Errorf("cmdCloudStart: %v", err) } }) if !strings.Contains(out, "export OPENAI_BASE_URL=http://198.51.100.1:8000/v1") || @@ -163,8 +163,8 @@ func TestRemoteStart_FlagAfterPositional(t *testing.T) { } } -// Remote env command prints exports for a running endpoint. -func TestRemoteEnv_PrintsExports(t *testing.T) { +// Cloud env command prints exports for a running endpoint. +func TestCloudEnv_PrintsExports(t *testing.T) { isolateConfig(t) stubAWSEnv(t) server := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) { @@ -175,11 +175,11 @@ func TestRemoteEnv_PrintsExports(t *testing.T) { w.Write([]byte(`{"base_url":"http://198.51.100.1:8000/v1","api_key":"sk-remote"}`)) })) defer server.Close() - writeRemoteConfig(t, server.URL) + writeCloudConfig(t, server.URL) out := captureStdout(t, func() { - if err := cmdRemoteEnv([]string{"--env", "default"}); err != nil { - t.Errorf("cmdRemoteEnv: %v", err) + if err := cmdCloudEnv([]string{"--env", "default"}); err != nil { + t.Errorf("cmdCloudEnv: %v", err) } }) if !strings.Contains(out, "export OPENAI_BASE_URL=http://198.51.100.1:8000/v1") || @@ -188,7 +188,7 @@ func TestRemoteEnv_PrintsExports(t *testing.T) { } } -func TestRemoteStatus_PrintsState(t *testing.T) { +func TestCloudStatus_PrintsState(t *testing.T) { isolateConfig(t) stubAWSEnv(t) server := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) { @@ -196,7 +196,7 @@ func TestRemoteStatus_PrintsState(t *testing.T) { w.Write([]byte(`{"state":"running","healthy":true,"base_url":"http://198.51.100.1:8000/v1"}`)) })) defer server.Close() - writeRemoteConfig(t, server.URL) + writeCloudConfig(t, server.URL) out := captureStdout(t, func() { if err := cmdStatus([]string{"--env", "default"}); err != nil { @@ -214,9 +214,9 @@ func TestRemoteStatus_PrintsState(t *testing.T) { } } -// Remote status includes the spinloop version from the stats Lambda when the +// Cloud status includes the spinloop version from the stats Lambda when the // instance is running, so the operator can verify the release without SSH. -func TestRemoteStatus_PrintsVersion(t *testing.T) { +func TestCloudStatus_PrintsVersion(t *testing.T) { isolateConfig(t) stubAWSEnv(t) server := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) { @@ -229,11 +229,11 @@ func TestRemoteStatus_PrintsVersion(t *testing.T) { } })) defer server.Close() - path := must1(remote.EnvConfigPath("default")) + path := must1(cloud.EnvConfigPath("default")) if err := os.MkdirAll(filepath.Dir(path), 0o700); err != nil { t.Fatal(err) } - data, err := json.Marshal(remote.Config{ + data, err := json.Marshal(cloud.Config{ StartURL: server.URL + "/start", StopURL: server.URL + "/stop", EnvURL: server.URL + "/env", @@ -257,7 +257,7 @@ func TestRemoteStatus_PrintsVersion(t *testing.T) { } } -func TestRemoteStop_PrintsState(t *testing.T) { +func TestCloudStop_PrintsState(t *testing.T) { isolateConfig(t) stubAWSEnv(t) server := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) { @@ -268,11 +268,11 @@ func TestRemoteStop_PrintsState(t *testing.T) { w.Write([]byte(`{"state":"stopping"}`)) })) defer server.Close() - writeRemoteConfig(t, server.URL) + writeCloudConfig(t, server.URL) out := captureStdout(t, func() { - if err := cmdRemoteStop([]string{"--env", "default"}); err != nil { - t.Errorf("cmdRemoteStop: %v", err) + if err := cmdCloudStop([]string{"--env", "default"}); err != nil { + t.Errorf("cmdCloudStop: %v", err) } }) if !strings.Contains(out, "stopping") { @@ -280,7 +280,7 @@ func TestRemoteStop_PrintsState(t *testing.T) { } } -func TestRemotePause_PrintsState(t *testing.T) { +func TestCloudPause_PrintsState(t *testing.T) { isolateConfig(t) stubAWSEnv(t) var gotAction string @@ -293,17 +293,17 @@ func TestRemotePause_PrintsState(t *testing.T) { w.Write([]byte(`{"state":"stopping"}`)) })) defer server.Close() - writeRemoteConfig(t, server.URL) + writeCloudConfig(t, server.URL) out := captureStdout(t, func() { - if err := cmdRemotePause([]string{"--env", "default"}); err != nil { - t.Errorf("cmdRemotePause: %v", err) + if err := cmdCloudPause([]string{"--env", "default"}); err != nil { + t.Errorf("cmdCloudPause: %v", err) } }) if gotAction != "pause" { t.Errorf("pause must ask the stop Lambda for its pause mode, got action=%q", gotAction) } - for _, want := range []string{"state: stopping", "spinloop remote start", "spinloop remote stop"} { + for _, want := range []string{"state: stopping", "spinloop cloud start", "spinloop cloud stop"} { if !strings.Contains(out, want) { t.Errorf("pause output missing %q:\n%s", want, out) } @@ -330,20 +330,20 @@ func restartHandler(statusState string, gotForce *string, stopCalls, wakeCalls * } } -// A bare `remote restart` dispatches through the tree, stops then wakes, and +// A bare `cloud restart` dispatches through the tree, stops then wakes, and // prints the base URL as confirmation the address is unchanged. -func TestRemoteRestart_Flow(t *testing.T) { +func TestCloudRestart_Flow(t *testing.T) { isolateConfig(t) stubAWSEnv(t) var force string var stops, wakes int server := httptest.NewServer(restartHandler("running", &force, &stops, &wakes)) defer server.Close() - writeRemoteConfig(t, server.URL) + writeCloudConfig(t, server.URL) out := captureStdout(t, func() { - if err := run([]string{"remote", "restart", "--env", "default"}); err != nil { - t.Errorf("remote restart: %v", err) + if err := run([]string{"cloud", "restart", "--env", "default"}); err != nil { + t.Errorf("cloud restart: %v", err) } }) if !strings.Contains(out, "base url: http://198.51.100.1:8000/v1") { @@ -358,7 +358,7 @@ func TestRemoteRestart_Flow(t *testing.T) { } // --force (long and short) marks the stop forced on the way over. -func TestRemoteRestart_ForceFlag(t *testing.T) { +func TestCloudRestart_ForceFlag(t *testing.T) { for _, flag := range []string{"--force", "-F"} { t.Run(flag, func(t *testing.T) { isolateConfig(t) @@ -367,11 +367,11 @@ func TestRemoteRestart_ForceFlag(t *testing.T) { var stops, wakes int server := httptest.NewServer(restartHandler("running", &force, &stops, &wakes)) defer server.Close() - writeRemoteConfig(t, server.URL) + writeCloudConfig(t, server.URL) out := captureStdout(t, func() { - if err := cmdRemoteRestart([]string{"--env", "default", flag}); err != nil { - t.Errorf("remote restart %s: %v", flag, err) + if err := cmdCloudRestart([]string{"--env", "default", flag}); err != nil { + t.Errorf("cloud restart %s: %v", flag, err) } }) if !strings.Contains(out, "base url: http://198.51.100.1:8000/v1") { @@ -388,18 +388,18 @@ func TestRemoteRestart_ForceFlag(t *testing.T) { } // --timeout parses as a duration like start's, and does not error. -func TestRemoteRestart_TimeoutFlag(t *testing.T) { +func TestCloudRestart_TimeoutFlag(t *testing.T) { isolateConfig(t) stubAWSEnv(t) var force string var stops, wakes int server := httptest.NewServer(restartHandler("running", &force, &stops, &wakes)) defer server.Close() - writeRemoteConfig(t, server.URL) + writeCloudConfig(t, server.URL) out := captureStdout(t, func() { - if err := cmdRemoteRestart([]string{"--env", "default", "--timeout", "5m"}); err != nil { - t.Errorf("remote restart --timeout: %v", err) + if err := cmdCloudRestart([]string{"--env", "default", "--timeout", "5m"}); err != nil { + t.Errorf("cloud restart --timeout: %v", err) } }) if !strings.Contains(out, "base url: http://198.51.100.1:8000/v1") { @@ -409,18 +409,18 @@ func TestRemoteRestart_TimeoutFlag(t *testing.T) { // Restarting an environment that is already stopped behaves as a start: the // pause-style stop is a no-op, and the wake still brings the endpoint back. -func TestRemoteRestart_AlreadyStoppedBehavesAsStart(t *testing.T) { +func TestCloudRestart_AlreadyStoppedBehavesAsStart(t *testing.T) { isolateConfig(t) stubAWSEnv(t) var force string var stops, wakes int server := httptest.NewServer(restartHandler("stopped", &force, &stops, &wakes)) defer server.Close() - writeRemoteConfig(t, server.URL) + writeCloudConfig(t, server.URL) out := captureStdout(t, func() { - if err := cmdRemoteRestart([]string{"--env", "default"}); err != nil { - t.Errorf("remote restart on a stopped environment: %v", err) + if err := cmdCloudRestart([]string{"--env", "default"}); err != nil { + t.Errorf("cloud restart on a stopped environment: %v", err) } }) if !strings.Contains(out, "base url: http://198.51.100.1:8000/v1") { @@ -434,7 +434,7 @@ func TestRemoteRestart_AlreadyStoppedBehavesAsStart(t *testing.T) { // A failed status check does not gate the restart: the stop Lambda is correct // for every state, so the command skips the status line and goes ahead, and the // status line stays absent rather than claiming a state it never read. -func TestRemoteRestart_StatusFailureDoesNotGate(t *testing.T) { +func TestCloudRestart_StatusFailureDoesNotGate(t *testing.T) { isolateConfig(t) stubAWSEnv(t) server := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) { @@ -451,12 +451,12 @@ func TestRemoteRestart_StatusFailureDoesNotGate(t *testing.T) { w.Write([]byte(`{"state":"ready","base_url":"http://198.51.100.1:8000/v1","api_key":"sk-test"}`)) })) defer server.Close() - writeRemoteConfig(t, server.URL) + writeCloudConfig(t, server.URL) out := captureStdout(t, func() { errOut := captureStderr(t, func() { - if err := cmdRemoteRestart([]string{"--env", "default"}); err != nil { - t.Errorf("remote restart after a failed status check: %v", err) + if err := cmdCloudRestart([]string{"--env", "default"}); err != nil { + t.Errorf("cloud restart after a failed status check: %v", err) } }) if strings.Contains(errOut, "the instance is") { @@ -471,7 +471,7 @@ func TestRemoteRestart_StatusFailureDoesNotGate(t *testing.T) { // When the stop took effect and the wake then fails, the recovery hint reaches // the user as the command's error: the instance is stopped, and start brings // it back. -func TestRemoteRestart_WakeFailureReportsRecovery(t *testing.T) { +func TestCloudRestart_WakeFailureReportsRecovery(t *testing.T) { isolateConfig(t) stubAWSEnv(t) server := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) { @@ -484,40 +484,40 @@ func TestRemoteRestart_WakeFailureReportsRecovery(t *testing.T) { w.Write([]byte(`{"state":"terminated","message":"cannot start"}`)) })) defer server.Close() - writeRemoteConfig(t, server.URL) + writeCloudConfig(t, server.URL) - err := cmdRemoteRestart([]string{"--env", "default"}) + err := cmdCloudRestart([]string{"--env", "default"}) if err == nil { t.Fatal("expected a wake failure error") } - if !strings.Contains(err.Error(), "stopped") || !strings.Contains(err.Error(), "spinloop remote start") { + if !strings.Contains(err.Error(), "stopped") || !strings.Contains(err.Error(), "spinloop cloud start") { t.Errorf("expected the recovery hint in the error, got %v", err) } } // The parent fallback names restart in both its usage line and its -// unknown-subcommand list, so a mistyped or bare `remote` points to it. -func TestRemote_RestartInGeneratedHelp(t *testing.T) { +// unknown-subcommand list, so a mistyped or bare `cloud` points to it. +func TestCloud_RestartInGeneratedHelp(t *testing.T) { isolateConfig(t) // A regression from when the usage was a hand-rolled list: restart had // to be added to two places by hand. The help now comes from the tree. out := captureStdout(t, func() { - if err := run([]string{"remote"}); err != nil { - t.Fatalf("bare remote should show its help, got %v", err) + if err := run([]string{"cloud"}); err != nil { + t.Fatalf("bare cloud should show its help, got %v", err) } }) if !strings.Contains(out, "restart") { - t.Errorf("bare remote help should name restart, got:\n%s", out) + t.Errorf("bare cloud help should name restart, got:\n%s", out) } } // A Spinloop is read only when named with --spinloop/-O — never implicitly // from the working directory — for its ENV instructions, which here name the // control plane a command with no other configuration would not find. -func TestRemote_SpinloopDiscovery(t *testing.T) { +func TestCloud_SpinloopDiscovery(t *testing.T) { isolateConfig(t) // no per-user config exists, so success proves the ENV reached it stubAWSEnv(t) - unsetEnvOnCleanup(t, "SPINLOOP_REMOTE_START_URL", "SPINLOOP_REMOTE_STOP_URL", "SPINLOOP_REMOTE_REGION") + unsetEnvOnCleanup(t, "SPINLOOP_CLOUD_START_URL", "SPINLOOP_CLOUD_STOP_URL", "SPINLOOP_CLOUD_REGION") server := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) { w.Header().Set("Content-Type", "application/json") w.Write([]byte(`{"state":"running","healthy":true,"base_url":"http://198.51.100.1:8000/v1"}`)) @@ -526,7 +526,7 @@ func TestRemote_SpinloopDiscovery(t *testing.T) { dir := t.TempDir() t.Chdir(dir) - spinloopFile := fmt.Sprintf("PROVIDER openai-compatible\nENV SPINLOOP_REMOTE_START_URL=%s\nENV SPINLOOP_REMOTE_STOP_URL=%s\nENV SPINLOOP_REMOTE_REGION=eu-west-1\n", server.URL, server.URL) + spinloopFile := fmt.Sprintf("PROVIDER openai-compatible\nENV SPINLOOP_CLOUD_START_URL=%s\nENV SPINLOOP_CLOUD_STOP_URL=%s\nENV SPINLOOP_CLOUD_REGION=eu-west-1\n", server.URL, server.URL) if err := os.WriteFile("Spinloop", []byte(spinloopFile), 0o600); err != nil { t.Fatal(err) } @@ -545,7 +545,7 @@ func TestRemote_SpinloopDiscovery(t *testing.T) { // .env; it never selects an environment by itself. One with no ENV lines // supplies nothing, so with no --env or --fleet either the command fails on // the missing target, not on the Spinloop. -func TestRemote_ExplicitSpinloopDoesNotNameAnEnvironment(t *testing.T) { +func TestCloud_ExplicitSpinloopDoesNotNameAnEnvironment(t *testing.T) { isolateConfig(t) t.Chdir(t.TempDir()) if err := os.WriteFile("Spinloop", []byte("PROVIDER ollama\n"), 0o600); err != nil { @@ -557,7 +557,7 @@ func TestRemote_ExplicitSpinloopDoesNotNameAnEnvironment(t *testing.T) { } } -func TestRemote_SpinloopFallsBackToTheUserConfig(t *testing.T) { +func TestCloud_SpinloopFallsBackToTheUserConfig(t *testing.T) { isolateConfig(t) stubAWSEnv(t) server := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) { @@ -565,7 +565,7 @@ func TestRemote_SpinloopFallsBackToTheUserConfig(t *testing.T) { w.Write([]byte(`{"state":"stopped","healthy":false,"base_url":"http://198.51.100.1:8000/v1"}`)) })) defer server.Close() - writeRemoteConfig(t, server.URL) + writeCloudConfig(t, server.URL) t.Chdir(t.TempDir()) if err := os.WriteFile("Spinloop", []byte("PROVIDER ollama\n"), 0o600); err != nil { @@ -581,7 +581,7 @@ func TestRemote_SpinloopFallsBackToTheUserConfig(t *testing.T) { } } -func TestRemote_IgnoresLowercaseSpinloopFile(t *testing.T) { +func TestCloud_IgnoresLowercaseSpinloopFile(t *testing.T) { // On case-insensitive filesystems a stat of "Spinloop" matches a file named // "spinloop" (e.g. the built binary in this repo's root); discovery must not // try to parse it. @@ -592,7 +592,7 @@ func TestRemote_IgnoresLowercaseSpinloopFile(t *testing.T) { w.Write([]byte(`{"state":"stopped","healthy":false,"base_url":"http://198.51.100.1:8000/v1"}`)) })) defer server.Close() - writeRemoteConfig(t, server.URL) + writeCloudConfig(t, server.URL) t.Chdir(t.TempDir()) if err := os.WriteFile("spinloop", []byte{0xcf, 0xfa, 0xed, 0xfe}, 0o755); err != nil { @@ -608,7 +608,7 @@ func TestRemote_IgnoresLowercaseSpinloopFile(t *testing.T) { } } -func TestRemoteMetrics_Running(t *testing.T) { +func TestCloudMetrics_Running(t *testing.T) { isolateConfig(t) stubAWSEnv(t) server := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) { @@ -640,8 +640,8 @@ func TestRemoteMetrics_Running(t *testing.T) { }`)) })) defer server.Close() - writeRemoteConfig(t, server.URL) - t.Setenv("SPINLOOP_REMOTE_STATS_URL", server.URL) + writeCloudConfig(t, server.URL) + t.Setenv("SPINLOOP_CLOUD_STATS_URL", server.URL) out := captureStdout(t, func() { if err := cmdMetrics([]string{"--env", "default", "--format=table"}); err != nil { @@ -671,7 +671,7 @@ func TestRemoteMetrics_Running(t *testing.T) { } } -func TestRemoteMetrics_Stopped(t *testing.T) { +func TestCloudMetrics_Stopped(t *testing.T) { isolateConfig(t) stubAWSEnv(t) server := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) { @@ -684,8 +684,8 @@ func TestRemoteMetrics_Stopped(t *testing.T) { }`)) })) defer server.Close() - writeRemoteConfig(t, server.URL) - t.Setenv("SPINLOOP_REMOTE_STATS_URL", server.URL) + writeCloudConfig(t, server.URL) + t.Setenv("SPINLOOP_CLOUD_STATS_URL", server.URL) out := captureStdout(t, func() { if err := cmdMetrics([]string{"--env", "default", "--format=table"}); err != nil { @@ -700,7 +700,7 @@ func TestRemoteMetrics_Stopped(t *testing.T) { } } -func TestRemoteMetrics_WithErrors(t *testing.T) { +func TestCloudMetrics_WithErrors(t *testing.T) { isolateConfig(t) stubAWSEnv(t) server := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) { @@ -713,8 +713,8 @@ func TestRemoteMetrics_WithErrors(t *testing.T) { }`)) })) defer server.Close() - writeRemoteConfig(t, server.URL) - t.Setenv("SPINLOOP_REMOTE_STATS_URL", server.URL) + writeCloudConfig(t, server.URL) + t.Setenv("SPINLOOP_CLOUD_STATS_URL", server.URL) errOut := captureStderr(t, func() { if err := cmdMetrics([]string{"--env", "default"}); err != nil { @@ -728,7 +728,7 @@ func TestRemoteMetrics_WithErrors(t *testing.T) { } } -func TestRemoteMetrics_DefaultFormat(t *testing.T) { +func TestCloudMetrics_DefaultFormat(t *testing.T) { isolateConfig(t) stubAWSEnv(t) server := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) { @@ -748,8 +748,8 @@ func TestRemoteMetrics_DefaultFormat(t *testing.T) { }`)) })) defer server.Close() - writeRemoteConfig(t, server.URL) - t.Setenv("SPINLOOP_REMOTE_STATS_URL", server.URL) + writeCloudConfig(t, server.URL) + t.Setenv("SPINLOOP_CLOUD_STATS_URL", server.URL) out := captureStdout(t, func() { if err := cmdMetrics([]string{"--env", "default"}); err != nil { @@ -769,7 +769,7 @@ func TestRemoteMetrics_DefaultFormat(t *testing.T) { } } -func TestRemoteMetrics_JsonFormat(t *testing.T) { +func TestCloudMetrics_JsonFormat(t *testing.T) { isolateConfig(t) stubAWSEnv(t) server := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) { @@ -785,8 +785,8 @@ func TestRemoteMetrics_JsonFormat(t *testing.T) { }`)) })) defer server.Close() - writeRemoteConfig(t, server.URL) - t.Setenv("SPINLOOP_REMOTE_STATS_URL", server.URL) + writeCloudConfig(t, server.URL) + t.Setenv("SPINLOOP_CLOUD_STATS_URL", server.URL) out := captureStdout(t, func() { if err := cmdMetrics([]string{"--env", "default", "--format=json"}); err != nil { @@ -813,7 +813,7 @@ func TestRemoteMetrics_JsonFormat(t *testing.T) { } } -func TestRemoteMetrics_JsonFormatWithCost(t *testing.T) { +func TestCloudMetrics_JsonFormatWithCost(t *testing.T) { isolateConfig(t) stubAWSEnv(t) server := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) { @@ -827,8 +827,8 @@ func TestRemoteMetrics_JsonFormatWithCost(t *testing.T) { }`)) })) defer server.Close() - writeRemoteConfig(t, server.URL) - t.Setenv("SPINLOOP_REMOTE_STATS_URL", server.URL) + writeCloudConfig(t, server.URL) + t.Setenv("SPINLOOP_CLOUD_STATS_URL", server.URL) out := captureStdout(t, func() { if err := cmdMetrics([]string{"--env", "default", "--format=json", "--cost"}); err != nil { @@ -847,10 +847,10 @@ func TestRemoteMetrics_JsonFormatWithCost(t *testing.T) { } } -func TestRemoteMetrics_InvalidFormat(t *testing.T) { +func TestCloudMetrics_InvalidFormat(t *testing.T) { isolateConfig(t) stubAWSEnv(t) - writeRemoteConfig(t, "http://localhost:0") + writeCloudConfig(t, "http://localhost:0") err := cmdMetrics([]string{"--env", "default", "--format=csv"}) if err == nil || !strings.Contains(err.Error(), "format") { @@ -858,7 +858,7 @@ func TestRemoteMetrics_InvalidFormat(t *testing.T) { } } -func TestRemoteMetrics_BarFormat(t *testing.T) { +func TestCloudMetrics_BarFormat(t *testing.T) { isolateConfig(t) stubAWSEnv(t) server := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) { @@ -878,8 +878,8 @@ func TestRemoteMetrics_BarFormat(t *testing.T) { }`)) })) defer server.Close() - writeRemoteConfig(t, server.URL) - t.Setenv("SPINLOOP_REMOTE_STATS_URL", server.URL) + writeCloudConfig(t, server.URL) + t.Setenv("SPINLOOP_CLOUD_STATS_URL", server.URL) out := captureStdout(t, func() { err := cmdMetrics([]string{"--env", "default", "--format=bar"}) @@ -911,7 +911,7 @@ func TestRemoteMetrics_BarFormat(t *testing.T) { } } -func TestRemoteMetrics_BarFormatStopped(t *testing.T) { +func TestCloudMetrics_BarFormatStopped(t *testing.T) { isolateConfig(t) stubAWSEnv(t) server := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) { @@ -923,8 +923,8 @@ func TestRemoteMetrics_BarFormatStopped(t *testing.T) { }`)) })) defer server.Close() - writeRemoteConfig(t, server.URL) - t.Setenv("SPINLOOP_REMOTE_STATS_URL", server.URL) + writeCloudConfig(t, server.URL) + t.Setenv("SPINLOOP_CLOUD_STATS_URL", server.URL) out := captureStdout(t, func() { err := cmdMetrics([]string{"--env", "default", "--format=bar"}) @@ -978,7 +978,7 @@ func TestFormatBytes(t *testing.T) { } } -func TestRemoteMetrics_WatchMode(t *testing.T) { +func TestCloudMetrics_WatchMode(t *testing.T) { isolateConfig(t) stubAWSEnv(t) callCount := 0 @@ -994,8 +994,8 @@ func TestRemoteMetrics_WatchMode(t *testing.T) { w.Write([]byte(`{"environment": "dev", "state": "running", "uptimeSeconds": 100}`)) })) defer server.Close() - writeRemoteConfig(t, server.URL) - t.Setenv("SPINLOOP_REMOTE_STATS_URL", server.URL) + writeCloudConfig(t, server.URL) + t.Setenv("SPINLOOP_CLOUD_STATS_URL", server.URL) // Short interval so the test completes quickly. oldInterval := metricsWatchInterval @@ -1020,7 +1020,7 @@ func TestRemoteMetrics_WatchMode(t *testing.T) { } } -func TestRemoteMetrics_WatchShortFlag(t *testing.T) { +func TestCloudMetrics_WatchShortFlag(t *testing.T) { isolateConfig(t) stubAWSEnv(t) callCount := 0 @@ -1035,8 +1035,8 @@ func TestRemoteMetrics_WatchShortFlag(t *testing.T) { w.Write([]byte(`{"environment": "dev", "state": "stopped"}`)) })) defer server.Close() - writeRemoteConfig(t, server.URL) - t.Setenv("SPINLOOP_REMOTE_STATS_URL", server.URL) + writeCloudConfig(t, server.URL) + t.Setenv("SPINLOOP_CLOUD_STATS_URL", server.URL) oldInterval := metricsWatchInterval metricsWatchInterval = 10 * time.Millisecond @@ -1056,7 +1056,7 @@ func TestRemoteMetrics_WatchShortFlag(t *testing.T) { } } -func TestRemoteMetrics_MultiGPU(t *testing.T) { +func TestCloudMetrics_MultiGPU(t *testing.T) { isolateConfig(t) stubAWSEnv(t) server := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) { @@ -1071,8 +1071,8 @@ func TestRemoteMetrics_MultiGPU(t *testing.T) { }`)) })) defer server.Close() - writeRemoteConfig(t, server.URL) - t.Setenv("SPINLOOP_REMOTE_STATS_URL", server.URL) + writeCloudConfig(t, server.URL) + t.Setenv("SPINLOOP_CLOUD_STATS_URL", server.URL) out := captureStdout(t, func() { if err := cmdMetrics([]string{"--env", "default", "--format=table"}); err != nil { @@ -1086,7 +1086,7 @@ func TestRemoteMetrics_MultiGPU(t *testing.T) { } } -func TestRemoteMetrics_JsonStopped(t *testing.T) { +func TestCloudMetrics_JsonStopped(t *testing.T) { isolateConfig(t) stubAWSEnv(t) server := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) { @@ -1094,8 +1094,8 @@ func TestRemoteMetrics_JsonStopped(t *testing.T) { w.Write([]byte(`{"environment": "dev", "state": "stopped"}`)) })) defer server.Close() - writeRemoteConfig(t, server.URL) - t.Setenv("SPINLOOP_REMOTE_STATS_URL", server.URL) + writeCloudConfig(t, server.URL) + t.Setenv("SPINLOOP_CLOUD_STATS_URL", server.URL) out := captureStdout(t, func() { if err := cmdMetrics([]string{"--env", "default", "--format=json"}); err != nil { @@ -1115,7 +1115,7 @@ func TestRemoteMetrics_JsonStopped(t *testing.T) { } } -func TestRemoteMetrics_JsonWithErrors(t *testing.T) { +func TestCloudMetrics_JsonWithErrors(t *testing.T) { isolateConfig(t) stubAWSEnv(t) server := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) { @@ -1128,8 +1128,8 @@ func TestRemoteMetrics_JsonWithErrors(t *testing.T) { }`)) })) defer server.Close() - writeRemoteConfig(t, server.URL) - t.Setenv("SPINLOOP_REMOTE_STATS_URL", server.URL) + writeCloudConfig(t, server.URL) + t.Setenv("SPINLOOP_CLOUD_STATS_URL", server.URL) out := captureStdout(t, func() { if err := cmdMetrics([]string{"--env", "default", "--format=json"}); err != nil { @@ -1152,7 +1152,7 @@ func TestRemoteMetrics_JsonWithErrors(t *testing.T) { // The start probe runs after the endpoint is ready. When the probe connects, // no warning is printed. -func TestRemoteStart_ProbeSucceedsNoWarning(t *testing.T) { +func TestCloudStart_ProbeSucceedsNoWarning(t *testing.T) { isolateConfig(t) stubAWSEnv(t) @@ -1179,11 +1179,11 @@ func TestRemoteStart_ProbeSucceedsNoWarning(t *testing.T) { w.Write([]byte(fmt.Sprintf(`{"state":"ready","base_url":"%s","api_key":"sk-test"}`, baseURL))) })) defer server.Close() - writeRemoteConfig(t, server.URL) + writeCloudConfig(t, server.URL) stderr := captureStderr(t, func() { - if err := cmdRemoteStart([]string{"--env", "default"}); err != nil { - t.Errorf("cmdRemoteStart: %v", err) + if err := cmdCloudStart([]string{"--env", "default"}); err != nil { + t.Errorf("cmdCloudStart: %v", err) } }) if strings.Contains(stderr, "not reachable") { @@ -1192,7 +1192,7 @@ func TestRemoteStart_ProbeSucceedsNoWarning(t *testing.T) { } // When the probe fails, start warns but still exits 0. -func TestRemoteStart_ProbeFailsWarns(t *testing.T) { +func TestCloudStart_ProbeFailsWarns(t *testing.T) { isolateConfig(t) stubAWSEnv(t) @@ -1200,9 +1200,9 @@ func TestRemoteStart_ProbeFailsWarns(t *testing.T) { detectPublicCIDRFn = func(context.Context) (string, error) { return "203.0.113.5/32", nil } t.Cleanup(func() { detectPublicCIDRFn = origDetect }) - origProbe := remote.ProbeTimeout - remote.ProbeTimeout = 100 * time.Millisecond - t.Cleanup(func() { remote.ProbeTimeout = origProbe }) + origProbe := cloud.ProbeTimeout + cloud.ProbeTimeout = 100 * time.Millisecond + t.Cleanup(func() { cloud.ProbeTimeout = origProbe }) baseURL := "http://192.0.2.1:8000/v1" // unreachable server := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) { @@ -1210,10 +1210,10 @@ func TestRemoteStart_ProbeFailsWarns(t *testing.T) { w.Write([]byte(fmt.Sprintf(`{"state":"ready","base_url":"%s","api_key":"sk-test"}`, baseURL))) })) defer server.Close() - writeRemoteConfig(t, server.URL) + writeCloudConfig(t, server.URL) errOut := captureStderr(t, func() { - err := cmdRemoteStart([]string{"--env", "default"}) + err := cmdCloudStart([]string{"--env", "default"}) if err != nil { t.Fatalf("start should exit 0 after a probe warning, got %v", err) } @@ -1228,7 +1228,7 @@ func TestRemoteStart_ProbeFailsWarns(t *testing.T) { } // When the probe fails and IP detection also fails, the hint uses a placeholder. -func TestRemoteStart_ProbeFailsIPDetectFails(t *testing.T) { +func TestCloudStart_ProbeFailsIPDetectFails(t *testing.T) { isolateConfig(t) stubAWSEnv(t) @@ -1236,9 +1236,9 @@ func TestRemoteStart_ProbeFailsIPDetectFails(t *testing.T) { detectPublicCIDRFn = func(context.Context) (string, error) { return "", fmt.Errorf("network error") } t.Cleanup(func() { detectPublicCIDRFn = origDetect }) - origProbe := remote.ProbeTimeout - remote.ProbeTimeout = 100 * time.Millisecond - t.Cleanup(func() { remote.ProbeTimeout = origProbe }) + origProbe := cloud.ProbeTimeout + cloud.ProbeTimeout = 100 * time.Millisecond + t.Cleanup(func() { cloud.ProbeTimeout = origProbe }) baseURL := "http://192.0.2.1:8000/v1" server := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) { @@ -1246,10 +1246,10 @@ func TestRemoteStart_ProbeFailsIPDetectFails(t *testing.T) { w.Write([]byte(fmt.Sprintf(`{"state":"ready","base_url":"%s","api_key":"sk-test"}`, baseURL))) })) defer server.Close() - writeRemoteConfig(t, server.URL) + writeCloudConfig(t, server.URL) errOut := captureStderr(t, func() { - err := cmdRemoteStart([]string{"--env", "default"}) + err := cmdCloudStart([]string{"--env", "default"}) if err != nil { t.Fatalf("start should exit 0 even when probe and IP detection both fail, got %v", err) } @@ -1263,12 +1263,12 @@ func TestRemoteStart_ProbeFailsIPDetectFails(t *testing.T) { } } -// TestRemoteMetrics_WatchBuffersBeforeClear verifies the fetch-before-clear +// TestCloudMetrics_WatchBuffersBeforeClear verifies the fetch-before-clear // invariant: metrics are rendered into a buffer first, then the screen is // cleared and the buffer is written. This eliminates the blank-frame flash // that occurs when you clear the screen before you have content to show. -// Regression for: io.Writer refactor was lost when remote.go was reverted. -func TestRemoteMetrics_WatchBuffersBeforeClear(t *testing.T) { +// Regression for: io.Writer refactor was lost when cloud.go was reverted. +func TestCloudMetrics_WatchBuffersBeforeClear(t *testing.T) { isolateConfig(t) stubAWSEnv(t) callCount := 0 @@ -1283,8 +1283,8 @@ func TestRemoteMetrics_WatchBuffersBeforeClear(t *testing.T) { w.Write([]byte(`{"environment": "dev", "state": "running", "uptimeSeconds": 10}`)) })) defer server.Close() - writeRemoteConfig(t, server.URL) - t.Setenv("SPINLOOP_REMOTE_STATS_URL", server.URL) + writeCloudConfig(t, server.URL) + t.Setenv("SPINLOOP_CLOUD_STATS_URL", server.URL) oldInterval := metricsWatchInterval metricsWatchInterval = 50 * time.Millisecond @@ -1326,8 +1326,8 @@ func TestRemoteMetrics_WatchBuffersBeforeClear(t *testing.T) { } } -// TestRemoteKeep sets the retention deadline. -func TestRemoteKeep_PrintsDeadline(t *testing.T) { +// TestCloudKeep sets the retention deadline. +func TestCloudKeep_PrintsDeadline(t *testing.T) { isolateConfig(t) stubAWSEnv(t) server := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) { @@ -1341,19 +1341,19 @@ func TestRemoteKeep_PrintsDeadline(t *testing.T) { w.Write([]byte(`{"environment":"test"}`)) })) defer server.Close() - writeRemoteConfig(t, server.URL) + writeCloudConfig(t, server.URL) // Also need to write the update URL. - path := must1(remote.EnvConfigPath("default")) + path := must1(cloud.EnvConfigPath("default")) data := must1(os.ReadFile(path)) - var cfg remote.Config + var cfg cloud.Config json.Unmarshal(data, &cfg) cfg.UpdateURL = server.URL os.WriteFile(path, must1(json.Marshal(cfg)), 0o600) out := captureStdout(t, func() { - if err := cmdRemoteKeep([]string{"--env", "default", "4h"}); err != nil { - t.Errorf("cmdRemoteKeep: %v", err) + if err := cmdCloudKeep([]string{"--env", "default", "4h"}); err != nil { + t.Errorf("cmdCloudKeep: %v", err) } }) if !strings.Contains(out, "retain until:") { @@ -1361,24 +1361,24 @@ func TestRemoteKeep_PrintsDeadline(t *testing.T) { } } -// TestRemoteKeep_MissingDuration fails. -func TestRemoteKeep_MissingDuration(t *testing.T) { - err := cmdRemoteKeep([]string{"--env", "default"}) +// TestCloudKeep_MissingDuration fails. +func TestCloudKeep_MissingDuration(t *testing.T) { + err := cmdCloudKeep([]string{"--env", "default"}) if err == nil || !strings.Contains(err.Error(), "usage") { t.Errorf("expected usage error, got %v", err) } } -// TestRemoteKeep_InvalidDuration fails. -func TestRemoteKeep_InvalidDuration(t *testing.T) { - err := cmdRemoteKeep([]string{"--env", "default", "4hours"}) +// TestCloudKeep_InvalidDuration fails. +func TestCloudKeep_InvalidDuration(t *testing.T) { + err := cmdCloudKeep([]string{"--env", "default", "4hours"}) if err == nil || !strings.Contains(err.Error(), "invalid duration") { t.Errorf("expected duration parse error, got %v", err) } } -// TestRemoteStart_KeepFlag passes the retainUntil parameter. -func TestRemoteStart_KeepFlag(t *testing.T) { +// TestCloudStart_KeepFlag passes the retainUntil parameter. +func TestCloudStart_KeepFlag(t *testing.T) { isolateConfig(t) stubAWSEnv(t) var gotRetainUntil string @@ -1388,11 +1388,11 @@ func TestRemoteStart_KeepFlag(t *testing.T) { w.Write([]byte(`{"state":"ready","base_url":"http://198.51.100.1:8000/v1","api_key":"sk-test"}`)) })) defer server.Close() - writeRemoteConfig(t, server.URL) + writeCloudConfig(t, server.URL) // Probe reachability will fail, but that's stderr and doesn't affect the test. out := captureStdout(t, func() { - cmdRemoteStart([]string{"--env", "default", "--keep", "2h"}) + cmdCloudStart([]string{"--env", "default", "--keep", "2h"}) }) // The keep deadline should be reported on stderr (via progress). // We can check that the request included the retainUntil parameter. diff --git a/cmd/spinloop/code.go b/cmd/spinloop/code.go index c9acb67b..9a18a70e 100644 --- a/cmd/spinloop/code.go +++ b/cmd/spinloop/code.go @@ -3,9 +3,9 @@ package main import ( "github.com/spf13/cobra" + "github.com/spinloop-ai/spinloop/internal/cloud" "github.com/spinloop-ai/spinloop/internal/fleet" "github.com/spinloop-ai/spinloop/internal/harness" - "github.com/spinloop-ai/spinloop/internal/remote" "github.com/spinloop-ai/spinloop/internal/spinloop" ) @@ -67,16 +67,16 @@ func openLaunchCmd(use, short, long string) *cobra.Command { // A .env beside the applied Spinloop is where its keys live, so the // launched agent is given the same ones. Without a Spinloop there is // no such file and only the environment (plus any provider key - // spinloop resolves) is passed on. remoteResp carries the live key of - // a remote endpoint, fetched while applying so the config is + // spinloop resolves) is passed on. cloudResp carries the live key of + // a cloud endpoint, fetched while applying so the config is // written knowing it will be there. var envDir string var sel spinloop.Selection - var remoteResp *remote.Response + var cloudResp *cloud.Response var choice *fleet.Choice if spinloopPath.set { var err error - sel, envDir, remoteResp, choice, err = applyBeforeLaunch(spinloopPath, providers, h, rest, route) + sel, envDir, cloudResp, choice, err = applyBeforeLaunch(spinloopPath, providers, h, rest, route) if err != nil { return err } @@ -86,7 +86,7 @@ func openLaunchCmd(use, short, long string) *cobra.Command { // actually deployed there instead of doing nothing with the // flag. var err error - sel, envDir, remoteResp, choice, err = applyFromEnvironment(providers, h, route) + sel, envDir, cloudResp, choice, err = applyFromEnvironment(providers, h, route) if err != nil { return err } @@ -96,12 +96,12 @@ func openLaunchCmd(use, short, long string) *cobra.Command { // Spinloop; a fleet that does not still does, and this says // so. var err error - sel, envDir, remoteResp, choice, err = applyFromGateway(providers, h, route) + sel, envDir, cloudResp, choice, err = applyFromGateway(providers, h, route) if err != nil { return err } } - return launchAgent(h, rest, providers, envDir, remoteResp, sel, spinloopPath.set, choice) + return launchAgent(h, rest, providers, envDir, cloudResp, sel, spinloopPath.set, choice) }, } fs := c.Flags() diff --git a/cmd/spinloop/commands.go b/cmd/spinloop/commands.go index b2de2ebb..c30837ca 100644 --- a/cmd/spinloop/commands.go +++ b/cmd/spinloop/commands.go @@ -14,10 +14,10 @@ import ( "github.com/spf13/cobra" "github.com/spf13/pflag" + "github.com/spinloop-ai/spinloop/internal/cloud" "github.com/spinloop-ai/spinloop/internal/fleet" "github.com/spinloop-ai/spinloop/internal/harness" "github.com/spinloop-ai/spinloop/internal/opencode" - "github.com/spinloop-ai/spinloop/internal/remote" "github.com/spinloop-ai/spinloop/internal/spinloop" ) @@ -91,7 +91,7 @@ has been.`, ) root.AddCommand(fleetCmd()) - root.AddCommand(remoteCmd()) + root.AddCommand(cloudCmd()) defaultFlagCompletions(root) return root } @@ -233,11 +233,11 @@ for a custom one. Each subcommand's --help says what it does.`, // args forwarded, and the environment the apply step reported as the source of // every key the agent is given. Both launch commands end here, so the agent a // launch is given can only differ the way the apply that preceded it did. -func launchAgent(h harness.Harness, rest []string, providers, envDir string, remoteResp *remote.Response, sel spinloop.Selection, worn bool, choice *fleet.Choice) error { - // The resolver the launch uses knows the remote key too, so every +func launchAgent(h harness.Harness, rest []string, providers, envDir string, cloudResp *cloud.Response, sel spinloop.Selection, worn bool, choice *fleet.Choice) error { + // The resolver the launch uses knows the cloud key too, so every // key the agent is given comes from the same place the apply step // reported. - resolveKey := remoteLaunchResolver(opencode.EnvResolver(envDir), remoteResp) + resolveKey := cloudLaunchResolver(opencode.EnvResolver(envDir), cloudResp) // Launch the harness, forwarding stdio and any trailing args. bin := h.Command() @@ -245,9 +245,9 @@ func launchAgent(h harness.Harness, rest []string, providers, envDir string, rem cmd.Stdin = os.Stdin cmd.Stdout = os.Stdout cmd.Stderr = os.Stderr - cmd.Env = harnessEnv(providers, resolveKey, remoteResp) + cmd.Env = harnessEnv(providers, resolveKey, cloudResp) // A routed launch points the agent at the node that was chosen. As - // on the remote path, an explicit setting in the environment already + // on the cloud path, an explicit setting in the environment already // won: routing fills what is unset rather than overriding a // deliberate choice. if choice != nil { @@ -260,7 +260,7 @@ func launchAgent(h harness.Harness, rest []string, providers, envDir string, rem // agent: its adjacent .env fills any gaps left above, and its ENV // instructions override everything. These shape only the child's // environment — spinloop never mutates its own — and follow the same - // precedence the remote commands use: ENV > process environment > + // precedence the cloud commands use: ENV > process environment > // .env. if worn { cmd.Env = overlayLocalEnv(cmd.Env, sel, envDir) @@ -313,7 +313,7 @@ func versionCmd() *cobra.Command { } } -// groupFallback is the RunE a command group (fleet, remote, seed) gets: bare, +// groupFallback is the RunE a command group (fleet, cloud, seed) gets: bare, // it shows the group's own help — the one cobra generates from the tree, so // its subcommand list cannot drift from the tree — and a word that is not a // subcommand is cobra's own unknown-command error. The help sentinel is @@ -353,9 +353,9 @@ var movedSubcommands = map[string]string{ "fleet dashboard": "dashboard", "fleet metrics": "metrics", "fleet logs": "logs", - "remote status": "status --env ", - "remote metrics": "metrics --env ", - "remote logs": "logs --env ", + "cloud status": "status --env ", + "cloud metrics": "metrics --env ", + "cloud logs": "logs --env ", } // fleetCmd builds the fleet parent and its subcommands. The parent does @@ -368,7 +368,7 @@ func fleetCmd() *cobra.Command { default; --fleet names another). Observation is fleet-wide (status, metrics, logs, and dashboard — the live tiled view); start and stop take one or more node names, or --all for the whole fleet, and with neither they list the -fleet and touch nothing; deploy provisions kind: remote nodes' AWS +fleet and touch nothing; deploy provisions kind: cloud nodes' AWS environments the same way. A node that fails is a rendered row, never an error — only a problem with the fleet file itself fails a command.`, Args: groupArgs, @@ -385,35 +385,34 @@ error — only a problem with the fleet file itself fails a command.`, return fleet } -// remoteCmd builds the remote parent and its subcommands. The parent does +// cloudCmd builds the cloud parent and its subcommands. The parent does // nothing itself — see groupFallback. -func remoteCmd() *cobra.Command { - remote := &cobra.Command{ - Use: "remote", - Short: "control the remote GPU inference instance", +func cloudCmd() *cobra.Command { + cloud := &cobra.Command{ + Use: "cloud", + Short: "control the cloud GPU inference instance", Long: `runs the model on a cloud GPU that exists only while you use it, from -the same Spinloop. The endpoint's URLs come from the Spinloop's REMOTE — a bare -name selects an environment under ~/.config/spinloop/remotes//, a path -names a file — falling back to the default environment. Each subcommand's ---help says what that step does.`, +the same Spinloop. Each environment is registered under +~/.config/spinloop/clouds// and selected with --env . Each +subcommand's --help says what that step does.`, Args: groupArgs, SilenceErrors: true, SilenceUsage: true, RunE: groupFallback, } - remote.AddCommand( - remoteBootstrapCmd(), - remoteAuthCmd(), - remoteBakeCmd(), - remoteStartCmd(), - remotePauseCmd(), - remoteRestartCmd(), - remoteStopCmd(), - remoteDeployCmd(), - remoteSeedCmd(), - remoteEnvCmd(), - remoteListCmd(), - remoteKeepCmd(), + cloud.AddCommand( + cloudBootstrapCmd(), + cloudAuthCmd(), + cloudBakeCmd(), + cloudStartCmd(), + cloudPauseCmd(), + cloudRestartCmd(), + cloudStopCmd(), + cloudDeployCmd(), + cloudSeedCmd(), + cloudEnvCmd(), + cloudListCmd(), + cloudKeepCmd(), ) - return remote + return cloud } diff --git a/cmd/spinloop/complete.go b/cmd/spinloop/complete.go index 258323c8..2e3ca2a7 100644 --- a/cmd/spinloop/complete.go +++ b/cmd/spinloop/complete.go @@ -18,12 +18,12 @@ import ( "github.com/spf13/cobra" "github.com/spf13/pflag" "github.com/spinloop-ai/spinloop/internal/catalog" + "github.com/spinloop-ai/spinloop/internal/cloud" "github.com/spinloop-ai/spinloop/internal/config" "github.com/spinloop-ai/spinloop/internal/daemon" "github.com/spinloop-ai/spinloop/internal/discovery" "github.com/spinloop-ai/spinloop/internal/harness" "github.com/spinloop-ai/spinloop/internal/opencode" - "github.com/spinloop-ai/spinloop/internal/remote" ) // completionShells lists the supported shells in a stable order, for the @@ -197,7 +197,7 @@ func compNoValues(_ *cobra.Command, _ []string, _ string) ([]string, cobra.Shell // compEnvs offers the registered environment names. An unreadable registry // yields no candidates rather than an error, as completion must never fail. func compEnvs(_ *cobra.Command, _ []string, _ string) ([]string, cobra.ShellCompDirective) { - envs, err := remote.ListEnvironments() + envs, err := cloud.ListEnvironments() if err != nil { return nil, cobra.ShellCompDirectiveNoFileComp } @@ -252,7 +252,7 @@ func fileSlot(_ *cobra.Command, args []string, _ string) ([]string, cobra.ShellC return nil, cobra.ShellCompDirectiveDefault } -// keepSlot is `remote keep [spinloop]`: a duration first, then the +// keepSlot is `cloud keep [spinloop]`: a duration first, then the // optional Spinloop. func keepSlot(_ *cobra.Command, args []string, _ string) ([]string, cobra.ShellCompDirective) { switch len(args) { diff --git a/cmd/spinloop/complete_test.go b/cmd/spinloop/complete_test.go index 7f88f7ad..50f23641 100644 --- a/cmd/spinloop/complete_test.go +++ b/cmd/spinloop/complete_test.go @@ -256,7 +256,7 @@ func TestComplete_SpinloopCommandsOfferAliasesAndPaths(t *testing.T) { slots := [][]string{ {"harness", "apply", ""}, {"harness", "unapply", ""}, {"serve", ""}, {"alias", ""}, - {"harness", "open", ""}, {"fleet", "route", ""}, {"remote", "deploy", ""}, {"remote", "start", ""}, + {"harness", "open", ""}, {"fleet", "route", ""}, {"cloud", "deploy", ""}, {"cloud", "start", ""}, } for _, words := range slots { got, directive := complete(t, words...) diff --git a/cmd/spinloop/dashboard.go b/cmd/spinloop/dashboard.go index ce90a286..9aa2b4b2 100644 --- a/cmd/spinloop/dashboard.go +++ b/cmd/spinloop/dashboard.go @@ -21,12 +21,12 @@ resource gauges, the token counters — repainted on an interval. A single registered environment (--env) opens as a board of one. The view is read-only apart from four keys: s starts the selected node, k -keeps a remote environment for a duration you type — it asks how long, +keeps a cloud environment for a duration you type — it asks how long, pre-filled with 4h, and reports the deadline the control plane set when the keep is done — a abandons a start still in flight on it (the wait ends, the node is free again — a wake the cloud is carrying goes on), x stops it after a confirmation. The arrow keys move the selection, r forces a refresh, q or -Ctrl+C leaves. The keep key shows only for a node that can be kept — a remote +Ctrl+C leaves. The keep key shows only for a node that can be kept — a cloud environment — and a kept environment's tile and detail view carry its deadline beside the last-active line, whatever the engine's state. diff --git a/cmd/spinloop/dashboard_keep_test.go b/cmd/spinloop/dashboard_keep_test.go index 4cc8da0c..0f784dfe 100644 --- a/cmd/spinloop/dashboard_keep_test.go +++ b/cmd/spinloop/dashboard_keep_test.go @@ -71,7 +71,7 @@ func (n *keeperDashNode) Keep(ctx context.Context, d time.Duration) (string, err // test drives. func keeperModel(node *keeperDashNode) *dashModel { return &dashModel{ - entries: []dashEntry{{name: node.f.Name(), kind: fleet.KindRemote, node: node}}, + entries: []dashEntry{{name: node.f.Name(), kind: fleet.KindCloud, node: node}}, results: []fleet.NodeResult{{Name: node.f.Name()}}, actions: make([]dashAction, 1), width: 120, height: 40, @@ -313,7 +313,7 @@ func TestDashKeepFailureShowsItsReason(t *testing.T) { err error want string }{ - {errors.New("no update_url configured: the remote deployment needs to be updated for keep support"), "no update_url"}, + {errors.New("no update_url configured: the cloud deployment needs to be updated for keep support"), "no update_url"}, {errors.New("keep returned HTTP 404: no running instance"), "no running instance"}, } { t.Run(tc.want, func(t *testing.T) { @@ -359,16 +359,16 @@ func TestDashAbortDrivesNothingOnAKeep(t *testing.T) { } } -// The keep hint shows only where the key would drive something: a remote node -// shows it, a local node hides it, and a busy remote node hides it. The start -// and stop entries sit beside it by the node's own state: a stopped remote +// The keep hint shows only where the key would drive something: a cloud node +// shows it, a local node hides it, and a busy cloud node hides it. The start +// and stop entries sit beside it by the node's own state: a stopped cloud // environment shows keep and start, a running one shows keep and stop, and a // busy one shows neither. func TestDashKeepHintOnlyWhereItDrivesSomething(t *testing.T) { read := func(state string) fleet.NodeResult { return fleet.NodeResult{Name: "env", Outcome: fleet.OutcomeOK, Metrics: metrics.Stats{State: state}} } - t.Run("stopped remote shows keep and start", func(t *testing.T) { + t.Run("stopped cloud shows keep and start", func(t *testing.T) { node := &keeperDashNode{f: newFakeDashNode("stopped")} m := keeperModel(node) m.results[0] = read("stopped") @@ -379,7 +379,7 @@ func TestDashKeepHintOnlyWhereItDrivesSomething(t *testing.T) { t.Errorf("detail hint:\ngot: %q\nwant: %q", got, want) } }) - t.Run("running remote shows keep and stop", func(t *testing.T) { + t.Run("running cloud shows keep and stop", func(t *testing.T) { node := &keeperDashNode{f: newFakeDashNode("running")} m := keeperModel(node) m.results[0] = read("running") @@ -402,7 +402,7 @@ func TestDashKeepHintOnlyWhereItDrivesSomething(t *testing.T) { t.Errorf("grid hint:\ngot: %q\nwant: %q", got, want) } }) - t.Run("busy remote hides it", func(t *testing.T) { + t.Run("busy cloud hides it", func(t *testing.T) { node := &keeperDashNode{f: newFakeDashNode("stopped")} m := keeperModel(node) m.results[0] = read("stopped") diff --git a/cmd/spinloop/dashboard_model.go b/cmd/spinloop/dashboard_model.go index 6c0f4cad..8604362b 100644 --- a/cmd/spinloop/dashboard_model.go +++ b/cmd/spinloop/dashboard_model.go @@ -25,12 +25,12 @@ import ( // variable so a test never waits for a slow node. var dashboardRefreshInterval = 5 * time.Second -// dashboardRemoteRefreshInterval is the cadence for kind: remote +// dashboardCloudRefreshInterval is the cadence for kind: cloud // environments instead. Each of their statuses is a signed call through the // cloud control plane — a Lambda invocation, not a socket on the // sideboard — so the board refreshes the local machines on the tick and the // cloud environments on this slower deadline. Variable, for the same reason. -var dashboardRemoteRefreshInterval = 60 * time.Second +var dashboardCloudRefreshInterval = 60 * time.Second // dashEntry is one fleet-file entry and what it resolved to. Node is nil // when the entry could not become a node at all — its token reference names @@ -122,7 +122,7 @@ type dashTickMsg time.Time // them are drawn is decided by each reading's own time, not by the round's: // see the message's handling in Update. type dashRefreshMsg struct { - remote bool // the cloud group's round + cloud bool // the cloud group's round idx []int results []fleet.NodeResult } @@ -233,7 +233,7 @@ func (m *dashModel) Update(msg tea.Msg) (tea.Model, tea.Cmd) { // than overlapping the next. return m, tea.Batch(append([]tea.Cmd{dashTickCmd()}, m.startRounds()...)...) case dashRefreshMsg: - if msg.remote { + if msg.cloud { m.slowBusy = false } else { m.fastBusy = false @@ -418,8 +418,8 @@ func (m *dashModel) Update(msg tea.Msg) (tea.Model, tea.Cmd) { // group with nothing due starts nothing. func (m *dashModel) startRounds() []tea.Cmd { var cmds []tea.Cmd - for _, remote := range []bool{false, true} { - if cmd := m.refreshRemoteGroup(remote); cmd != nil { + for _, cloud := range []bool{false, true} { + if cmd := m.refreshCloudGroup(cloud); cmd != nil { cmds = append(cmds, cmd) } } @@ -433,8 +433,8 @@ func (m *dashModel) startRounds() []tea.Cmd { // scale of minutes — hold for an environment nobody is touching and hold for // neither one the operator has just started. func dashNodeInterval(kind string, a dashAction) time.Duration { - if a.verb == "" && kind == fleet.KindRemote { - return dashboardRemoteRefreshInterval + if a.verb == "" && kind == fleet.KindCloud { + return dashboardCloudRefreshInterval } return dashboardRefreshInterval } @@ -472,9 +472,9 @@ func (m *dashModel) scheduleRead(i int, at time.Time) { m.nextReadAt[i] = at } -// refreshRemoteGroup starts one round over the due nodes of one group of live -// nodes — the local daemon machines (remote false) or the cloud environments -// (remote true) — and returns the round's command, or nil when nothing in the +// refreshCloudGroup starts one round over the due nodes of one group of live +// nodes — the local daemon machines (cloud false) or the cloud environments +// (cloud true) — and returns the round's command, or nil when nothing in the // group is due or a round is already in flight there. Starting the round // spends each read node's deadline: each moves to one of its own intervals // away, so a node the operator is acting on comes round again on the short @@ -486,12 +486,12 @@ func (m *dashModel) scheduleRead(i int, at time.Time) { // an action must not shorten what the call is given to answer in. Each node // answers independently — the fan-out calls them concurrently — so one slow // node delays no other, and a slow cloud round stretches only its own group. -func (m *dashModel) refreshRemoteGroup(remote bool) tea.Cmd { +func (m *dashModel) refreshCloudGroup(cloud bool) tea.Cmd { now := time.Now() idx := make([]int, 0, len(m.entries)) nodes := make([]fleet.Node, 0, len(m.entries)) for i, e := range m.entries { - if e.node == nil || (e.kind == fleet.KindRemote) != remote { + if e.node == nil || (e.kind == fleet.KindCloud) != cloud { continue } if !m.isDue(i, now) { @@ -503,7 +503,7 @@ func (m *dashModel) refreshRemoteGroup(remote bool) tea.Cmd { if len(nodes) == 0 { return nil } - if remote { + if cloud { if m.slowBusy { return nil } @@ -517,17 +517,17 @@ func (m *dashModel) refreshRemoteGroup(remote bool) tea.Cmd { for _, i := range idx { m.scheduleRead(i, now.Add(dashNodeInterval(m.entries[i].kind, m.actions[i]))) } - interval := m.intervalFor(remote) + interval := m.intervalFor(cloud) return func() tea.Msg { ctx, cancel := context.WithTimeout(context.Background(), interval) defer cancel() - return dashRefreshMsg{remote: remote, idx: idx, results: fleet.FanOutNodes(ctx, fleet.MetricsCall, nodes)} + return dashRefreshMsg{cloud: cloud, idx: idx, results: fleet.FanOutNodes(ctx, fleet.MetricsCall, nodes)} } } -func (m *dashModel) intervalFor(remote bool) time.Duration { - if remote { - return dashboardRemoteRefreshInterval +func (m *dashModel) intervalFor(cloud bool) time.Duration { + if cloud { + return dashboardCloudRefreshInterval } return dashboardRefreshInterval } @@ -623,7 +623,7 @@ func (m *dashModel) beginAction(verb string) tea.Cmd { // beginKeep sets off a keep of the selected node for the confirmed duration. // It reuses beginAction's scaffolding — one action per node, the tile's // spinner, the short read interval for the duration of the call — but the call -// itself is a keep, not a start or stop: the node must be a Keeper (a remote +// itself is a keep, not a start or stop: the node must be a Keeper (a cloud // environment), and the call returns the deadline it set rather than an engine // state. A node that is not a Keeper, or that already has an action in flight, // is driven by nothing, and its reason lands on the status line. @@ -642,7 +642,7 @@ func (m *dashModel) beginKeep(d time.Duration) tea.Cmd { } keeper, ok := e.node.(fleet.Keeper) if !ok { - m.statusLine = e.name + ": keep needs a remote environment" + m.statusLine = e.name + ": keep needs a cloud environment" return nil } spin := !m.actionInFlight() @@ -783,7 +783,7 @@ func (m dashModel) nodeRunning() bool { } // keepOffered reports whether the keep key would do anything for the node under -// the cursor: the node must support a keep (a remote environment) and have +// the cursor: the node must support a keep (a cloud environment) and have // nothing in flight. A local daemon node has no retention tag to set, and a // node already acting takes no second action. The footer uses this to include // the keep hint only where it would drive something, the same way canAbort diff --git a/cmd/spinloop/dashboard_render.go b/cmd/spinloop/dashboard_render.go index 2cb13611..a28bd775 100644 --- a/cmd/spinloop/dashboard_render.go +++ b/cmd/spinloop/dashboard_render.go @@ -238,8 +238,8 @@ const dashStaleThreshold = 3 // dashStaleAfter is how old a reading of a node of this kind may be before the // panel says so. func dashStaleAfter(kind string) time.Duration { - if kind == fleet.KindRemote { - return dashStaleThreshold * dashboardRemoteRefreshInterval + if kind == fleet.KindCloud { + return dashStaleThreshold * dashboardCloudRefreshInterval } return dashStaleThreshold * dashboardRefreshInterval } @@ -332,7 +332,7 @@ func dashReadingAge(r fleet.NodeResult, now time.Time, staleAfter time.Duration) // weights; then an answer that carries no state at all is unknown; then a node // that answered with nothing serving is not serving, a faded dot rather than // the green of a node that is up and serving — idle, the daemon with nothing -// started; stopped, a daemon engine that was stopped; undeployed, a remote +// started; stopped, a daemon engine that was stopped; undeployed, a cloud // environment with no instance at all; anything else, including a running // engine the daemon reports no readiness for at all (an older daemon, or a // runner with no known health check), is healthy, so this degrades to the @@ -420,7 +420,7 @@ func dashTileReportBody(w io.Writer, m metrics.Stats, resources bool, gauge bool if line := dashTileServingLine(m); line != "" { fmt.Fprintln(w, line) } - // The active figure and, for a kept remote environment, the relative keep + // The active figure and, for a kept cloud environment, the relative keep // after it — one line, from the same read, whatever the engine's state. A // read without either — a local node, an unkept or lapsed environment — // draws nothing. diff --git a/cmd/spinloop/flagparse_test.go b/cmd/spinloop/flagparse_test.go index cd476b04..e0462d1b 100644 --- a/cmd/spinloop/flagparse_test.go +++ b/cmd/spinloop/flagparse_test.go @@ -29,7 +29,7 @@ func TestPflagParseForms(t *testing.T) { "apply": func() error { return cmdApply([]string{"--nope"}) }, "serve": func() error { return cmdServe([]string{"--nope"}) }, "fleet metrics": func() error { return cmdMetrics([]string{"--nope"}) }, - "remote start": func() error { return cmdRemoteStart([]string{"--env", "default", "--nope"}) }, + "cloud start": func() error { return cmdCloudStart([]string{"--env", "default", "--nope"}) }, "daemon": func() error { return cmdDaemon([]string{"--nope"}) }, } { if err := call(); err == nil || !strings.Contains(err.Error(), "unknown flag: --nope") { @@ -51,7 +51,7 @@ func TestPflagParseForms(t *testing.T) { // the proof that parsing continued past it. for name, call := range map[string]func() error{ "fleet metrics": func() error { return cmdMetrics([]string{"someNode", "--nope"}) }, - "remote env": func() error { return cmdRemoteEnv([]string{"--env", "default", "somePath", "--nope"}) }, + "cloud env": func() error { return cmdCloudEnv([]string{"--env", "default", "somePath", "--nope"}) }, } { if err := call(); err == nil || !strings.Contains(err.Error(), "unknown flag: --nope") { t.Errorf("%s --nope = %v, want unknown-flag error", name, err) diff --git a/cmd/spinloop/fleet.go b/cmd/spinloop/fleet.go index feca6765..cacdc322 100644 --- a/cmd/spinloop/fleet.go +++ b/cmd/spinloop/fleet.go @@ -60,7 +60,7 @@ func fleetRow(r fleet.NodeResult) (state, serving string) { if !r.OK() { return string(r.Outcome), r.Detail() } - // The shared facts come from the same source the remote status view reads, + // The shared facts come from the same source the cloud status view reads, // so the two cannot word or compute them differently. // A node that runs on an instance reports its release outside the status // reply, so the instance's answer fills what the reply left empty rather @@ -150,10 +150,10 @@ func renderFleetMetrics(w io.Writer, results []fleet.NodeResult, format string) } stats := r.Metrics renderMetricsHeader(w, r, format) - // Before the continue, for the same reason the remote formats show it + // Before the continue, for the same reason the cloud formats show it // before theirs: a node whose engine has stopped still has a useful // answer to "when did it last do anything?" — and, for a retained - // remote environment, "how long is it kept?". + // cloud environment, "how long is it kept?". // The table spells its facts as key-value lines, so its active line is // spelled that way too; the compact formats indent theirs under the // header. @@ -164,7 +164,7 @@ func renderFleetMetrics(w io.Writer, results []fleet.NodeResult, format string) } switch format { case "bar": - // No state gate, for the same reason the remote bar format has + // No state gate, for the same reason the cloud bar format has // none: a stopped node's retained history says what its engine // was doing until it stopped, and a stopped node's current // reading carries no figures for it to fall back on. @@ -301,7 +301,7 @@ func renderFleetMetricsJSON(w io.Writer, results []fleet.NodeResult) error { // deploy config that source derives (StartWith) — telling the daemon what to // run, exactly as a routed wake already does for the Spinloop being // launched. A kind: daemon node with no resolvable source, and a kind: -// remote node regardless, get a plain start: a remote environment's +// cloud node regardless, get a plain start: a cloud environment's // StartWith always refuses a config, since what it serves is fixed at // deploy time. func fleetStartCmd() *cobra.Command { @@ -451,10 +451,10 @@ func runFleetDrive(verb string, cfg *fleet.Config, all bool, names []string, cal return nil } -// fleetDeployCmd creates the AWS environment for one or more kind: remote -// nodes, or every kind: remote node with --all, deriving what each serves +// fleetDeployCmd creates the AWS environment for one or more kind: cloud +// nodes, or every kind: cloud node with --all, deriving what each serves // from its resolved Spinloop source — the same derivation and registration -// a standalone `spinloop remote deploy` performs for one file, so the two +// a standalone `spinloop cloud deploy` performs for one file, so the two // can never disagree about what a given Spinloop deploys. func fleetDeployCmd() *cobra.Command { var ( @@ -471,12 +471,12 @@ func fleetDeployCmd() *cobra.Command { ) c := &cobra.Command{ Use: "deploy", - Short: "create the AWS environment for one or more remote nodes", - Long: `deploys the AWS environment for each named kind: remote node, or -every kind: remote node with --all, deriving what to serve from each node's + Short: "create the AWS environment for one or more cloud nodes", + Long: `deploys the AWS environment for each named kind: cloud node, or +every kind: cloud node with --all, deriving what to serve from each node's own Spinloop source: its file field, or its name resolved as a registered alias or a same-named subdirectory beside the fleet file. Reuses the same -derivation, consent, and registration behaviour as "spinloop remote deploy".`, +derivation, consent, and registration behaviour as "spinloop cloud deploy".`, Args: cobra.ArbitraryArgs, SilenceErrors: true, SilenceUsage: true, @@ -496,7 +496,7 @@ derivation, consent, and registration behaviour as "spinloop remote deploy".`, fs := c.Flags() fs.StringVarP(&path, "fleet", "f", "", fleetFileUsage) fs.StringVar(&envName, "env", "", envFlagTargetUsage) - fs.BoolVar(&all, "all", false, "deploy every kind: remote node in the fleet") + fs.BoolVar(&all, "all", false, "deploy every kind: cloud node in the fleet") fs.BoolVarP(&dryRun, "dry-run", "n", false, "print the config that would be deployed, without sending it") fs.BoolVar(&overwrite, "overwrite", false, "proceed against an already-registered or live environment") fs.BoolVar(&reseed, "reseed", false, "re-fetch the weights even if they are already in S3 (starts a ~20-minute seed)") @@ -520,21 +520,21 @@ func runFleetDeploy(target fleetTarget, all bool, names []string, opts deployOpt return fmt.Errorf("spinloop fleet deploy: --all is ambiguous with node names") } - remoteNames := make([]string, 0, len(cfg.Nodes)) + cloudNames := make([]string, 0, len(cfg.Nodes)) for _, n := range cfg.Nodes { - if n.Kind == fleet.KindRemote { - remoteNames = append(remoteNames, n.Name) + if n.Kind == fleet.KindCloud { + cloudNames = append(cloudNames, n.Name) } } var targets []string switch { case all: - targets = remoteNames + targets = cloudNames case len(names) == 0: return fmt.Errorf( "spinloop fleet deploy needs a node, or --all: %s", - strings.Join(remoteNames, ", ")) + strings.Join(cloudNames, ", ")) default: for _, name := range names { entry, ok := cfg.Node(name) @@ -542,7 +542,7 @@ func runFleetDeploy(target fleetTarget, all bool, names []string, opts deployOpt return fmt.Errorf("no node %q in %s (known nodes: %s)", name, cfg.Path, strings.Join(cfg.Names(), ", ")) } - if entry.Kind != fleet.KindRemote { + if entry.Kind != fleet.KindCloud { return fmt.Errorf( "node %q is kind %q: fleet deploy provisions cloud environments, and %[1]s is not one", name, entry.Kind) @@ -705,7 +705,7 @@ func (r fleetDeployResult) text() string { } // deployOneNode resolves and deploys a single targeted node. The node's own -// name is the registered environment the deploy creates: a kind: remote node +// name is the registered environment the deploy creates: a kind: cloud node // is only driveable by the fleet commands under that name, so there is no // override. It never returns an error itself — a bad node becomes a // fleetDeployResult, so the caller's fan-out can label it without aborting diff --git a/cmd/spinloop/fleet_dashboard.go b/cmd/spinloop/fleet_dashboard.go index 51c59999..d428f694 100644 --- a/cmd/spinloop/fleet_dashboard.go +++ b/cmd/spinloop/fleet_dashboard.go @@ -65,7 +65,7 @@ func dashModelFor(target fleetTarget) (dashModel, error) { // the terminal on the way out, whatever key got here. The program holds the // model by pointer: a value model would drop the mutations Init makes (the // program does not read Init's receiver back), and the first round's answers -// — expensive cloud calls, for remote environments — must not be discarded. +// — expensive cloud calls, for cloud environments — must not be discarded. func runDashProgram(m dashModel) error { prog := tea.NewProgram(&m, tea.WithAltScreen()) // The calls behind in-flight actions report their status lines from diff --git a/cmd/spinloop/fleet_dashboard_test.go b/cmd/spinloop/fleet_dashboard_test.go index 6f698951..5bf683af 100644 --- a/cmd/spinloop/fleet_dashboard_test.go +++ b/cmd/spinloop/fleet_dashboard_test.go @@ -196,7 +196,7 @@ func (f *fakeDashNode) Logs(ctx context.Context, offset int64, limit int) (daemo // door and runs it to completion. func startFastRound(t *testing.T, m *dashModel) (dashRefreshMsg, bool) { t.Helper() - cmd := m.refreshRemoteGroup(false) + cmd := m.refreshCloudGroup(false) if cmd == nil { return dashRefreshMsg{}, false } @@ -441,7 +441,7 @@ func TestDashTileStoppedByteStable(t *testing.T) { if got := dashTestTile("idle", r, false, dashAction{}); got != want { t.Errorf("stopped tile mismatch:\n%q\nwant:\n%q", got, want) } - // A remote environment with no instance at all reports undeployed and + // A cloud environment with no instance at all reports undeployed and // keeps its deployment: the serving line rides on the same shape, and // the dot is the faded one, not the green of a serving node. u := fleet.NodeResult{ @@ -698,7 +698,7 @@ func TestDashTileStaleReadingShowsItsAgeAndRecovers(t *testing.T) { r := fleet.NodeResult{Name: "dev-1", Outcome: fleet.OutcomeOK, Metrics: metrics.Stats{State: "running", Ready: "ready"}, At: now.Add(-4 * time.Minute)} - staleAfter := dashStaleAfter(fleet.KindRemote) // three minutes, on the minute cadence + staleAfter := dashStaleAfter(fleet.KindCloud) // three minutes, on the minute cadence got := dashTile("dev-1", r, false, dashAction{}, now, staleAfter, false) want := dashTileExpected([]string{ dashExpectedHeader("dev-1 running · 4m 0s ago", dashUnknown), @@ -1411,13 +1411,13 @@ func TestDashModelRefreshOrdersReadingsByTheirTime(t *testing.T) { } } -// A kind: remote environment refreshes on its own slower deadline: the round +// A kind: cloud environment refreshes on its own slower deadline: the round // starts only when that time comes, starting it spends the deadline, and the // manual refresh brings it forward whatever the deadline says. -func TestDashModelRemoteCadence(t *testing.T) { - orig := dashboardRemoteRefreshInterval - dashboardRemoteRefreshInterval = time.Minute - defer func() { dashboardRemoteRefreshInterval = orig }() +func TestDashModelCloudCadence(t *testing.T) { + orig := dashboardCloudRefreshInterval + dashboardCloudRefreshInterval = time.Minute + defer func() { dashboardCloudRefreshInterval = orig }() local := newFakeDashNode("running") r1 := newFakeDashNode("running") @@ -1425,8 +1425,8 @@ func TestDashModelRemoteCadence(t *testing.T) { m := &dashModel{ entries: []dashEntry{ {name: "local", kind: fleet.KindDaemon, node: local}, - {name: "r1", kind: fleet.KindRemote, node: r1}, - {name: "r2", kind: fleet.KindRemote, node: r2}, + {name: "r1", kind: fleet.KindCloud, node: r1}, + {name: "r2", kind: fleet.KindCloud, node: r2}, }, results: make([]fleet.NodeResult, 3), actions: make([]dashAction, 3), @@ -1479,17 +1479,17 @@ func TestDashModelRemoteCadence(t *testing.T) { // kind, and returns to its own cadence once the action settles. Its // neighbours in the same group keep their own cadence throughout. func TestDashModelActedOnNodeIsReadMoreOften(t *testing.T) { - orig := dashboardRemoteRefreshInterval - dashboardRemoteRefreshInterval = time.Minute - defer func() { dashboardRemoteRefreshInterval = orig }() + orig := dashboardCloudRefreshInterval + dashboardCloudRefreshInterval = time.Minute + defer func() { dashboardCloudRefreshInterval = orig }() r1, r2 := newFakeDashNode("stopped"), newFakeDashNode("running") hold := make(chan struct{}) r1.hold = hold m := &dashModel{ entries: []dashEntry{ - {name: "r1", kind: fleet.KindRemote, node: r1}, - {name: "r2", kind: fleet.KindRemote, node: r2}, + {name: "r1", kind: fleet.KindCloud, node: r1}, + {name: "r2", kind: fleet.KindCloud, node: r2}, }, results: make([]fleet.NodeResult, 2), actions: make([]dashAction, 2), @@ -1656,8 +1656,8 @@ func TestDashModelConcurrentStarts(t *testing.T) { b := newFakeDashNode("stopped") m := &dashModel{ entries: []dashEntry{ - {name: "a", kind: fleet.KindRemote, node: a}, - {name: "b", kind: fleet.KindRemote, node: b}, + {name: "a", kind: fleet.KindCloud, node: a}, + {name: "b", kind: fleet.KindCloud, node: b}, }, results: make([]fleet.NodeResult, 2), actions: make([]dashAction, 2), @@ -1734,7 +1734,7 @@ func TestDashModelLandedRoundShowsBesideInFlightAction(t *testing.T) { dashFixNow(t, dashTestClock) f := newFakeDashNode("stopped") m := &dashModel{ - entries: []dashEntry{{name: "a", kind: fleet.KindRemote, node: f}}, + entries: []dashEntry{{name: "a", kind: fleet.KindCloud, node: f}}, results: make([]fleet.NodeResult, 1), actions: make([]dashAction, 1), width: 120, height: 40, @@ -1750,7 +1750,7 @@ func TestDashModelLandedRoundShowsBesideInFlightAction(t *testing.T) { Since: dashTestClock, RetryAt: dashTestClock.Add(time.Second)}}) m = next.(*dashModel) // The cloud round lands while the start is in flight. - cmd := m.refreshRemoteGroup(true) + cmd := m.refreshCloudGroup(true) if cmd == nil { t.Fatal("the cloud round did not start") } @@ -1860,7 +1860,7 @@ func TestDashModelQuitDuringConfirmation(t *testing.T) { func TestDashModelLateRoundDoesNotOverwriteAPostActionReport(t *testing.T) { node := newFakeDashNode("stopped") m := &dashModel{ - entries: []dashEntry{{name: "a", kind: fleet.KindRemote, node: node}}, + entries: []dashEntry{{name: "a", kind: fleet.KindCloud, node: node}}, results: make([]fleet.NodeResult, 1), actions: make([]dashAction, 1), width: 120, height: 40, @@ -1869,13 +1869,13 @@ func TestDashModelLateRoundDoesNotOverwriteAPostActionReport(t *testing.T) { // The start finishes, and the report that follows it lands. after := fleet.NodeResult{Name: "a", Outcome: fleet.OutcomeOK, Metrics: metrics.Stats{State: "running"}, At: issued.Add(3 * time.Second)} - next, _ := m.Update(dashRefreshMsg{remote: true, idx: []int{0}, results: []fleet.NodeResult{after}}) + next, _ := m.Update(dashRefreshMsg{cloud: true, idx: []int{0}, results: []fleet.NodeResult{after}}) m = next.(*dashModel) // The round issued before the action finished now answers, carrying the // node as it was then. before := fleet.NodeResult{Name: "a", Outcome: fleet.OutcomeOK, Metrics: metrics.Stats{State: "stopped"}, At: issued} - next, _ = m.Update(dashRefreshMsg{remote: true, idx: []int{0}, results: []fleet.NodeResult{before}}) + next, _ = m.Update(dashRefreshMsg{cloud: true, idx: []int{0}, results: []fleet.NodeResult{before}}) m = next.(*dashModel) if got := m.results[0].Metrics.State; got != "running" { t.Errorf("the late round repainted the node's older report: state = %q, want running", got) @@ -1887,14 +1887,14 @@ func TestDashModelLateRoundDoesNotOverwriteAPostActionReport(t *testing.T) { func TestDashModelSlowRoundInFlightGuard(t *testing.T) { node := newFakeDashNode("running") m := &dashModel{ - entries: []dashEntry{{name: "a", kind: fleet.KindRemote, node: node}}, + entries: []dashEntry{{name: "a", kind: fleet.KindCloud, node: node}}, results: make([]fleet.NodeResult, 1), actions: make([]dashAction, 1), width: 120, height: 40, } m.slowBusy = true m.scheduleRead(0, time.Time{}) // the node is due - if cmd := m.refreshRemoteGroup(true); cmd != nil { + if cmd := m.refreshCloudGroup(true); cmd != nil { t.Fatal("started a second slow round over one in flight") } } @@ -2005,7 +2005,7 @@ func TestDashModelAbortsAnInFlightStart(t *testing.T) { f := newFakeDashNode("stopped") f.hold = hold // the start stays in flight until released or cancelled m := &dashModel{ - entries: []dashEntry{{name: "a", kind: fleet.KindRemote, node: f}}, + entries: []dashEntry{{name: "a", kind: fleet.KindCloud, node: f}}, results: []fleet.NodeResult{{Name: "a"}}, actions: make([]dashAction, 1), width: 120, height: 40, @@ -2095,7 +2095,7 @@ func TestDashModelAbortOnAnIdleNodeDrivesNothing(t *testing.T) { func TestDashModelRacingSuccessIsReportedAsSuccess(t *testing.T) { node := newFakeDashNode("stopped") m := &dashModel{ - entries: []dashEntry{{name: "a", kind: fleet.KindRemote, node: node}}, + entries: []dashEntry{{name: "a", kind: fleet.KindCloud, node: node}}, results: []fleet.NodeResult{{Name: "a"}}, actions: make([]dashAction, 1), width: 120, height: 40, @@ -2721,9 +2721,9 @@ func TestDashFooterNamesOnlyTheKeysTheNodeTakes(t *testing.T) { "esc back f follow", }, { - "a stopped remote environment", + "a stopped cloud environment", dashModel{ - entries: []dashEntry{{name: "env", kind: fleet.KindRemote, node: kept}}, + entries: []dashEntry{{name: "env", kind: fleet.KindCloud, node: kept}}, results: []fleet.NodeResult{read("stopped")}, actions: make([]dashAction, 1), }, @@ -2731,9 +2731,9 @@ func TestDashFooterNamesOnlyTheKeysTheNodeTakes(t *testing.T) { "esc back s start k keep f follow", }, { - "a running remote environment", + "a running cloud environment", dashModel{ - entries: []dashEntry{{name: "env", kind: fleet.KindRemote, node: kept}}, + entries: []dashEntry{{name: "env", kind: fleet.KindCloud, node: kept}}, results: []fleet.NodeResult{read("running")}, actions: make([]dashAction, 1), }, diff --git a/cmd/spinloop/fleet_deploy_test.go b/cmd/spinloop/fleet_deploy_test.go index 809fe499..d5f594c0 100644 --- a/cmd/spinloop/fleet_deploy_test.go +++ b/cmd/spinloop/fleet_deploy_test.go @@ -11,8 +11,8 @@ import ( "testing" "github.com/aws/aws-sdk-go-v2/aws" + "github.com/spinloop-ai/spinloop/internal/cloud" "github.com/spinloop-ai/spinloop/internal/config" - "github.com/spinloop-ai/spinloop/internal/remote" ) // fleetDeployServer answers every environment's deploy call with a @@ -34,15 +34,15 @@ func fleetDeployServer(t *testing.T) *httptest.Server { // --overwrite) unless overridden by the caller after this returns. func stubFleetDeploySeams(t *testing.T, server *httptest.Server) { t.Helper() - origDiscover, origStatus, origDetect := deployDiscoverFn, remoteStatusFn, detectPublicCIDRFn - t.Cleanup(func() { deployDiscoverFn, remoteStatusFn, detectPublicCIDRFn = origDiscover, origStatus, origDetect }) - deployDiscoverFn = func(context.Context, aws.Config, string) (remote.ControlPlane, error) { - return remote.ControlPlane{Config: remote.Config{ + origDiscover, origStatus, origDetect := deployDiscoverFn, cloudStatusFn, detectPublicCIDRFn + t.Cleanup(func() { deployDiscoverFn, cloudStatusFn, detectPublicCIDRFn = origDiscover, origStatus, origDetect }) + deployDiscoverFn = func(context.Context, aws.Config, string) (cloud.ControlPlane, error) { + return cloud.ControlPlane{Config: cloud.Config{ StartURL: server.URL, StopURL: server.URL, DeployURL: server.URL, Region: "us-east-1", }}, nil } - remoteStatusFn = func(context.Context, remote.Config) (*remote.Response, error) { - return &remote.Response{StatusCode: 200, State: "undeployed"}, nil + cloudStatusFn = func(context.Context, cloud.Config) (*cloud.Response, error) { + return &cloud.Response{StatusCode: 200, State: "undeployed"}, nil } detectPublicCIDRFn = func(context.Context) (string, error) { return "203.0.113.7/32", nil } } @@ -58,17 +58,17 @@ func writeFleetDeploySetup(t *testing.T) string { dir := writeFleetFile(t, ` nodes: - name: gpu-a - kind: remote + kind: cloud file: ./gpu-a.Spinloop - name: gpu-b - kind: remote + kind: cloud file: ./gpu-b.Spinloop - name: aliased - kind: remote + kind: cloud - name: subdir-env - kind: remote + kind: cloud - name: no-source - kind: remote + kind: cloud - name: studio host: studio.local `) @@ -154,7 +154,7 @@ func TestCmdFleetDeployNoTargetIsAnError(t *testing.T) { } for _, want := range []string{"gpu-a", "gpu-b", "aliased"} { if !strings.Contains(err.Error(), want) { - t.Errorf("error %q should list the remote nodes, missing %q", err, want) + t.Errorf("error %q should list the cloud nodes, missing %q", err, want) } } // studio (kind: daemon) must not be offered as a deploy target. @@ -229,8 +229,8 @@ func TestCmdFleetDeployResolvedButUndeployableSpinloopFailsOnlyThatNode(t *testi // gpu-a's own file, with a REMOTE line — readSpinloop rejects it, but // resolveNodeSpinloop has already succeeded by the time it does. - staleRemote := filepath.Join(dir, "gpu-a.Spinloop") - if err := os.WriteFile(staleRemote, []byte("PROVIDER llamacpp\nMODEL org/m:Q4\nCONTEXT 8192\nREMOTE gpu-a\n"), 0o600); err != nil { + staleCloud := filepath.Join(dir, "gpu-a.Spinloop") + if err := os.WriteFile(staleCloud, []byte("PROVIDER llamacpp\nMODEL org/m:Q4\nCONTEXT 8192\nREMOTE gpu-a\n"), 0o600); err != nil { t.Fatal(err) } @@ -291,7 +291,7 @@ func TestCmdFleetDeployGuardDoesNotBlockSiblings(t *testing.T) { server := fleetDeployServer(t) stubFleetDeploySeams(t, server) - if err := remote.SaveEnvironment("gpu-a", remote.Config{StartURL: "https://s", StopURL: "https://x", Region: "us-east-1"}); err != nil { + if err := cloud.SaveEnvironment("gpu-a", cloud.Config{StartURL: "https://s", StopURL: "https://x", Region: "us-east-1"}); err != nil { t.Fatal(err) } @@ -313,9 +313,9 @@ func TestCmdFleetDeployDryRunTouchesNothing(t *testing.T) { called := false origDiscover := deployDiscoverFn t.Cleanup(func() { deployDiscoverFn = origDiscover }) - deployDiscoverFn = func(context.Context, aws.Config, string) (remote.ControlPlane, error) { + deployDiscoverFn = func(context.Context, aws.Config, string) (cloud.ControlPlane, error) { called = true - return remote.ControlPlane{}, fmt.Errorf("must not be called") + return cloud.ControlPlane{}, fmt.Errorf("must not be called") } out := captureStdout(t, func() { @@ -342,11 +342,11 @@ func TestCmdFleetDeployInstanceTypePerNode(t *testing.T) { dir := writeFleetFile(t, ` nodes: - name: typed - kind: remote + kind: cloud file: ./typed.Spinloop instance-type: g6e.2xlarge - name: untyped - kind: remote + kind: cloud file: ./untyped.Spinloop `) writeSpinloop := func(name string) { @@ -359,10 +359,10 @@ nodes: writeSpinloop("typed.Spinloop") writeSpinloop("untyped.Spinloop") - deployDiscoverFn = func(context.Context, aws.Config, string) (remote.ControlPlane, error) { - return remote.ControlPlane{}, fmt.Errorf("must not be called") + deployDiscoverFn = func(context.Context, aws.Config, string) (cloud.ControlPlane, error) { + return cloud.ControlPlane{}, fmt.Errorf("must not be called") } - t.Cleanup(func() { deployDiscoverFn = remote.DiscoverControlPlane }) + t.Cleanup(func() { deployDiscoverFn = cloud.DiscoverControlPlane }) // --dry-run touches nothing, so no deploy seams need stubbing. outTyped := captureStdout(t, func() { @@ -388,7 +388,7 @@ nodes: // registered, under the isolated config dir this test's HOME points at. func mustEnvConfigPath(t *testing.T, env string) string { t.Helper() - path, err := remote.EnvConfigPath(env) + path, err := cloud.EnvConfigPath(env) if err != nil { t.Fatal(err) } diff --git a/cmd/spinloop/fleet_test.go b/cmd/spinloop/fleet_test.go index 007eb693..14a16303 100644 --- a/cmd/spinloop/fleet_test.go +++ b/cmd/spinloop/fleet_test.go @@ -499,12 +499,12 @@ func (f *fakeFleetNode) Logs(context.Context, int64, int) (daemon.LogsResponse, return daemon.LogsResponse{}, nil } -// A kind: remote node's start always uses a plain start, never StartWith — +// A kind: cloud node's start always uses a plain start, never StartWith — // StartWith refuses a config for that kind unconditionally (see -// remoteNode.StartWith), so fleetStartCall must not even attempt it. -func TestFleetStartCallRemoteNodeUsesPlainStart(t *testing.T) { +// cloudNode.StartWith), so fleetStartCall must not even attempt it. +func TestFleetStartCallCloudNodeUsesPlainStart(t *testing.T) { t.Setenv("SPINLOOP_CONFIG_DIR", t.TempDir()) - dir := writeFleetFile(t, "nodes:\n - name: gpu-env\n kind: remote\n file: ./gpu-env.Spinloop\n") + dir := writeFleetFile(t, "nodes:\n - name: gpu-env\n kind: cloud\n file: ./gpu-env.Spinloop\n") if err := os.WriteFile(filepath.Join(dir, "gpu-env.Spinloop"), []byte("PROVIDER llamacpp\nMODEL org/m:Q4\n"), 0o600); err != nil { t.Fatal(err) } @@ -515,10 +515,10 @@ func TestFleetStartCallRemoteNodeUsesPlainStart(t *testing.T) { node := &fakeFleetNode{name: "gpu-env"} r := fleetStartCall(cfg)(context.Background(), node) if !r.OK() { - t.Fatalf("fleetStartCall on a remote node = %+v", r) + t.Fatalf("fleetStartCall on a cloud node = %+v", r) } if node.startWithCalls != 0 { - t.Errorf("StartWith was called %d times for a remote node, want 0", node.startWithCalls) + t.Errorf("StartWith was called %d times for a cloud node, want 0", node.startWithCalls) } if node.startCalls != 1 { t.Errorf("Start was called %d times, want 1", node.startCalls) @@ -558,7 +558,7 @@ func TestFleetStartCallDaemonNodeUsesStartWith(t *testing.T) { func TestFleetStartCallResolvedButBrokenSourceNeverStarts(t *testing.T) { cases := map[string]string{ "unparseable Spinloop": "this is not a Spinloop\x00\x01", - // deployConfigForNode shares remote deploy's runnerFor, which only + // deployConfigForNode shares cloud deploy's runnerFor, which only // accepts llamacpp/vllm — the same limit that already applies to a // routed wake. An MLX (or any other) provider is resolved but // cannot be turned into a deploy config. diff --git a/cmd/spinloop/follow.go b/cmd/spinloop/follow.go index b7fdd6d7..8309245d 100644 --- a/cmd/spinloop/follow.go +++ b/cmd/spinloop/follow.go @@ -9,7 +9,7 @@ import ( // followUntilInterrupted runs a polling loop under a context that an interrupt // cancels, and treats that cancellation as success — the user asking a follow -// to stop is not a failure. Both `spinloop remote logs -f` and `spinloop fleet +// to stop is not a failure. Both `spinloop cloud logs -f` and `spinloop fleet // logs -f` follow output this way, and the wiring is the part they genuinely // share: what they poll, and what they do with the answers, is different // enough that a common loop would fit neither. diff --git a/cmd/spinloop/follow_test.go b/cmd/spinloop/follow_test.go index 4dd4efaa..7ed4da83 100644 --- a/cmd/spinloop/follow_test.go +++ b/cmd/spinloop/follow_test.go @@ -6,7 +6,7 @@ import ( "testing" ) -// followUntilInterrupted is the wiring both `remote logs -f` and `fleet logs +// followUntilInterrupted is the wiring both `cloud logs -f` and `fleet logs // -f` reach, and nothing exercised it: each command's tests drive the polling // loop directly, so the wrapper that installs the signal handler and decides // what a cancelled follow returns was never entered. diff --git a/cmd/spinloop/gateway.go b/cmd/spinloop/gateway.go index 338a101a..0d500c84 100644 --- a/cmd/spinloop/gateway.go +++ b/cmd/spinloop/gateway.go @@ -154,8 +154,8 @@ func newGatewayServer(fleetPath, listen, apiToken, apiTokenFile string, maxReque if _, err := cfg.Token(entry); err != nil { return nil, nil, err } - if entry.Kind == fleet.KindRemote { - if _, err := cfg.RemoteEngineToken(entry); err != nil { + if entry.Kind == fleet.KindCloud { + if _, err := cfg.CloudEngineToken(entry); err != nil { return nil, nil, err } } else if _, err := cfg.EngineToken(entry); err != nil { diff --git a/cmd/spinloop/harness_remote_test.go b/cmd/spinloop/harness_cloud_test.go similarity index 88% rename from cmd/spinloop/harness_remote_test.go rename to cmd/spinloop/harness_cloud_test.go index 1209f144..5b10d50a 100644 --- a/cmd/spinloop/harness_remote_test.go +++ b/cmd/spinloop/harness_cloud_test.go @@ -8,19 +8,19 @@ import ( "strings" "testing" + "github.com/spinloop-ai/spinloop/internal/cloud" "github.com/spinloop-ai/spinloop/internal/fleet" "github.com/spinloop-ai/spinloop/internal/harness" - "github.com/spinloop-ai/spinloop/internal/remote" "github.com/spinloop-ai/spinloop/internal/spinloop" ) -// remoteSpinloopDir registers an environment served by envURL and writes a +// cloudSpinloopDir registers an environment served by envURL and writes a // Spinloop in a fresh directory, returning the directory. baseURL is the // endpoint address the environment records — a public one, so the apply has a // key to warn about when it has none. -func remoteSpinloopDir(t *testing.T, name, envURL, baseURL string) string { +func cloudSpinloopDir(t *testing.T, name, envURL, baseURL string) string { t.Helper() - registerEnv(t, name, remote.Config{ + registerEnv(t, name, cloud.Config{ StartURL: envURL, StopURL: envURL, EnvURL: envURL, @@ -62,7 +62,7 @@ func deployedEnvServer(t *testing.T) *httptest.Server { })) } -// The key of a remote endpoint is known only to the control plane, and +// The key of a cloud endpoint is known only to the control plane, and // `spinloop harness` fetches it and hands it to the agent it launches. The apply // that precedes the launch must therefore not warn that no key is set: it is // about to be. @@ -73,7 +73,7 @@ func TestApplyBeforeLaunch_RemoteKeySilencesTheMissingKeyWarning(t *testing.T) { server := envServer(t) defer server.Close() - dir := remoteSpinloopDir(t, "dev-1", server.URL, "http://198.51.100.1:8000/v1") + dir := cloudSpinloopDir(t, "dev-1", server.URL, "http://198.51.100.1:8000/v1") h, _ := harness.Lookup("opencode") var sel spinloopSelectionResult @@ -95,7 +95,7 @@ func TestApplyBeforeLaunch_RemoteKeySilencesTheMissingKeyWarning(t *testing.T) { } // ...and it reaches the launched agent's environment. - env := harnessEnv("", remoteLaunchResolver(func(string) string { return "" }, sel.resp), sel.resp) + env := harnessEnv("", cloudLaunchResolver(func(string) string { return "" }, sel.resp), sel.resp) if got, _ := envValue(env, "OPENAI_API_KEY"); got != "sk-remote" { t.Errorf("launched agent's OPENAI_API_KEY = %q, want sk-remote", got) } @@ -117,7 +117,7 @@ func TestHarness_EnvFlagAfterTheAlias(t *testing.T) { server := envServer(t) defer server.Close() - dir := remoteSpinloopDir(t, "dev-1", server.URL, "http://198.51.100.1:8000/v1") + dir := cloudSpinloopDir(t, "dev-1", server.URL, "http://198.51.100.1:8000/v1") captureStdout(t, func() { if err := cmdAlias([]string{"-n", "dev-3", dir}); err != nil { t.Fatalf("cmdAlias: %v", err) @@ -154,10 +154,10 @@ func TestHarness_EnvFlagAfterTheAlias(t *testing.T) { t.Fatalf("could not read the launched agent's env: %v", err) } if v, _ := envValue(strings.Split(string(env), "\n"), "OPENAI_API_KEY"); v != "sk-remote" { - t.Errorf("launched agent's OPENAI_API_KEY = %q, want the fetched remote key", v) + t.Errorf("launched agent's OPENAI_API_KEY = %q, want the fetched cloud key", v) } if v, _ := envValue(strings.Split(string(env), "\n"), "OPENAI_BASE_URL"); v != "http://198.51.100.1:8000/v1" { - t.Errorf("launched agent's OPENAI_BASE_URL = %q, want the remote endpoint's address", v) + t.Errorf("launched agent's OPENAI_BASE_URL = %q, want the cloud endpoint's address", v) } } @@ -172,7 +172,7 @@ func TestHarness_BareEnvAutoConfiguresAndLaunches(t *testing.T) { server := deployedEnvServer(t) defer server.Close() - registerEnv(t, "dev-3", remote.Config{ + registerEnv(t, "dev-3", cloud.Config{ StartURL: server.URL, StopURL: server.URL, EnvURL: server.URL, BaseURL: "http://198.51.100.1:8000/v1", Region: "eu-west-1", Environment: "dev-3", }) @@ -202,7 +202,7 @@ func TestHarness_AppliedSpinloopStillWinsOverDeployConfig(t *testing.T) { server := deployedEnvServer(t) // deploy-config's servedName is "q3" defer server.Close() - dir := remoteSpinloopDir(t, "dev-3", server.URL, "http://198.51.100.1:8000/v1") + dir := cloudSpinloopDir(t, "dev-3", server.URL, "http://198.51.100.1:8000/v1") mustWrite(t, filepath.Join(dir, "Spinloop"), "PROVIDER llamacpp\nALIAS from-spinloop\n") forwarded, _ := launchedArgs(t, []string{dir, "--env", "dev-3", "--", "run"}) @@ -222,7 +222,7 @@ func TestHarness_AppliedSpinloopStillWinsOverDeployConfig(t *testing.T) { type spinloopSelectionResult struct { selection spinloop.Selection envDir string - resp *remote.Response + resp *cloud.Response err error } @@ -246,7 +246,7 @@ func TestApplyBeforeLaunch_FailsWhenNoKeyCanBeHad(t *testing.T) { server := failingEnvServer(t) defer server.Close() - dir := remoteSpinloopDir(t, "dev-1", server.URL, "http://198.51.100.1:8000/v1") + dir := cloudSpinloopDir(t, "dev-1", server.URL, "http://198.51.100.1:8000/v1") h, _ := harness.Lookup("opencode") var res spinloopSelectionResult @@ -259,7 +259,7 @@ func TestApplyBeforeLaunch_FailsWhenNoKeyCanBeHad(t *testing.T) { if res.err == nil { t.Fatal("a launch that cannot authenticate should fail") } - for _, want := range []string{"could not fetch the API key for dev-1", "spinloop remote start --env dev-1", "OPENAI_API_KEY"} { + for _, want := range []string{"could not fetch the API key for dev-1", "spinloop cloud start --env dev-1", "OPENAI_API_KEY"} { if !strings.Contains(res.err.Error(), want) { t.Errorf("error should mention %q, got: %v", want, res.err) } @@ -279,7 +279,7 @@ func TestApplyBeforeLaunch_CarriesOnWhenTheKeyIsAlreadySet(t *testing.T) { server := failingEnvServer(t) defer server.Close() - dir := remoteSpinloopDir(t, "dev-1", server.URL, "http://198.51.100.1:8000/v1") + dir := cloudSpinloopDir(t, "dev-1", server.URL, "http://198.51.100.1:8000/v1") h, _ := harness.Lookup("opencode") var res spinloopSelectionResult @@ -313,7 +313,7 @@ func TestApplyBeforeLaunch_CountsAnEnvInstructionAsTheKey(t *testing.T) { server := failingEnvServer(t) defer server.Close() - dir := remoteSpinloopDir(t, "dev-1", server.URL, "http://198.51.100.1:8000/v1") + dir := cloudSpinloopDir(t, "dev-1", server.URL, "http://198.51.100.1:8000/v1") mustWrite(t, filepath.Join(dir, "Spinloop"), "PROVIDER llamacpp\nALIAS q3\nENV OPENAI_API_KEY=sk-from-spinloop\n") @@ -338,7 +338,7 @@ func TestApplyBeforeLaunch_AnnouncesTheFetch(t *testing.T) { server := envServer(t) defer server.Close() - dir := remoteSpinloopDir(t, "dev-1", server.URL, "http://198.51.100.1:8000/v1") + dir := cloudSpinloopDir(t, "dev-1", server.URL, "http://198.51.100.1:8000/v1") h, _ := harness.Lookup("opencode") var stdout string @@ -357,7 +357,7 @@ func TestApplyBeforeLaunch_AnnouncesTheFetch(t *testing.T) { } } -// A Spinloop that names no remote contacts nothing and reports nothing. +// A Spinloop that names no cloud contacts nothing and reports nothing. func TestApplyBeforeLaunch_LocalSpinloopFetchesNothing(t *testing.T) { isolateConfig(t) @@ -376,7 +376,7 @@ func TestApplyBeforeLaunch_LocalSpinloopFetchesNothing(t *testing.T) { t.Fatalf("applyBeforeLaunch: %v", res.err) } if res.resp != nil { - t.Errorf("a local Spinloop should fetch no remote environment, got %+v", res.resp) + t.Errorf("a local Spinloop should fetch no cloud environment, got %+v", res.resp) } if strings.Contains(stderr, "could not fetch") { t.Errorf("nothing was fetched, so nothing should be reported:\n%s", stderr) @@ -391,8 +391,8 @@ func TestApplyBeforeLaunch_EnvAndFleetConflict(t *testing.T) { server := envServer(t) defer server.Close() - dir := remoteSpinloopDir(t, "dev-1", server.URL, "http://198.51.100.1:8000/v1") - fleetDir := writeFleetFile(t, "nodes:\n - name: gpu-a\n kind: remote\n") + dir := cloudSpinloopDir(t, "dev-1", server.URL, "http://198.51.100.1:8000/v1") + fleetDir := writeFleetFile(t, "nodes:\n - name: gpu-a\n kind: cloud\n") h, _ := harness.Lookup("opencode") _, _, _, _, err := applyBeforeLaunch(spinloopPathFlag{set: true, path: dir}, "", h, nil, @@ -416,14 +416,14 @@ func TestApplyFromEnvironment_ConfiguresFromDeployConfig(t *testing.T) { server := deployedEnvServer(t) defer server.Close() - registerEnv(t, "dev-3", remote.Config{ + registerEnv(t, "dev-3", cloud.Config{ StartURL: server.URL, StopURL: server.URL, EnvURL: server.URL, BaseURL: "http://198.51.100.1:8000/v1", Region: "eu-west-1", Environment: "dev-3", }) h, _ := harness.Lookup("opencode") var sel spinloop.Selection - var resp *remote.Response + var resp *cloud.Response var choice *fleet.Choice var err error captureStderr(t, func() { @@ -465,7 +465,7 @@ func TestApplyFromEnvironment_UnrecognisedRunnerFails(t *testing.T) { w.Write([]byte(`{"base_url":"http://198.51.100.1:8000/v1","api_key":"sk-remote","deployed":true,"runner":"bogus","servedName":"q3"}`)) })) defer server.Close() - registerEnv(t, "dev-3", remote.Config{ + registerEnv(t, "dev-3", cloud.Config{ StartURL: server.URL, StopURL: server.URL, EnvURL: server.URL, Region: "eu-west-1", Environment: "dev-3", }) @@ -488,7 +488,7 @@ func TestApplyFromEnvironment_FailsWithNothingDeployed(t *testing.T) { server := envServer(t) // base_url/api_key only — no deploy-config fields defer server.Close() - registerEnv(t, "dev-3", remote.Config{ + registerEnv(t, "dev-3", cloud.Config{ StartURL: server.URL, StopURL: server.URL, EnvURL: server.URL, Region: "eu-west-1", Environment: "dev-3", }) @@ -502,7 +502,7 @@ func TestApplyFromEnvironment_FailsWithNothingDeployed(t *testing.T) { if err == nil { t.Fatal("a launch with nothing to auto-configure from should fail") } - for _, want := range []string{"dev-3", "remote deploy", "remote bootstrap"} { + for _, want := range []string{"dev-3", "cloud deploy", "cloud bootstrap"} { if !strings.Contains(err.Error(), want) { t.Errorf("error should mention %q, got: %v", want, err) } @@ -512,16 +512,16 @@ func TestApplyFromEnvironment_FailsWithNothingDeployed(t *testing.T) { } } -// remoteLaunchResolver widens a lookup rather than replacing it: an exported +// cloudLaunchResolver widens a lookup rather than replacing it: an exported // key or one from the .env is the user's own and still wins. -func TestRemoteLaunchResolver_KeepsTheLocalValue(t *testing.T) { +func TestCloudLaunchResolver_KeepsTheLocalValue(t *testing.T) { base := func(name string) string { if name == "OPENAI_API_KEY" { return "sk-local" } return "" } - resolve := remoteLaunchResolver(base, &remote.Response{APIKey: "sk-remote"}) + resolve := cloudLaunchResolver(base, &cloud.Response{APIKey: "sk-remote"}) if got := resolve("OPENAI_API_KEY"); got != "sk-local" { t.Errorf("resolved %q, want the local value to win", got) } @@ -529,20 +529,20 @@ func TestRemoteLaunchResolver_KeepsTheLocalValue(t *testing.T) { t.Errorf("resolved an unrelated variable as %q", got) } // With nothing fetched the lookup is the base one, unchanged. - if got := remoteLaunchResolver(base, nil)("OPENAI_API_KEY"); got != "sk-local" { - t.Errorf("resolved %q with no remote response, want sk-local", got) + if got := cloudLaunchResolver(base, nil)("OPENAI_API_KEY"); got != "sk-local" { + t.Errorf("resolved %q with no cloud response, want sk-local", got) } } -// `spinloop remote env` is meant to be eval'd, so nothing but export lines may +// `spinloop cloud env` is meant to be eval'd, so nothing but export lines may // reach stdout — an alias, which is reported, is the case that broke it. -func TestRemoteEnv_StdoutIsEvalSafe(t *testing.T) { +func TestCloudEnv_StdoutIsEvalSafe(t *testing.T) { isolateConfig(t) stubAWSEnv(t) server := envServer(t) defer server.Close() - dir := remoteSpinloopDir(t, "dev-1", server.URL, "http://198.51.100.1:8000/v1") + dir := cloudSpinloopDir(t, "dev-1", server.URL, "http://198.51.100.1:8000/v1") captureStdout(t, func() { if err := cmdAlias([]string{dir}); err != nil { t.Fatalf("cmdAlias: %v", err) @@ -553,14 +553,14 @@ func TestRemoteEnv_StdoutIsEvalSafe(t *testing.T) { var stdout string captureStderr(t, func() { stdout = captureStdout(t, func() { - if err := cmdRemoteEnv([]string{"q3", "--env", "dev-1"}); err != nil { - t.Fatalf("cmdRemoteEnv: %v", err) + if err := cmdCloudEnv([]string{"q3", "--env", "dev-1"}); err != nil { + t.Fatalf("cmdCloudEnv: %v", err) } }) }) for _, line := range strings.Split(strings.TrimSpace(stdout), "\n") { if line != "" && !strings.HasPrefix(line, "export ") { - t.Errorf("stdout line %q would break `eval $(spinloop remote env)`:\n%s", line, stdout) + t.Errorf("stdout line %q would break `eval $(spinloop cloud env)`:\n%s", line, stdout) } } if !strings.Contains(stdout, "export OPENAI_API_KEY=sk-remote") { diff --git a/cmd/spinloop/harness_test.go b/cmd/spinloop/harness_test.go index 8627fe33..77f63a6f 100644 --- a/cmd/spinloop/harness_test.go +++ b/cmd/spinloop/harness_test.go @@ -7,7 +7,7 @@ import ( "strings" "testing" - "github.com/spinloop-ai/spinloop/internal/remote" + "github.com/spinloop-ai/spinloop/internal/cloud" "github.com/spinloop-ai/spinloop/internal/spinloop" ) @@ -288,7 +288,7 @@ func TestCode_SameLaunchAsHarnessOpen(t *testing.T) { // harness from what the environment reports as deployed. env := deployedEnvServer(t) defer env.Close() - registerEnv(t, "dev-1", remote.Config{ + registerEnv(t, "dev-1", cloud.Config{ StartURL: env.URL, StopURL: env.URL, EnvURL: env.URL, BaseURL: "http://198.51.100.1:8000/v1", Region: "eu-west-1", Environment: "dev-1", }) diff --git a/cmd/spinloop/last_active_test.go b/cmd/spinloop/last_active_test.go index 079f4490..39b39427 100644 --- a/cmd/spinloop/last_active_test.go +++ b/cmd/spinloop/last_active_test.go @@ -9,15 +9,15 @@ import ( "testing" ) -// The last-active figure appears in four places — `remote metrics` in bar, -// table and json, `remote status`, and `fleet metrics` — and every one of them +// The last-active figure appears in four places — `cloud metrics` in bar, +// table and json, `cloud status`, and `fleet metrics` — and every one of them // gates on the timestamp rather than the seconds. These tests hold that line, // because gating on the seconds hides the busiest engine there is: the daemon // omits idleSeconds at zero, so an engine working this instant sends a // timestamp and nothing else. // statsServer stands in for the stats Lambda, replying with whatever the test -// wants `spinloop remote metrics` to render. +// wants `spinloop cloud metrics` to render. func statsServer(t *testing.T, body string) { t.Helper() isolateConfig(t) @@ -27,8 +27,8 @@ func statsServer(t *testing.T, body string) { fmt.Fprint(w, body) })) t.Cleanup(server.Close) - writeRemoteConfig(t, server.URL) - t.Setenv("SPINLOOP_REMOTE_STATS_URL", server.URL) + writeCloudConfig(t, server.URL) + t.Setenv("SPINLOOP_CLOUD_STATS_URL", server.URL) } const runningWithActivity = `{ @@ -44,7 +44,7 @@ const runningWithActivity = `{ "idleSeconds": 125 }` -func TestRemoteMetricsBarShowsLastActive(t *testing.T) { +func TestCloudMetricsBarShowsLastActive(t *testing.T) { statsServer(t, runningWithActivity) out := captureStdout(t, func() { @@ -65,7 +65,7 @@ func TestRemoteMetricsBarShowsLastActive(t *testing.T) { } } -func TestRemoteMetricsTableShowsLastActive(t *testing.T) { +func TestCloudMetricsTableShowsLastActive(t *testing.T) { statsServer(t, runningWithActivity) out := captureStdout(t, func() { @@ -82,7 +82,7 @@ func TestRemoteMetricsTableShowsLastActive(t *testing.T) { } } -func TestRemoteMetricsJSONCarriesLastActive(t *testing.T) { +func TestCloudMetricsJSONCarriesLastActive(t *testing.T) { statsServer(t, runningWithActivity) out := captureStdout(t, func() { @@ -114,7 +114,7 @@ func TestRemoteMetricsJSONCarriesLastActive(t *testing.T) { // A stopped endpoint draws no bars and no token block, but when it last did // work is exactly what a stopped endpoint is worth asking about. -func TestRemoteMetricsStoppedStillShowsLastActive(t *testing.T) { +func TestCloudMetricsStoppedStillShowsLastActive(t *testing.T) { statsServer(t, `{ "environment": "dev", "state": "stopped", @@ -200,7 +200,7 @@ func TestLastActiveOmittedWithoutATimestamp(t *testing.T) { } // statusServer stands in for the start Lambda's GET branch, which is what -// `spinloop remote status` calls. +// `spinloop cloud status` calls. func statusServer(t *testing.T, body string) { t.Helper() isolateConfig(t) @@ -210,10 +210,10 @@ func statusServer(t *testing.T, body string) { fmt.Fprint(w, body) })) t.Cleanup(server.Close) - writeRemoteConfig(t, server.URL) + writeCloudConfig(t, server.URL) } -func TestRemoteStatusShowsLastActive(t *testing.T) { +func TestCloudStatusShowsLastActive(t *testing.T) { statusServer(t, `{ "state": "running", "healthy": true, @@ -243,7 +243,7 @@ func TestRemoteStatusShowsLastActive(t *testing.T) { } } -func TestRemoteStatusZeroIdleStillRenders(t *testing.T) { +func TestCloudStatusZeroIdleStillRenders(t *testing.T) { statusServer(t, `{"state": "running", "healthy": true, "lastActiveAt": "2026-08-10T10:00:00Z"}`) out := captureStdout(t, func() { @@ -258,7 +258,7 @@ func TestRemoteStatusZeroIdleStillRenders(t *testing.T) { // A stopped instance cannot be asked — reaching the daemon needs a running // box — so the control plane sends nothing and the command claims nothing. -func TestRemoteStatusOmitsLastActiveWhenAbsent(t *testing.T) { +func TestCloudStatusOmitsLastActiveWhenAbsent(t *testing.T) { statusServer(t, `{"state": "stopped", "healthy": false}`) out := captureStdout(t, func() { diff --git a/cmd/spinloop/logs.go b/cmd/spinloop/logs.go index 3172d98a..2ae71419 100644 --- a/cmd/spinloop/logs.go +++ b/cmd/spinloop/logs.go @@ -10,8 +10,8 @@ import ( "time" "github.com/spf13/cobra" + "github.com/spinloop-ai/spinloop/internal/cloud" "github.com/spinloop-ai/spinloop/internal/fleet" - "github.com/spinloop-ai/spinloop/internal/remote" ) func logsCmd() *cobra.Command { @@ -80,7 +80,7 @@ func cmdLogs(args []string) error { return execCmd(logsCmd(), args) } // any node is contacted. func validateLogQuery(q fleet.LogQuery, format string, limit int) error { switch q.Source { - case "", remote.LogSourceEngine, remote.LogSourceBoot, remote.LogSourceAll: + case "", cloud.LogSourceEngine, cloud.LogSourceBoot, cloud.LogSourceAll: default: return fmt.Errorf("--source must be engine, boot or all, got %q", q.Source) } diff --git a/cmd/spinloop/main.go b/cmd/spinloop/main.go index bc63657e..6bd42a52 100644 --- a/cmd/spinloop/main.go +++ b/cmd/spinloop/main.go @@ -45,6 +45,7 @@ import ( "github.com/spf13/cobra" "github.com/spf13/pflag" "github.com/spinloop-ai/spinloop/internal/catalog" + "github.com/spinloop-ai/spinloop/internal/cloud" "github.com/spinloop-ai/spinloop/internal/config" "github.com/spinloop-ai/spinloop/internal/contextsize" "github.com/spinloop-ai/spinloop/internal/discovery" @@ -52,7 +53,6 @@ import ( "github.com/spinloop-ai/spinloop/internal/harness" "github.com/spinloop-ai/spinloop/internal/lucinate" "github.com/spinloop-ai/spinloop/internal/opencode" - "github.com/spinloop-ai/spinloop/internal/remote" "github.com/spinloop-ai/spinloop/internal/spinloop" "github.com/spinloop-ai/spinloop/internal/spinloopsrc" ) @@ -77,7 +77,7 @@ func main() { os.Args = append(os.Args, "") } } - remote.SetCLIVersion(version) + cloud.SetCLIVersion(version) if err := run(os.Args[1:]); err != nil { fmt.Fprintln(os.Stderr, "Error:", err) os.Exit(1) @@ -193,7 +193,7 @@ func envFileDir(spinloopPath string) string { // the registered environment the selection is applied against, empty for a // local apply. resolve looks up API key variables — normally // opencode.EnvResolver of the Spinloop's local directory, but `spinloop -// harness` widens it with the key it fetched from a remote endpoint, which it +// harness` widens it with the key it fetched from a cloud endpoint, which it // is about to put in the launched agent's environment. gatewayLabel waives the // model-or-alias requirement below and renames the provider: non-empty only // when the selection is routed at a gateway with no model or alias of its own, @@ -221,22 +221,22 @@ func applySelection(sel spinloop.Selection, h harness.Harness, spinloopPath, env // removeSelection too, so apply and unapply stay symmetric. The flag is an // explicit name, so a missing registration is a mistake to report rather // than a config to wait for. - var envCfg *remote.Config + var envCfg *cloud.Config if envName != "" { - if !remote.IsEnvName(envName) { + if !cloud.IsEnvName(envName) { return fmt.Errorf("%q is not an environment name: an environment name is a plain identifier, with no path", envName) } - envPath, err := remote.EnvConfigPath(envName) + envPath, err := cloud.EnvConfigPath(envName) if err != nil { return err } if _, err := os.Stat(envPath); err != nil { if os.IsNotExist(err) { - return fmt.Errorf("environment %q is not registered: run `spinloop remote deploy --env %q` to create it", envName, envName) + return fmt.Errorf("environment %q is not registered: run `spinloop cloud deploy --env %q` to create it", envName, envName) } return err } - cfg, err := remote.LoadEnvironment(envName, viperGetenv()) + cfg, err := cloud.LoadEnvironment(envName, viperGetenv()) if err != nil { return err } @@ -245,7 +245,7 @@ func applySelection(sel spinloop.Selection, h harness.Harness, spinloopPath, env // The provider is now keyed on the environment; label it so it reads // distinctly from a local engine of the same kind in a model picker // (e.g. "llama.cpp (dev-2)" rather than another bare "llama.cpp"). - sel.DisplayName = catalog.RemoteProviderLabel(p.Name, envName) + sel.DisplayName = catalog.CloudProviderLabel(p.Name, envName) } else if gatewayLabel != "" { // A gateway-routed selection with no model of its own carries the // catalogue's shared "openai-compatible" id, which every gateway a @@ -257,12 +257,12 @@ func applySelection(sel spinloop.Selection, h harness.Harness, spinloopPath, env // gateway in its text at all, so a user searching a harness's model // picker for "gateway" would find nothing. sel.Provider = gatewayProviderKey(gatewayLabel) - sel.DisplayName = catalog.RemoteProviderLabel("Gateway", gatewayLabel) + sel.DisplayName = catalog.CloudProviderLabel("Gateway", gatewayLabel) } // A Spinloop applied against an environment states no BASEURL: the address // belongs to the deployment, which records it in the environment's - // registered remote.json. Take it from there — but only when the Spinloop + // registered cloud.json. Take it from there — but only when the Spinloop // stated none, so a hand-written BASEURL still wins. // The harness reports the base URL it wrote, so this needs no announcement // of its own beyond naming where it came from. @@ -340,7 +340,7 @@ func applySelection(sel spinloop.Selection, h harness.Harness, spinloopPath, env // convention that only cmdX functions report anything. The alternative is to // repeat that reporting at all four call sites, where one omission would leave // the user guessing which file was read. The line goes to stderr because -// `spinloop remote env` writes shell exports to stdout for `eval`, which a stray +// `spinloop cloud env` writes shell exports to stdout for `eval`, which a stray // prose line would break. func readSpinloop(usage, path string) (spinloop.Selection, string, error) { if path == "" { @@ -497,9 +497,9 @@ func applyCmd() *cobra.Command { Short: "apply a Spinloop file (defaults to ./Spinloop)", Long: `applies a Spinloop file — a declarative, Dockerfile-style description of one provider selection — as if you had run the equivalent add. With --env, -applies it against a registered remote environment: the provider is keyed on +applies it against a registered cloud environment: the provider is keyed on the environment's name and, absent a BASEURL, takes its address from the -environment's registered remote.json.`, +environment's registered cloud.json.`, Args: cobra.ArbitraryArgs, SilenceErrors: true, SilenceUsage: true, @@ -529,7 +529,7 @@ environment's registered remote.json.`, fs.StringVar(&providers, "providers", "", "path to a providers.yaml override") fs.StringVarP(&output, "output", "o", "", "max output tokens (overrides the Spinloop's OUTPUT)") fs.StringVarP(&harnessName, "harness", "H", "", "which harness to configure") - fs.StringVarP(&envName, "env", "e", "", "apply against this registered environment (its name keys the provider; its remote.json supplies the base URL)") + fs.StringVarP(&envName, "env", "e", "", "apply against this registered environment (its name keys the provider; its cloud.json supplies the base URL)") compRegister(c, "env", compEnvs) fs.SetInterspersed(false) c.ValidArgsFunction = aliasSlot @@ -939,7 +939,7 @@ func exportLimit(sel spinloop.Selection, st harness.ProviderState, values map[st // finds nothing and says so. func removeSelection(sel spinloop.Selection, h harness.Harness, spinloopPath, envName string) error { if envName != "" { - if !remote.IsEnvName(envName) { + if !cloud.IsEnvName(envName) { return fmt.Errorf("%q is not an environment name: an environment name is a plain identifier, with no path", envName) } sel.Provider = envName @@ -1027,17 +1027,17 @@ func cmdConfig(args []string) error { // export always wins. A catalogue that cannot be loaded is not fatal: launching // the agent matters more than the keys, and it will report its own auth error. // -// remoteResp carries the live API key and base URL from a running remote +// cloudResp carries the live API key and base URL from a running cloud // endpoint. When present, OPENAI_API_KEY and OPENAI_BASE_URL are injected so -// the harness can reach the remote without the user exporting them manually. -func harnessEnv(providersPath string, resolve func(string) string, remoteResp *remote.Response) []string { +// the harness can reach the cloud without the user exporting them manually. +func harnessEnv(providersPath string, resolve func(string) string, cloudResp *cloud.Response) []string { env := os.Environ() - if remoteResp != nil { + if cloudResp != nil { if os.Getenv("OPENAI_API_KEY") == "" { - env = append(env, "OPENAI_API_KEY="+remoteResp.APIKey) + env = append(env, "OPENAI_API_KEY="+cloudResp.APIKey) } if os.Getenv("OPENAI_BASE_URL") == "" { - env = append(env, "OPENAI_BASE_URL="+remoteResp.BaseURL) + env = append(env, "OPENAI_BASE_URL="+cloudResp.BaseURL) } } cat, err := catalog.LoadFrom(catalog.ResolveCatalogPath(providersPath)) @@ -1098,7 +1098,7 @@ func lucinateLaunchKey(providersPath string, resolve func(string) string, sel sp // non-empty value. It differs from setEnvIfAbsent in treating an exported but // empty variable as unset — for an address or a key, "" is not a deliberate // choice worth preserving, it is a gap, and this is the rule harnessEnv already -// applies to the remote endpoint's values. +// applies to the cloud endpoint's values. func setEnvIfBlank(env []string, key, value string) []string { prefix := key + "=" for i, kv := range env { @@ -1132,7 +1132,7 @@ func setEnvIfAbsent(env []string, key, value string) []string { // dir is the Spinloop's directory. base already holds spinloop's process // environment and any provider key it resolved, so the `.env` only fills genuine // gaps and ENV alone can override an exported variable — the precedence is -// ENV > process environment > `.env`, the same rule the remote commands follow. +// ENV > process environment > `.env`, the same rule the cloud commands follow. // A `.env` that cannot be read is not fatal; the agent launches without it. func overlayLocalEnv(base []string, sel spinloop.Selection, dir string) []string { out := append([]string(nil), base...) @@ -1282,14 +1282,14 @@ func namesAnSpinloopOrAlias(arg string) bool { // will be forwarded to the harness, inspected only to catch a path that was // meant for the flag. // It returns the applied Spinloop's directory and selection, so the launched -// agent can be given the same keys the apply resolved, along with the remote +// agent can be given the same keys the apply resolved, along with the cloud // endpoint's live environment when --env names an environment. // -// The remote key is fetched before the apply, not after, so the apply resolves +// The cloud key is fetched before the apply, not after, so the apply resolves // against the environment the agent will actually run with. Fetching it // afterwards left the apply warning that no key was set while the launch was // about to supply one. -func applyBeforeLaunch(f spinloopPathFlag, providers string, h harness.Harness, rest []string, route routeOptions) (spinloop.Selection, string, *remote.Response, *fleet.Choice, error) { +func applyBeforeLaunch(f spinloopPathFlag, providers string, h harness.Harness, rest []string, route routeOptions) (spinloop.Selection, string, *cloud.Response, *fleet.Choice, error) { // The flag's value has to be attached, so `--spinloop ./dev/Spinloop` (or // `--spinloop q3`) would otherwise apply ./Spinloop and quietly hand the path // or alias to the harness. @@ -1300,11 +1300,11 @@ func applyBeforeLaunch(f spinloopPathFlag, providers string, h harness.Harness, if err != nil { return spinloop.Selection{}, "", nil, nil, err } - sel, envDir, remoteResp, choice, err := applyRoutedSpinloop(sel, path, providers, h, route, false) + sel, envDir, cloudResp, choice, err := applyRoutedSpinloop(sel, path, providers, h, route, false) if err != nil { return spinloop.Selection{}, "", nil, nil, err } - return sel, envDir, remoteResp, choice, nil + return sel, envDir, cloudResp, choice, nil } // applyFromEnvironment configures the harness with no Spinloop at all: a bare @@ -1312,7 +1312,7 @@ func applyBeforeLaunch(f spinloopPathFlag, providers string, h harness.Harness, // counterpart for that case — same return shape, same launch continuation — // except there is no Spinloop to read, so routing and the provider selection // come entirely from the named environment's live deploy-config. -func applyFromEnvironment(providers string, h harness.Harness, route routeOptions) (spinloop.Selection, string, *remote.Response, *fleet.Choice, error) { +func applyFromEnvironment(providers string, h harness.Harness, route routeOptions) (spinloop.Selection, string, *cloud.Response, *fleet.Choice, error) { return applyRoutedSpinloop(spinloop.Selection{}, "", providers, h, route, true) } @@ -1337,8 +1337,8 @@ const gatewayProviderID = "openai-compatible" // applyRoutedSpinloop's already — this supplies only the selection that says // "a gateway is the endpoint", which is the one thing a Spinloop would // otherwise have carried. -func applyFromGateway(providers string, h harness.Harness, route routeOptions) (spinloop.Selection, string, *remote.Response, *fleet.Choice, error) { - fail := func(err error) (spinloop.Selection, string, *remote.Response, *fleet.Choice, error) { +func applyFromGateway(providers string, h harness.Harness, route routeOptions) (spinloop.Selection, string, *cloud.Response, *fleet.Choice, error) { + fail := func(err error) (spinloop.Selection, string, *cloud.Response, *fleet.Choice, error) { return spinloop.Selection{}, "", nil, nil, err } cfg, err := fleet.Resolve(route.fleetPath) @@ -1354,7 +1354,7 @@ func applyFromGateway(providers string, h harness.Harness, route routeOptions) ( // applyRoutedSpinloop routes an already-read Spinloop and applies it to the // harness that is about to be launched: routing first, so a launch that cannot -// find a node leaves the harness config exactly as it was, then the remote +// find a node leaves the harness config exactly as it was, then the cloud // fetch and the apply themselves. `spinloop harness` reads its Spinloop on the // way in; `spinloop fleet harness` reads its own, because with none it fails on // its own terms. Both then run this one path. @@ -1367,7 +1367,7 @@ func applyFromGateway(providers string, h harness.Harness, route routeOptions) ( // nothing deployed (or an env Lambda predating this) fails the launch rather // than reaching applySelection's generic "needs a model or an alias" error, // so the message names the actual cause and how to fix it. -func applyRoutedSpinloop(sel spinloop.Selection, path string, providers string, h harness.Harness, route routeOptions, autoConfigure bool) (spinloop.Selection, string, *remote.Response, *fleet.Choice, error) { +func applyRoutedSpinloop(sel spinloop.Selection, path string, providers string, h harness.Harness, route routeOptions, autoConfigure bool) (spinloop.Selection, string, *cloud.Response, *fleet.Choice, error) { // As for apply, --providers overrides the catalogue the selection resolves // against (a Spinloop never names one). sel.Providers = providers @@ -1391,7 +1391,7 @@ func applyRoutedSpinloop(sel spinloop.Selection, path string, providers string, } localResolve := opencode.EnvResolver(envDir) // Routing runs before the apply, and before anything is printed about - // applying, for the reason the remote fetch does: a launch that cannot find + // applying, for the reason the cloud fetch does: a launch that cannot find // a node must leave the harness config exactly as it was. choice, err := routeThroughFleet(sel, path, route) if err != nil { @@ -1399,7 +1399,7 @@ func applyRoutedSpinloop(sel spinloop.Selection, path string, providers string, } if choice != nil { // The chosen node's address is what the apply writes, in the slot a - // remote endpoint's address is written to. + // cloud endpoint's address is written to. sel.BaseURL = choice.BaseURL } // label names what is being applied in the messages below: the Spinloop's @@ -1420,30 +1420,30 @@ func applyRoutedSpinloop(sel spinloop.Selection, path string, providers string, } // Before the apply, so a launch that cannot authenticate stops without // having rewritten the harness config. - remoteResp, err := fetchRemoteEnv(sel, route.envName, localResolve) + cloudResp, err := fetchCloudEnv(sel, route.envName, localResolve) if err != nil { return spinloop.Selection{}, "", nil, nil, err } if autoConfigure { - if remoteResp == nil || !remoteResp.Deployed || remoteResp.Runner == "" || remoteResp.ServedName == "" { + if cloudResp == nil || !cloudResp.Deployed || cloudResp.Runner == "" || cloudResp.ServedName == "" { return spinloop.Selection{}, "", nil, nil, fmt.Errorf( "nothing is deployed to environment %q to configure the harness with: "+ - "run `spinloop remote deploy --env %s` to deploy one, "+ - "or `spinloop remote bootstrap` to update the control plane if %s already has something deployed", + "run `spinloop cloud deploy --env %s` to deploy one, "+ + "or `spinloop cloud bootstrap` to update the control plane if %s already has something deployed", route.envName, route.envName, route.envName) } - provider, err := providerForRunner(remoteResp.Runner) + provider, err := providerForRunner(cloudResp.Runner) if err != nil { return spinloop.Selection{}, "", nil, nil, err } sel.Provider = provider - sel.Alias = remoteResp.ServedName - if remoteResp.ContextSize > 0 { - sel.Context = strconv.Itoa(remoteResp.ContextSize) + sel.Alias = cloudResp.ServedName + if cloudResp.ContextSize > 0 { + sel.Context = strconv.Itoa(cloudResp.ContextSize) } fmt.Printf("Configuring from what is deployed to %s.\n\n", route.envName) } - resolve := remoteLaunchResolver(localResolve, remoteResp) + resolve := cloudLaunchResolver(localResolve, cloudResp) if choice != nil && choice.APIKey != "" { resolve = fleetLaunchResolver(resolve, choice.APIKey) } @@ -1494,7 +1494,7 @@ func applyRoutedSpinloop(sel spinloop.Selection, path string, providers string, return spinloop.Selection{}, "", nil, nil, err } fmt.Println() - return sel, envDir, remoteResp, choice, nil + return sel, envDir, cloudResp, choice, nil } // gatewayModelsTimeout bounds fetchGatewayModels, so a gateway that never @@ -1573,25 +1573,25 @@ func slugify(s string) string { // fleetLaunchResolver extends a lookup with the engine key of the node a launch // was routed to. The apply then writes a config knowing the key will be there, -// exactly as the remote path does — the missing-key warning is left for when a +// exactly as the cloud path does — the missing-key warning is left for when a // key really is missing. func fleetLaunchResolver(base func(string) string, key string) func(string) string { return func(name string) string { if v := base(name); v != "" { return v } - if name == remoteAPIKeyEnv { + if name == cloudAPIKeyEnv { return key } return "" } } -// remoteEnvTimeout bounds the call that fetches a remote endpoint's key, so a +// cloudEnvTimeout bounds the call that fetches a cloud endpoint's key, so a // control plane that never answers delays the launch rather than blocking it. -const remoteEnvTimeout = 30 * time.Second +const cloudEnvTimeout = 30 * time.Second -// fetchRemoteEnv returns the live base URL and API key of the endpoint the +// fetchCloudEnv returns the live base URL and API key of the endpoint the // environment named by --env serves, or nil when the flag is absent. The // endpoint is started with a key that only the control plane knows, so this is // the one place it can come from; `spinloop harness` puts it in the @@ -1606,43 +1606,43 @@ const remoteEnvTimeout = 30 * time.Second // An environment whose configuration is not registered is a different failure: // the name points at nothing on this machine, so it is reported as-is rather // than downgraded. -func fetchRemoteEnv(sel spinloop.Selection, envName string, resolve func(string) string) (*remote.Response, error) { +func fetchCloudEnv(sel spinloop.Selection, envName string, resolve func(string) string) (*cloud.Response, error) { if envName == "" { return nil, nil } - if !remote.IsEnvName(envName) { + if !cloud.IsEnvName(envName) { return nil, fmt.Errorf("%q is not an environment name: an environment name is a plain identifier, with no path", envName) } - envPath, err := remote.EnvConfigPath(envName) + envPath, err := cloud.EnvConfigPath(envName) if err != nil { return nil, err } if _, err := os.Stat(envPath); err != nil { if os.IsNotExist(err) { - return nil, fmt.Errorf("environment %q is not registered: run `spinloop remote deploy --env %q` to create it", envName, envName) + return nil, fmt.Errorf("environment %q is not registered: run `spinloop cloud deploy --env %q` to create it", envName, envName) } return nil, err } // The call crosses the network, and a cold control plane is not instant. fmt.Fprintf(os.Stderr, "Fetching the endpoint's environment from %s...\n", envName) - cfg, err := remote.LoadEnvironment(envName, viperGetenv()) + cfg, err := cloud.LoadEnvironment(envName, viperGetenv()) if err == nil { - ctx, cancel := context.WithTimeout(context.Background(), remoteEnvTimeout) + ctx, cancel := context.WithTimeout(context.Background(), cloudEnvTimeout) defer cancel() - var resp *remote.Response - if resp, err = remote.Env(ctx, cfg); err == nil { + var resp *cloud.Response + if resp, err = cloud.Env(ctx, cfg); err == nil { return resp, nil } } if localKey(sel, resolve) == "" { return nil, fmt.Errorf( "could not fetch the API key for %s: %w\n"+ - "Start the endpoint with `spinloop remote start --env %s` if it is stopped, or export %s yourself", - envName, err, envName, remoteAPIKeyEnv) + "Start the endpoint with `spinloop cloud start --env %s` if it is stopped, or export %s yourself", + envName, err, envName, cloudAPIKeyEnv) } fmt.Fprintf(os.Stderr, "Warning: could not fetch the API key for %s (%v).\nCarrying on with the %s already set here.\n", - envName, err, remoteAPIKeyEnv) + envName, err, cloudAPIKeyEnv) return nil, nil } @@ -1659,19 +1659,19 @@ func localKeyUnder(sel spinloop.Selection, resolve func(string) string, name str return resolve(name) } -// localKey resolves under the remote API key's variable — the one a REMOTE +// localKey resolves under the cloud API key's variable — the one a REMOTE // endpoint authenticates under. func localKey(sel spinloop.Selection, resolve func(string) string) string { - return localKeyUnder(sel, resolve, remoteAPIKeyEnv) + return localKeyUnder(sel, resolve, cloudAPIKeyEnv) } -// remoteLaunchResolver extends an environment-variable lookup with the key -// fetched from a running remote endpoint. `spinloop harness` gives that key to +// cloudLaunchResolver extends an environment-variable lookup with the key +// fetched from a running cloud endpoint. `spinloop harness` gives that key to // the agent it launches, so an apply on the same path should resolve it too: // the config it writes is complete, and the missing-key warning is left for // the case where the key really is missing. resp is nil when the Spinloop names -// no remote, or the fetch failed, and the lookup is then unchanged. -func remoteLaunchResolver(base func(string) string, resp *remote.Response) func(string) string { +// no cloud, or the fetch failed, and the lookup is then unchanged. +func cloudLaunchResolver(base func(string) string, resp *cloud.Response) func(string) string { if resp == nil || resp.APIKey == "" { return base } @@ -1679,7 +1679,7 @@ func remoteLaunchResolver(base func(string) string, resp *remote.Response) func( if v := base(name); v != "" { return v } - if name == remoteAPIKeyEnv { + if name == cloudAPIKeyEnv { return resp.APIKey } return "" diff --git a/cmd/spinloop/metrics_render.go b/cmd/spinloop/metrics_render.go index ec682634..c4dbb6da 100644 --- a/cmd/spinloop/metrics_render.go +++ b/cmd/spinloop/metrics_render.go @@ -1,8 +1,8 @@ -// Shared metrics rendering. Both `spinloop remote metrics` (one cloud endpoint) +// Shared metrics rendering. Both `spinloop cloud metrics` (one cloud endpoint) // and `spinloop fleet metrics` (a node per machine) display the same // internal/metrics stats, so the parts that draw those stats live here and // each caller supplies only its own heading — the environment and instance -// type for remote, the node name for fleet. +// type for cloud, the node name for fleet. package main @@ -118,7 +118,7 @@ func renderActiveKeyValue(w io.Writer, lastActiveAt string, idleSeconds int, ret } // validateMetricsFormat rejects a --format value the metrics commands do not -// understand, naming the ones they do. Both `remote metrics` and +// understand, naming the ones they do. Both `cloud metrics` and // `fleet metrics` run it before doing any work. func validateMetricsFormat(format string) error { switch format { diff --git a/cmd/spinloop/metrics_render_test.go b/cmd/spinloop/metrics_render_test.go index 63fb5b48..2c50db47 100644 --- a/cmd/spinloop/metrics_render_test.go +++ b/cmd/spinloop/metrics_render_test.go @@ -6,9 +6,9 @@ import ( "testing" "github.com/charmbracelet/lipgloss" + "github.com/spinloop-ai/spinloop/internal/cloud" "github.com/spinloop-ai/spinloop/internal/fleet" "github.com/spinloop-ai/spinloop/internal/metrics" - "github.com/spinloop-ai/spinloop/internal/remote" ) func ptrPct(v float64) *float64 { return &v } @@ -480,14 +480,14 @@ func TestFormatMetricsBarStoppedWithHistory(t *testing.T) { } func TestFormatMetricsBarRunning(t *testing.T) { - resp := &remote.StatsResponse{ + resp := &cloud.StatsResponse{ Environment: "prod", State: "running", InstanceType: "g5.xlarge", ModelID: "org/qwen:q4", Version: "0.4.3", LastActiveAt: "2026-08-21T10:00:00Z", IdleSeconds: 3, CPU: &metrics.CpuStat{Utilization: 62}, Memory: &metrics.MemoryStat{Total: 1000, Used: 300}, GPUs: []metrics.GpuStat{{Index: 0, Name: "H100", Utilization: 61, MemoryUsed: 80, MemoryTotal: 160}}, - Tokens: &remote.TokenStats{Running: 2, PromptTokens: 4096, GenerationTokens: 1024, Requests: ptrInt(17)}, + Tokens: &cloud.TokenStats{Running: 2, PromptTokens: 4096, GenerationTokens: 1024, Requests: ptrInt(17)}, History: []metrics.HistorySample{ {Time: 1, CPU: ptrPct(10), Mem: ptrPct(20), GPUs: []metrics.HistoryGPU{{Index: 0, Util: 50, Mem: ptrPct(50)}}}, {Time: 2, CPU: ptrPct(20), Mem: ptrPct(30), GPUs: []metrics.HistoryGPU{{Index: 0, Util: 61, Mem: ptrPct(50)}}}, @@ -521,10 +521,10 @@ func TestFormatMetricsBarRunning(t *testing.T) { // An engine family whose metrics expose no cumulative request counter yields // statistics without the figure, and the token block draws no line for it. func TestFormatMetricsBarNoRequestCount(t *testing.T) { - resp := &remote.StatsResponse{ + resp := &cloud.StatsResponse{ Environment: "prod", State: "running", InstanceType: "g5.xlarge", ModelID: "org/qwen:q4", Version: "0.4.3", - Tokens: &remote.TokenStats{Running: 2, PromptTokens: 4096, GenerationTokens: 1024}, + Tokens: &cloud.TokenStats{Running: 2, PromptTokens: 4096, GenerationTokens: 1024}, } var b bytes.Buffer if err := renderFleetMetrics(&b, nodeResultsFor(resp), "bar"); err != nil { @@ -540,7 +540,7 @@ func TestFormatMetricsBarNoRequestCount(t *testing.T) { } func TestFormatMetricsJSONCarriesHistory(t *testing.T) { - resp := &remote.StatsResponse{ + resp := &cloud.StatsResponse{ Environment: "prod", State: "running", CPU: &metrics.CpuStat{Utilization: 62}, History: []metrics.HistorySample{ @@ -571,7 +571,7 @@ func TestFormatMetricsJSONCarriesHistory(t *testing.T) { // nodeResultsFor turns a control-plane stats reply into the one node result a // fan-out would produce for it, so a test can state its input as the reply and // assert on what the renderer draws. -func nodeResultsFor(resp *remote.StatsResponse) []fleet.NodeResult { +func nodeResultsFor(resp *cloud.StatsResponse) []fleet.NodeResult { return []fleet.NodeResult{{ Name: resp.Environment, Outcome: fleet.OutcomeOK, Instance: fleet.Instance{ diff --git a/cmd/spinloop/palette.go b/cmd/spinloop/palette.go index 4467f25b..fd3e2182 100644 --- a/cmd/spinloop/palette.go +++ b/cmd/spinloop/palette.go @@ -1,6 +1,6 @@ // The colours and the spinner every surface of the CLI draws from — the // dashboard's tiles and title bar, `fleet deploy`'s progress lines, the -// resource bars `fleet metrics` and `remote metrics` print. They live together +// resource bars `fleet metrics` and `cloud metrics` print. They live together // so a second surface finds them rather than writing its own copy, which is // what left the same ten spinner frames declared twice in this package. // diff --git a/cmd/spinloop/read_spinloop_env.go b/cmd/spinloop/read_spinloop_env.go index 0fb38b5b..6df87f6a 100644 --- a/cmd/spinloop/read_spinloop_env.go +++ b/cmd/spinloop/read_spinloop_env.go @@ -2,10 +2,10 @@ // from --env or --fleet — and is read only for the environment it carries: its // ENV instructions and the .env beside it, applied before any control-plane or // daemon work. That is what lets AWS credentials, a profile, or the -// SPINLOOP_REMOTE_* overrides live in a project's Spinloop rather than in the +// SPINLOOP_CLOUD_* overrides live in a project's Spinloop rather than in the // shell that happens to be running the command. // -// The `remote` subcommands this replaces also consulted ./Spinloop when none +// The `cloud` subcommands this replaces also consulted ./Spinloop when none // was named. The verbs do not: a file sitting in the working directory should // not silently set environment variables for a command that reads a fleet, and // naming it is one flag. Nothing else about the rule changes. diff --git a/cmd/spinloop/read_spinloop_env_test.go b/cmd/spinloop/read_spinloop_env_test.go index 5f03e801..190b1ee9 100644 --- a/cmd/spinloop/read_spinloop_env_test.go +++ b/cmd/spinloop/read_spinloop_env_test.go @@ -8,7 +8,7 @@ import ( ) // A Spinloop's ENV instructions reach the command that was given it, which is -// what lets a project's AWS profile or SPINLOOP_REMOTE_* overrides live in the +// what lets a project's AWS profile or SPINLOOP_CLOUD_* overrides live in the // Spinloop rather than in whatever shell is running the command. func TestReadSpinloopEnv_AppliesENVInstructions(t *testing.T) { t.Setenv("SPINLOOP_READ_ENV_PROBE", "") diff --git a/cmd/spinloop/retain_render_test.go b/cmd/spinloop/retain_render_test.go index 75270c73..3bf6f04a 100644 --- a/cmd/spinloop/retain_render_test.go +++ b/cmd/spinloop/retain_render_test.go @@ -44,7 +44,7 @@ func aLineContaining(out string, phrases ...string) string { // A kept, active endpoint draws the keep after the active figure, on the same // line — the point of sharing the line is that "when did it last do anything" // and "how long is it kept" are read at a glance, not on two rows. -func TestRemoteMetricsBarKeepsOnTheActiveLine(t *testing.T) { +func TestCloudMetricsBarKeepsOnTheActiveLine(t *testing.T) { deadline := keepNow(t, 2*time.Hour) statsServer(t, `{ "environment": "dev", @@ -75,7 +75,7 @@ func TestRemoteMetricsBarKeepsOnTheActiveLine(t *testing.T) { } // The table format draws the same combined line as a key-value row. -func TestRemoteMetricsTableKeepsOnTheActiveRow(t *testing.T) { +func TestCloudMetricsTableKeepsOnTheActiveRow(t *testing.T) { deadline := keepNow(t, 2*time.Hour) statsServer(t, `{ "environment": "dev", @@ -137,7 +137,7 @@ func TestKeepDurationRendersRelatively(t *testing.T) { // A stopped environment can still be kept: the deadline is the control plane's, // not the engine's, so the keep survives the non-running short-circuit in both // formats. -func TestRemoteMetricsStoppedKeptStillShowsKeep(t *testing.T) { +func TestCloudMetricsStoppedKeptStillShowsKeep(t *testing.T) { deadline := keepNow(t, 4*time.Hour) for format := range map[string]bool{"bar": true, "table": true} { t.Run(format, func(t *testing.T) { @@ -162,7 +162,7 @@ func TestRemoteMetricsStoppedKeptStillShowsKeep(t *testing.T) { // No deadline on the read, no keep: the renderer does not invent one, and it // leaves the active figure (now just "active") in place. -func TestRemoteMetricsOmitsKeepWhenAbsent(t *testing.T) { +func TestCloudMetricsOmitsKeepWhenAbsent(t *testing.T) { for _, format := range []string{"bar", "table"} { t.Run(format, func(t *testing.T) { statsServer(t, `{ @@ -189,7 +189,7 @@ func TestRemoteMetricsOmitsKeepWhenAbsent(t *testing.T) { // The deadline rides the node's stats, so `fleet metrics` draws the keep after // the active figure whatever the node's state. A faked daemon carrying the field -// stands in for a kept remote environment on the same render path. +// stands in for a kept cloud environment on the same render path. func TestFleetMetricsShowsKeep(t *testing.T) { deadline := keepNow(t, 2*time.Hour) fleetNodeWithMetrics(t, map[string]any{ diff --git a/cmd/spinloop/root_dispatch_test.go b/cmd/spinloop/root_dispatch_test.go index cbd7959f..db116f58 100644 --- a/cmd/spinloop/root_dispatch_test.go +++ b/cmd/spinloop/root_dispatch_test.go @@ -80,7 +80,7 @@ func TestRoot_MovedSpellingsNameTheirNewHome(t *testing.T) { // TestRoot_MovedSubcommandSpellingsNameTheirNewHome pins the same signpost one // level down, for a subcommand a group used to have: "fleet metrics" and -// "remote status" and their siblings must each name the top-level verb that +// "cloud status" and their siblings must each name the top-level verb that // replaced them, rather than cobra's ordinary unknown-command error. func TestRoot_MovedSubcommandSpellingsNameTheirNewHome(t *testing.T) { isolateConfig(t) diff --git a/cmd/spinloop/route.go b/cmd/spinloop/route.go index 7dd1b047..552ca4a7 100644 --- a/cmd/spinloop/route.go +++ b/cmd/spinloop/route.go @@ -1,7 +1,7 @@ // Routing a launch through the fleet: choosing the node the agent talks to. It -// sits beside the remote path in main.go — both answer "where does this agent +// sits beside the cloud path in main.go — both answer "where does this agent // send its requests", one by asking a control plane and one by choosing a -// machine — and it runs before the apply for the same reason the remote fetch +// machine — and it runs before the apply for the same reason the cloud fetch // does: a failed route must leave the harness config alone. package main diff --git a/cmd/spinloop/serve.go b/cmd/spinloop/serve.go index 4ed091be..b3ca26fd 100644 --- a/cmd/spinloop/serve.go +++ b/cmd/spinloop/serve.go @@ -1,5 +1,5 @@ // Serve: launching a local inference server for a Spinloop. The engine is chosen -// by the Spinloop's PROVIDER, the same way `spinloop remote deploy` picks a cloud +// by the Spinloop's PROVIDER, the same way `spinloop cloud deploy` picks a cloud // runner, so one file describes both what dresses the harness and what serves // it. Kept out of main.go so the dispatch-coverage scan in complete_test.go only // ever sees run()'s own switch. diff --git a/cmd/spinloop/status.go b/cmd/spinloop/status.go index 1387213a..bea6adda 100644 --- a/cmd/spinloop/status.go +++ b/cmd/spinloop/status.go @@ -28,7 +28,7 @@ The target is a registered environment (--env), a fleet file (--fleet), or the fleet.yaml in the working directory. A node that cannot be reached is a row saying so, not a failure: one unreachable machine never blanks the rest. -An environment's endpoint address is spinloop remote env, and its retention +An environment's endpoint address is spinloop cloud env, and its retention deadline spinloop fleet metrics — neither is a column here, because neither applies to every node.`, Args: cobra.NoArgs, diff --git a/cmd/spinloop/status_render.go b/cmd/spinloop/status_render.go index 785d72c3..711c6abc 100644 --- a/cmd/spinloop/status_render.go +++ b/cmd/spinloop/status_render.go @@ -1,9 +1,9 @@ -// Shared status rendering. Both `spinloop remote status` (one cloud endpoint) and +// Shared status rendering. Both `spinloop cloud status` (one cloud endpoint) and // `spinloop fleet status` (a node per machine) report the same facts about an // spinloop-driven inference endpoint: its state, what it is serving, how long since // it last did work, and its spinloop version. Those facts live here, computed once, // so the two commands cannot word or compute them differently. Each command layers -// its own layout and any facts the other does not carry on top: the remote keeps +// its own layout and any facts the other does not carry on top: the cloud keeps // its key-value block and its endpoint's address and health, the fleet keeps its // one-node-per-row table. @@ -43,7 +43,7 @@ type statusFact struct { // servingText is the "what it serves" text: runner and model, then the uptime and // the since-last-work and version, in the order and wording both commands use. -// The fleet renders it as its table cell; the remote reads the same pieces for +// The fleet renders it as its table cell; the cloud reads the same pieces for // its lines. func (f statusFact) servingText() string { var serving string diff --git a/cmd/spinloop/status_render_test.go b/cmd/spinloop/status_render_test.go index 39e42a5e..2bd1c212 100644 --- a/cmd/spinloop/status_render_test.go +++ b/cmd/spinloop/status_render_test.go @@ -7,7 +7,7 @@ import ( "github.com/spinloop-ai/spinloop/internal/daemon" ) -// The shared status view is where `remote status` and `fleet status` agree on +// The shared status view is where `cloud status` and `fleet status` agree on // the facts they both carry; this pins the wording so the two cannot diverge. func TestStatusFactServingText(t *testing.T) { f := statusFact{ diff --git a/cmd/spinloop/target_test.go b/cmd/spinloop/target_test.go index 586cb141..146d2185 100644 --- a/cmd/spinloop/target_test.go +++ b/cmd/spinloop/target_test.go @@ -6,8 +6,8 @@ import ( "strings" "testing" + "github.com/spinloop-ai/spinloop/internal/cloud" "github.com/spinloop-ai/spinloop/internal/fleet" - "github.com/spinloop-ai/spinloop/internal/remote" ) // registerTargetEnv points the registry at a temp config directory and @@ -16,7 +16,7 @@ import ( func registerTargetEnv(t *testing.T, name string) { t.Helper() t.Setenv("SPINLOOP_CONFIG_DIR", t.TempDir()) - registerEnv(t, name, remote.Config{ + registerEnv(t, name, cloud.Config{ StartURL: "https://s.example/start", StopURL: "https://s.example/stop", Region: "us-east-1", @@ -39,8 +39,8 @@ func TestResolveFleetTarget(t *testing.T) { if len(cfg.Nodes) != 1 || cfg.Nodes[0].Name != "prod" { t.Fatalf("nodes = %+v, want just prod", cfg.Nodes) } - if cfg.Nodes[0].Kind != fleet.KindRemote { - t.Errorf("kind = %q, want %q", cfg.Nodes[0].Kind, fleet.KindRemote) + if cfg.Nodes[0].Kind != fleet.KindCloud { + t.Errorf("kind = %q, want %q", cfg.Nodes[0].Kind, fleet.KindCloud) } }) @@ -127,7 +127,7 @@ func TestResolveFleetTargetUnregisteredEnv(t *testing.T) { if err == nil { t.Fatal("want an error") } - if !strings.Contains(err.Error(), "remotes/nope/remote.json") { + if !strings.Contains(err.Error(), "clouds/nope/cloud.json") { t.Errorf("error %q does not name the environment's registry path", err) } } @@ -186,10 +186,10 @@ func TestDescribeTarget(t *testing.T) { } // Every fleet command that takes a target completes --env from the registered -// environments, the same source `remote --env` completes from. +// environments, the same source `cloud --env` completes from. func TestFleetCommandsCompleteEnv(t *testing.T) { registerTargetEnv(t, "prod") - registerEnv(t, "staging", remote.Config{ + registerEnv(t, "staging", cloud.Config{ StartURL: "https://s.example/start", StopURL: "https://s.example/stop", Region: "us-east-1", @@ -222,7 +222,7 @@ func TestFleetStatusEnvMatchesAOneNodeFile(t *testing.T) { stubAWSEnv(t) up := stateServer(t) t.Setenv("SPINLOOP_CONFIG_DIR", t.TempDir()) - registerEnv(t, "prod", remote.Config{ + registerEnv(t, "prod", cloud.Config{ StartURL: up.URL, StopURL: up.URL, StatsURL: up.URL, @@ -231,7 +231,7 @@ func TestFleetStatusEnvMatchesAOneNodeFile(t *testing.T) { }) // Named by a fleet file holding exactly that node. - writeFleetFile(t, "nodes:\n - name: prod\n kind: remote\n") + writeFleetFile(t, "nodes:\n - name: prod\n kind: cloud\n") fromFile := captureStdout(t, func() { if err := cmdStatus(nil); err != nil { t.Errorf("fleet status returned %v", err) @@ -260,7 +260,7 @@ func TestFleetStatusEnvNeedsNoFleetFile(t *testing.T) { stubAWSEnv(t) up := stateServer(t) t.Setenv("SPINLOOP_CONFIG_DIR", t.TempDir()) - registerEnv(t, "prod", remote.Config{ + registerEnv(t, "prod", cloud.Config{ StartURL: up.URL, StopURL: up.URL, StatsURL: up.URL, @@ -283,7 +283,7 @@ func TestFleetStatusEnvNeedsNoFleetFile(t *testing.T) { // resolver directly. func TestFleetStatusRefusesEnvAndFleet(t *testing.T) { t.Setenv("SPINLOOP_CONFIG_DIR", t.TempDir()) - registerEnv(t, "prod", remote.Config{ + registerEnv(t, "prod", cloud.Config{ StartURL: "https://s.example/start", StopURL: "https://s.example/stop", Region: "us-east-1", diff --git a/cmd/spinloop/viper_test.go b/cmd/spinloop/viper_test.go index 2d67fbc5..a7d03e66 100644 --- a/cmd/spinloop/viper_test.go +++ b/cmd/spinloop/viper_test.go @@ -7,7 +7,7 @@ import ( "strings" "testing" - "github.com/spinloop-ai/spinloop/internal/remote" + "github.com/spinloop-ai/spinloop/internal/cloud" ) // TestViperSpinloopAliasPrecedence pins the SPINLOOP_ALIAS resolution through the @@ -61,16 +61,16 @@ func TestViperDefaultSpinloopNamed(t *testing.T) { } } -// TestViperRemoteEnvPrecedence pins, for every SPINLOOP_REMOTE_* variable, the +// TestViperCloudEnvPrecedence pins, for every SPINLOOP_CLOUD_* variable, the // resolution the CLI's Viper gives: an exported variable beats the same key in -// the remote config file, and an unset variable falls through to the file. No +// the cloud config file, and an unset variable falls through to the file. No // control call is made — only the Config the commands would take is asserted. -func TestViperRemoteEnvPrecedence(t *testing.T) { +func TestViperCloudEnvPrecedence(t *testing.T) { isolateConfig(t) t.Chdir(t.TempDir()) // no ./Spinloop, so the per-user file is consulted stubAWSEnv(t) - file := remote.Config{ + file := cloud.Config{ StartURL: "https://file.example/start", StopURL: "https://file.example/stop", DeployURL: "https://file.example/deploy", @@ -80,7 +80,7 @@ func TestViperRemoteEnvPrecedence(t *testing.T) { Region: "us-east-1", Environment: "default", } - path := must1(remote.EnvConfigPath("default")) + path := must1(cloud.EnvConfigPath("default")) if err := os.MkdirAll(filepath.Dir(path), 0o700); err != nil { t.Fatal(err) } @@ -93,20 +93,20 @@ func TestViperRemoteEnvPrecedence(t *testing.T) { } const envValue = "https://env.example/wins" - legs := map[string]func(remote.Config) string{ - "SPINLOOP_REMOTE_START_URL": func(c remote.Config) string { return c.StartURL }, - "SPINLOOP_REMOTE_STOP_URL": func(c remote.Config) string { return c.StopURL }, - "SPINLOOP_REMOTE_DEPLOY_URL": func(c remote.Config) string { return c.DeployURL }, - "SPINLOOP_REMOTE_STATS_URL": func(c remote.Config) string { return c.StatsURL }, - "SPINLOOP_REMOTE_ENV_URL": func(c remote.Config) string { return c.EnvURL }, - "SPINLOOP_REMOTE_UPDATE_URL": func(c remote.Config) string { return c.UpdateURL }, - "SPINLOOP_REMOTE_REGION": func(c remote.Config) string { return c.Region }, + legs := map[string]func(cloud.Config) string{ + "SPINLOOP_CLOUD_START_URL": func(c cloud.Config) string { return c.StartURL }, + "SPINLOOP_CLOUD_STOP_URL": func(c cloud.Config) string { return c.StopURL }, + "SPINLOOP_CLOUD_DEPLOY_URL": func(c cloud.Config) string { return c.DeployURL }, + "SPINLOOP_CLOUD_STATS_URL": func(c cloud.Config) string { return c.StatsURL }, + "SPINLOOP_CLOUD_ENV_URL": func(c cloud.Config) string { return c.EnvURL }, + "SPINLOOP_CLOUD_UPDATE_URL": func(c cloud.Config) string { return c.UpdateURL }, + "SPINLOOP_CLOUD_REGION": func(c cloud.Config) string { return c.Region }, } // Unset variables fall through to the file. - cfg, err := resolveRemoteConfig("default", "") + cfg, err := resolveCloudConfig("default", "") if err != nil { - t.Fatalf("resolveRemoteConfig: %v", err) + t.Fatalf("resolveCloudConfig: %v", err) } for name, get := range legs { if got := get(cfg); got != get(file) { @@ -117,7 +117,7 @@ func TestViperRemoteEnvPrecedence(t *testing.T) { // Each exported variable wins over the file, one at a time. for name, get := range legs { t.Setenv(name, envValue) - cfg, err := resolveRemoteConfig("default", "") + cfg, err := resolveCloudConfig("default", "") if err != nil { t.Fatalf("%s set: %v", name, err) } diff --git a/docs/commands/remote.md b/docs/commands/cloud.md similarity index 82% rename from docs/commands/remote.md rename to docs/commands/cloud.md index 9bcb9440..765fc2e0 100644 --- a/docs/commands/remote.md +++ b/docs/commands/cloud.md @@ -1,21 +1,21 @@ -# spinloop remote +# spinloop cloud Run a model too big for your laptop on a GPU in the cloud, from the same [`Spinloop` file](../spinloop-file.md) you'd use locally — and only pay for it while you're using it. ```sh -spinloop remote bootstrap # once per account: deploy the control plane -spinloop remote auth # store or report the credential this machine signs with -spinloop remote bake # bake the runner AMI(s) an environment runs from -spinloop remote deploy # create an endpoint (environment) and tell it what to serve -spinloop remote start # boot it; with --print-env, prints the exports your agent needs +spinloop cloud bootstrap # once per account: deploy the control plane +spinloop cloud auth # store or report the credential this machine signs with +spinloop cloud bake # bake the runner AMI(s) an environment runs from +spinloop cloud deploy # create an endpoint (environment) and tell it what to serve +spinloop cloud start # boot it; with --print-env, prints the exports your agent needs spinloop status --env # is it up? is it healthy? spinloop logs --env # what did it say? (readable after it's gone) -spinloop remote pause # stop it now; a later start re-wakes it -spinloop remote restart # fresh engine, same address: stop it and wake it again -spinloop remote keep 4h # prevent the idle sweep from stopping it for 4 hours -spinloop remote stop # terminate it now, rather than waiting for the idle timer +spinloop cloud pause # stop it now; a later start re-wakes it +spinloop cloud restart # fresh engine, same address: stop it and wake it again +spinloop cloud keep 4h # prevent the idle sweep from stopping it for 4 hours +spinloop cloud stop # terminate it now, rather than waiting for the idle timer ``` The endpoint is the one @@ -27,71 +27,71 @@ enough that the pause is over. ## Bootstrapping the account Before any endpoint can run, the account-level control plane has to -exist — much like `cdk bootstrap`. `spinloop remote bootstrap` does it once per +exist — much like `cdk bootstrap`. `spinloop cloud bootstrap` does it once per account: it downloads the `remote/` CDK project (version-matched to your binary) and deploys the control plane — the EC2 Image Builder pipelines, the lifecycle Lambdas, and the shared weights bucket, roles and VPC — publishing them as -CloudFormation outputs that `spinloop remote deploy` discovers later. It bakes -**no** AMIs — that is the separate `spinloop remote bake` step below. +CloudFormation outputs that `spinloop cloud deploy` discovers later. It bakes +**no** AMIs — that is the separate `spinloop cloud bake` step below. ```sh -spinloop remote bootstrap # shows a consent plan, then deploys -spinloop remote bootstrap --dry-run # print the plan and do nothing -spinloop remote bootstrap --package-manager npm # use npm instead of pnpm +spinloop cloud bootstrap # shows a consent plan, then deploys +spinloop cloud bootstrap --dry-run # print the plan and do nothing +spinloop cloud bootstrap --package-manager npm # use npm instead of pnpm ``` Before deploying, bootstrap prints a plan — the target account and region, the control-plane resources, the cost, and the exact commands — and asks you to confirm (`--yes` skips the prompt). It creates **no** Elastic IP or instance and **no** -environment; those come from `spinloop remote deploy`. Re-running is safe: it +environment; those come from `spinloop cloud deploy`. Re-running is safe: it updates the control-plane stack and doesn't touch any live instance. It needs Node 22, a Node package manager, AWS credentials, and enough GPU vCPU quota for a later launch. By default bootstrap uses `pnpm` and falls back to `npm` when `pnpm` isn't on the path, logging which one it picked. To pin the choice, pass `--package-manager` -(`pnpm` or `npm`) or set `SPINLOOP_REMOTE_PACKAGE_MANAGER`; the flag wins over the +(`pnpm` or `npm`) or set `SPINLOOP_CLOUD_PACKAGE_MANAGER`; the flag wins over the env var. A pinned manager that isn't installed fails the preflight rather than -falling back. `spinloop remote bake` honours the same flags. +falling back. `spinloop cloud bake` honours the same flags. Bootstrap stamps the control plane with the version of the `spinloop` that deploys it, and every control plane response carries that version. When a later `remote` command sees a version different from its own, it prints one warning to stderr -naming both, and carries on; re-run `spinloop remote bootstrap` to bring the +naming both, and carries on; re-run `spinloop cloud bootstrap` to bring the control plane up to date. No warning appears for a control plane deployed before this was added, or when either side is a development build. ## Baking the AMIs Each engine runs from a baked AMI (driver + engine, no model). -`spinloop remote bake` starts a bake for each runner you name — both `llamacpp` +`spinloop cloud bake` starts a bake for each runner you name — both `llamacpp` and `vllm` when you name none — and **waits** until the AMI(s) are available, so -the command returns at the point `spinloop remote deploy` can go: +the command returns at the point `spinloop cloud deploy` can go: ```sh -spinloop remote bake # bake both engines' AMIs; waits (~20-40 min) -spinloop remote bake llamacpp # bake one engine's AMI -spinloop remote bake --no-wait # return once the bakes are queued +spinloop cloud bake # bake both engines' AMIs; waits (~20-40 min) +spinloop cloud bake llamacpp # bake one engine's AMI +spinloop cloud bake --no-wait # return once the bakes are queued ``` Bakes are slow (a builder instance runs for 20–40 minutes) and independent of the weight seed, so `--no-wait` lets them run in parallel — the command prints how to check on them. A bake deploys nothing: it needs the control plane's Image Builder pipelines, so if the control plane isn't deployed it fails telling -you to run `spinloop remote bootstrap` first. Re-bake only when the engine +you to run `spinloop cloud bootstrap` first. Re-bake only when the engine version or the driver changes; the model is **not** baked in, and a new AMI is picked up automatically once it is available. ## The usual flow ```sh -eval "$(spinloop remote start --env qwen3.6-27b --print-env)" # boots it (~10 min +eval "$(spinloop cloud start --env qwen3.6-27b --print-env)" # boots it (~10 min # from cold) and sets OPENAI_BASE_URL and # OPENAI_API_KEY spinloop harness apply --env qwen3.6-27b # point your agent at it spinloop harness open --env qwen3.6-27b # work -spinloop remote pause --env qwen3.6-27b # done for now: stopped, re-wakeable with start -spinloop remote stop --env qwen3.6-27b # done for good: terminate it +spinloop cloud pause --env qwen3.6-27b # done for now: stopped, re-wakeable with start +spinloop cloud stop --env qwen3.6-27b # done for good: terminate it ``` Every command that acts on an endpoint — `start`, `status`, `metrics`, `logs`, @@ -107,7 +107,7 @@ minutes and hours of GPU time. ## Pointing at your endpoint -`spinloop remote` needs the endpoint's control URLs, which its deployment prints. +`spinloop cloud` needs the endpoint's control URLs, which its deployment prints. Put them in a JSON file: ```json @@ -120,9 +120,9 @@ Put them in a JSON file: } ``` -`base_url` is the endpoint's own address, and it's optional — `remote` doesn't +`base_url` is the endpoint's own address, and it's optional — `cloud` doesn't need it, since `start` and `status` report the address themselves. It's there -for [`spinloop harness apply`](harness.md#spinloop-harness-apply): a Spinloop for a remote endpoint can leave out +for [`spinloop harness apply`](harness.md#spinloop-harness-apply): a Spinloop for a cloud endpoint can leave out `BASEURL` and let apply take the address from here, so the address stays with the deployment that owns it. A `BASEURL` in the Spinloop wins if you set one. @@ -133,11 +133,11 @@ spinloop status --env --env qwen3.6-27b-prod ``` The flag selects a **named environment** from the -per-user registry at `~/.config/spinloop/remotes//remote.json`. This keeps +per-user registry at `~/.config/spinloop/clouds//cloud.json`. This keeps deployment state per-user and per-machine: two projects name two environments without clobbering, and the Spinloop carries none of the URLs at all. -`spinloop remote deploy` registers an environment for you; you can also -create one by hand. A name is a plain identifier — `--env ./remote.json` fails, +`spinloop cloud deploy` registers an environment for you; you can also +create one by hand. A name is a plain identifier — `--env ./cloud.json` fails, saying an environment name has no path. **`--env` is required.** There is no environment a command falls back to: half @@ -150,31 +150,31 @@ if you like, and pass `--env default` to use it. A command may also be given a Spinloop path (or a [registered alias](alias.md), or a URL): its `ENV` lines and the `.env` beside it are read before the command signs its AWS calls, so credentials, region and -`SPINLOOP_REMOTE_*` overrides can travel with the Spinloop. The Spinloop does +`SPINLOOP_CLOUD_*` overrides can travel with the Spinloop. The Spinloop does not select the environment — that is the flag's job alone. -`spinloop remote env --env ` fetches a running endpoint's credentials +`spinloop cloud env --env ` fetches a running endpoint's credentials (`export OPENAI_BASE_URL`/`export OPENAI_API_KEY`, safe to `eval`) without booting it. It also reports what is deployed to the environment — runner, served model, context size — whenever something is: this is what lets [`spinloop harness open --env `](harness.md#launching-with-no-spinloop-at-all) configure the harness from a deployed environment with no Spinloop at all. It -appears once the control plane has been redeployed with `spinloop remote +appears once the control plane has been redeployed with `spinloop cloud bootstrap`; an older control plane simply omits it, and `spinloop harness open --env` with no Spinloop fails naming that as the fix. ## Listing environments ```sh -spinloop remote ls +spinloop cloud ls ``` lists each registered environment with its base URL and region, marking any -whose `remote.json` is missing or unreadable. It contacts no endpoint. +whose `cloud.json` is missing or unreadable. It contacts no endpoint. Requests are signed with an AWS credential resolved per region — explicit environment credentials or a named profile first, then the stored -control-plane credential from `spinloop remote auth --store`, then the usual +control-plane credential from `spinloop cloud auth --store`, then the usual chain of config files, SSO sessions, and instance metadata; see [credentials](#credentials). The endpoint's URLs require it. Beyond invoking those URLs, the only extra permission it wants is for @@ -183,7 +183,7 @@ endpoint. ## Credentials -Every `spinloop remote` command signs its requests with an AWS credential, +Every `spinloop cloud` command signs its requests with an AWS credential, resolved for the region the command targets, in this order: 1. Explicit credentials in the process environment (`AWS_ACCESS_KEY_ID` and @@ -193,13 +193,13 @@ resolved for the region the command targets, in this order: it outlives SSO log-ins, which is the point of it. 3. The standard chain: shared config files, SSO sessions, instance metadata. -`spinloop remote auth` manages the stored credential on this machine: +`spinloop cloud auth` manages the stored credential on this machine: ```sh -spinloop remote auth # what is stored (no AWS call) -spinloop remote auth --store # store one for the region; rotates it when stored -spinloop remote auth --store --region ap-southeast-2 -spinloop remote auth --clear # remove it and delete the access key +spinloop cloud auth # what is stored (no AWS call) +spinloop cloud auth --store # store one for the region; rotates it when stored +spinloop cloud auth --store --region ap-southeast-2 +spinloop cloud auth --clear # remove it and delete the access key ``` `--store` creates an access key for the control-plane user the stack makes @@ -207,7 +207,7 @@ spinloop remote auth --clear # remove it and delete the access key Keychain on macOS, Credential Manager on Windows, the Secret Service on Linux. Where no keystore is reachable it keeps it in an owner-only file under the spinloop config directory instead, and every report says which store it used; -`SPINLOOP_REMOTE_KEYSTORE=file` selects the file store even where a keystore +`SPINLOOP_CLOUD_KEYSTORE=file` selects the file store even where a keystore is, for a machine whose keystore is locked or unreachable. The secret is never printed. The key is scoped to day-to-day control only: invoke the control URLs, read the instance logs, discover the stack, price an instance, and manage @@ -227,12 +227,12 @@ If the AWS-side deletion cannot be made, the local entry is still removed and the failure reported. A control plane deployed before this capability has no control-plane user, so -`--store` against it fails naming `spinloop remote bootstrap` — re-run it to +`--store` against it fails naming `spinloop cloud bootstrap` — re-run it to add the user, then store. `bootstrap` and `bake` never consult the stored credential: they provision the control plane itself and run on the administrator's ambient credentials. The -fleet's operations on remote environments sign through the same resolution, so +fleet's operations on cloud environments sign through the same resolution, so a stored key covers them too. ## Checking on an endpoint @@ -276,8 +276,8 @@ did. ## Keeping an instance alive ```sh -spinloop remote keep 4h # retain for 4 hours from now -spinloop remote start --keep 2h # start and retain for 2 hours +spinloop cloud keep 4h # retain for 4 hours from now +spinloop cloud start --keep 2h # start and retain for 2 hours ``` `keep` sets the `Retain-Until` tag on the environment's instance, preventing @@ -300,8 +300,8 @@ version, or re-bootstrap). ## Restarting the engine ```sh -spinloop remote restart # fresh engine, same endpoint -spinloop remote restart --force # skip the graceful engine stop +spinloop cloud restart # fresh engine, same endpoint +spinloop cloud restart --force # skip the graceful engine stop ``` `restart` stops the instance the way `pause` does — without terminating it, so @@ -364,13 +364,13 @@ Reading logs needs one permission beyond the usual endpoint access: rather than reporting an empty log. If it reports that no log group exists, the control plane was deployed before -log shipping existed; `spinloop remote bootstrap` re-deploys it and the next +log shipping existed; `spinloop cloud bootstrap` re-deploys it and the next instance will ship. Logs already lost with a terminated instance can't be recovered — only what's shipped from then on. ## Creating an endpoint: `deploy` -`spinloop remote deploy` creates an **environment** on the bootstrapped control +`spinloop cloud deploy` creates an **environment** on the bootstrapped control plane and tells it what to serve. It reads the Spinloop and its preset — `PROVIDER` picks the engine, so the file that runs a model locally under [`spinloop serve`](serve.md) deploys the same model remotely — and `--env ` @@ -387,12 +387,12 @@ PRESET ./preset.ini # the model and its flags ``` ```sh -spinloop remote deploy --env qwen3.6-27b +spinloop cloud deploy --env qwen3.6-27b ``` Deploy discovers the control plane from the bootstrap stack's outputs, then provisions the environment's own Elastic IP, API key, ingress rule and state, -registers it under `~/.config/spinloop/remotes//`, and stores what to serve. +registers it under `~/.config/spinloop/clouds//`, and stores what to serve. Everything the endpoint sets itself — host, port, where the weights live, the API key, the context size, the alias — is dropped from the preset, so one preset works both locally and remotely without edits. @@ -427,10 +427,10 @@ Switching model, quantisation, or engine is an edit to those two files and one under a different `--env` name gets its own environment, side by side. ```sh -spinloop remote deploy --env qwen3.6-27b --dry-run # see what would be sent -spinloop remote deploy --env other-env path/to/Spinloop # deploy a different file -spinloop remote deploy --env qwen3.6-27b --overwrite # redeploy over the existing environment -spinloop remote deploy --env qwen3.6-27b --reseed # re-fetch weights already in S3 +spinloop cloud deploy --env qwen3.6-27b --dry-run # see what would be sent +spinloop cloud deploy --env other-env path/to/Spinloop # deploy a different file +spinloop cloud deploy --env qwen3.6-27b --overwrite # redeploy over the existing environment +spinloop cloud deploy --env qwen3.6-27b --reseed # re-fetch weights already in S3 ``` Deploy fetches the weights only when they are not in S3 already. `--reseed` @@ -446,10 +446,10 @@ beside the Spinloop) and sends the value: the deploy creates or **rotates** the environment's key, so the old value stops working — and the reply says which happened, never the value itself. -Deploying several environments this way means running `remote deploy` once -per Spinloop file. [`spinloop fleet deploy`](fleet.md#deploying-remote-nodes) +Deploying several environments this way means running `cloud deploy` once +per Spinloop file. [`spinloop fleet deploy`](fleet.md#deploying-cloud-nodes) does the same derivation, consent, and registration (`--api-key-env` -included) for every `kind: remote` node a fleet file names — or a chosen few +included) for every `kind: cloud` node a fleet file names — or a chosen few — in one command, each from its own resolved Spinloop source. ## Flags diff --git a/docs/commands/fleet.md b/docs/commands/fleet.md index bd3025be..82c2a31b 100644 --- a/docs/commands/fleet.md +++ b/docs/commands/fleet.md @@ -13,7 +13,7 @@ spinloop fleet route my-spinloop # which node a harness launch would pick spinloop fleet start gpu-box # start one or more nodes' engines spinloop fleet start --all # start every node in the fleet spinloop fleet stop gpu-box # stop one or more nodes' engines -spinloop fleet deploy --all # create every kind: remote node's AWS environment +spinloop fleet deploy --all # create every kind: cloud node's AWS environment ``` [`spinloop status`](status.md) and [`spinloop dashboard`](dashboard.md) are @@ -93,7 +93,7 @@ nodes: The file is found the way a `Spinloop` is: `./fleet.yaml` in the working directory, or `--fleet `. The full format reference — every field, a node's [Spinloop source](../fleet-file.md#a-nodes-spinloop-source), -[remote environments](../fleet-file.md#remote-environments), +[cloud environments](../fleet-file.md#cloud-environments), [`prefer`](../fleet-file.md#spreading-or-consolidating) and [`wake`](../fleet-file.md#waking), [tags](../fleet-file.md#tags) and [concurrency](../fleet-file.md#concurrency), the [gateway section](../fleet-file.md#gateway), @@ -130,7 +130,7 @@ has been started at all. `spinloop metrics` renders each node's engine and system metrics in the same `gauge` (default), `bar`, `table`, and `json` formats as -[`spinloop metrics --env `](remote.md) — they share the renderers, so a node in +[`spinloop metrics --env `](cloud.md) — they share the renderers, so a node in your fleet and a cloud endpoint look the same. `gauge` draws the current reading per series as a filled progress gauge; `--format=bar` draws each series as a sparkline of the node's daemon's retained history instead, and a @@ -145,7 +145,7 @@ node's engine has done some work. A node whose engine has *stopped* still shows it — the daemon keeps the record across a stop, and "how long since this did anything?" is worth more about a stopped engine than about a busy one. -A `kind: remote` environment carries a relative keep after that figure, on the +A `kind: cloud` environment carries a relative keep after that figure, on the same line — `active 2m 5s ago keep for 2h` — on the same omitted-when-absent terms: it shows how long the idle sweep will hold the box while the deadline is in the future, and is gone once it has passed or was never set. It is the same @@ -193,13 +193,13 @@ spinloop dashboard --fleet f.yaml # another fleet file | `r` | Force a refresh of every node, now | | `g` | Toggle every tile's resource series between bar (sparklines of the retained history) and gauge (the current reading) | | `s` | Start the selected node — without confirmation — shown only for a node that is not running, and only while it has no action in flight | -| `k` | Keep a remote environment for a duration you type — shown only for a node that can be kept, and only while it has no action in flight | +| `k` | Keep a cloud environment for a duration you type — shown only for a node that can be kept, and only while it has no action in flight | | `a` | Abandon a start in flight on the selected node — the wait ends, the node is free again (a stop in flight is not abortable) | | `x` | Stop the selected node — it asks first (`y` sends, `n` or `esc` cancel) — shown only for a node that is running, and only while it has no action in flight | | `q` or `Ctrl+C` | Leave | The board keeps its own cadence: local machines are read every two seconds, -and a [`kind: remote`](../fleet-file.md#remote-environments) environment every 60 — one +and a [`kind: cloud`](../fleet-file.md#cloud-environments) environment every 60 — one status call a minute, because its status is a signed control-plane call, not a local socket, and a cold instance changes state on the scale of minutes. `r` is due for every node whatever those deadlines say. @@ -222,7 +222,7 @@ refresh. A stop in flight is not abortable: it targets an engine already running rather than a cold wake with no deadline of its own, and `a` drives nothing while one is in progress. -`keep` is a remote-environment action: a local daemon has no idle sweep, so +`keep` is a cloud-environment action: a local daemon has no idle sweep, so there is no deadline to set, and the key does not show for one. Pressing it opens a prompt at the foot of the view, pre-filled with `4h`, asking how long the environment should be retained. The prompt is the confirmation — there is @@ -435,30 +435,30 @@ already tells a node what to run when it wakes one. When it does not resolve, daemon already happens to have configured. This is a breaking change: every fleet file with a `kind: daemon` node needs a `file` field, a matching alias, or a matching subdirectory added, or `fleet start` fails for that node. A -`kind: remote` node's `start` is unaffected either way. +`kind: cloud` node's `start` is unaffected either way. -## Deploying remote nodes +## Deploying cloud nodes -`fleet deploy` creates the AWS environment for one or more `kind: remote` +`fleet deploy` creates the AWS environment for one or more `kind: cloud` nodes — the step that otherwise has to happen outside the fleet file -entirely, one `spinloop remote deploy --env ` at a time, run from the +entirely, one `spinloop cloud deploy --env ` at a time, run from the directory holding each node's Spinloop: ```sh spinloop fleet deploy qwen # one node spinloop fleet deploy qwen llama # several -spinloop fleet deploy --all # every kind: remote node in the file +spinloop fleet deploy --all # every kind: cloud node in the file ``` Each node deploys from its own resolved [Spinloop source](../fleet-file.md#a-nodes-spinloop-source), reusing the exact derivation, consent, and -registration `spinloop remote deploy` uses for the same file — the two can +registration `spinloop cloud deploy` uses for the same file — the two can never disagree about what a given Spinloop deploys — and the environment each node creates is named after the node itself. A `kind: daemon` node named explicitly fails the command, explaining that `deploy` provisions cloud environments and that node is not one; `--all` only ever selects `kind: -remote` nodes, so a daemon node is never swept in by it. As with -`start`/`stop`, no node and no `--all` lists the fleet's `kind: remote` nodes +cloud` nodes, so a daemon node is never swept in by it. As with +`start`/`stop`, no node and no `--all` lists the fleet's `kind: cloud` nodes and deploys nothing, `--all` plus node names is refused as ambiguous, and several targeted nodes deploy independently — one node's guard or failure is reported against it alone. @@ -469,15 +469,15 @@ spinloop fleet deploy qwen --overwrite # redeploy over a registered environme ``` `--dry-run`, `--overwrite`, `--reseed`, `--allowed-cidr`, `--region`, and -`--spinloop-version` mean exactly what they mean on [`spinloop remote -deploy`](remote.md), applied per node. +`--spinloop-version` mean exactly what they mean on [`spinloop cloud +deploy`](cloud.md), applied per node. ## Flags | Flag | Meaning | | ---- | ------- | | `-f`, `--fleet ` | The fleet file (default `./fleet.yaml`) — `logs` takes it long-form only, since `-f` is its follow flag | -| `--all` | `start`/`stop`/`deploy`: act on every node (or every `kind: remote` node, for `deploy`) instead of named ones | +| `--all` | `start`/`stop`/`deploy`: act on every node (or every `kind: cloud` node, for `deploy`) instead of named ones | | `--node ` | `route` only: report this node rather than choosing one | | `--prefer` | `route` only: rank by `idle` or `active`, overriding the file | | `--format` | `metrics`: `gauge` (default), `bar`, `table`, or `json`; `logs`: `text` (default) or `json` | diff --git a/docs/commands/gateway.md b/docs/commands/gateway.md index d363f619..37a4faaa 100644 --- a/docs/commands/gateway.md +++ b/docs/commands/gateway.md @@ -73,7 +73,7 @@ picker and a second gateway does not overwrite this one. See | Path | Meaning | | ---- | ------- | | `GET /health` | That the gateway is up. It touches no node on purpose — it is how you tell the gateway down from the fleet down. | -| `GET /v1/models` | The OpenAI list of what a request can reach: what the running nodes report (the served name when a node reports one, else the model id), and, for a stopped node [waking can reach](#waking-a-node), the model it would start with — its own Spinloop source for a `kind: daemon` node, its own stats reply for a `kind: remote` one. Duplicates once. Nothing reachable is an empty list, not an error. | +| `GET /v1/models` | The OpenAI list of what a request can reach: what the running nodes report (the served name when a node reports one, else the model id), and, for a stopped node [waking can reach](#waking-a-node), the model it would start with — its own Spinloop source for a `kind: daemon` node, its own stats reply for a `kind: cloud` one. Duplicates once. Nothing reachable is an empty list, not an error. | | `POST /v1/chat/completions` | Routed to the node serving the request's `model`, the way a launch routes. | | `POST /v1/completions` | The same, for the completions endpoint. | | `GET /v1/fleet` | The fleet's [topology](#the-fleets-topology) — what a [`spinloop orchestrator`](orchestrator.md) reads to work its backlog. | @@ -83,7 +83,7 @@ not serve is refused with a `404` naming the ones it does. The list is what a request can reach, so it is bounded by what the gateway can start: a running node contributes only what it reports — a running engine is -never displaced to make room. A deployed-but-stopped `kind: remote` +never displaced to make room. A deployed-but-stopped `kind: cloud` environment contributes the model id its own stats reply reports — read directly from its stored deploy config, the way `spinloop metrics --env ` already reads it, since its status reply carries no such facts while @@ -106,7 +106,7 @@ The request's body goes out unmodified and streamed replies are flushed as they are produced, so a `stream: true` request streams through. The caller's authorisation never travels past the gateway: the engine is reached with the key its fleet entry names (`engineTokenEnv`, or the fleet-wide `apiKeyEnv` for -a `kind: remote` node), and an ungated engine is reached with none. The reply +a `kind: cloud` node), and an ungated engine is reached with none. The reply the engine gives is the reply the caller gets — the gateway never retries another node, and an upstream failure reaches the caller as an error naming the node. @@ -140,21 +140,21 @@ one whose stored config already matches is tried first: - A **`kind: daemon`** node is started with the config its own Spinloop source resolves to. -- A **deployed-but-stopped `kind: remote`** node is booted as it is: its own - stored deploy config — set by `spinloop remote deploy`, not by this wake — +- A **deployed-but-stopped `kind: cloud`** node is booted as it is: its own + stored deploy config — set by `spinloop cloud deploy`, not by this wake — decides what it serves, and the gateway pushes nothing new. An **undeployed** environment is never a candidate: it has nothing to serve - yet, and choosing what to deploy is `spinloop remote deploy`'s call, not a + yet, and choosing what to deploy is `spinloop cloud deploy`'s call, not a request's. The wait is bounded by `--wake-timeout` (default 5m); a timeout fails the request saying so and leaves the engine running, so a slow load — or, for a -remote node, a slow boot — is not thrown away. Concurrent requests for the +cloud node, a slow boot — is not thrown away. Concurrent requests for the same model wake at most one engine: the gateway coalesces two requests racing to wake the same node into a single start, so a request that arrives mid-wake joins the one already under way rather than starting a second engine of its own — a daemon node's control API would refuse the second -start anyway, but a remote environment's control plane does not, so this is +start anyway, but a cloud environment's control plane does not, so this is what keeps a burst of requests from booting (and billing for) more than one instance. diff --git a/docs/commands/harness.md b/docs/commands/harness.md index 13d39b96..e012f927 100644 --- a/docs/commands/harness.md +++ b/docs/commands/harness.md @@ -78,7 +78,7 @@ variable chooses which Spinloop, never whether you are configured. See `--env ` on its own — no leading alias or path, no `--spinloop`/`-O` — configures the harness from what is actually deployed to that -[environment](remote.md), rather than doing nothing with the flag: +[environment](cloud.md), rather than doing nothing with the flag: ```sh spinloop harness open --env dev-3 --prompt "..." # configured from dev-3's deployment, then launched @@ -89,8 +89,8 @@ the model, and its context size (when set) becomes the context window — the same result a Spinloop stating the matching `PROVIDER`/`ALIAS`/`CONTEXT` with `--env dev-3` would produce, without writing one. This is what makes the two-machine flow work: deploy from one machine -(`spinloop remote deploy --env dev-3`), then on any machine that -can reach the same environment — one that has its `remote.json` in the +(`spinloop cloud deploy --env dev-3`), then on any machine that +can reach the same environment — one that has its `cloud.json` in the registry, however it got there — run `spinloop harness open --env dev-3` with no Spinloop and get the same configuration, live, so a later redeploy is picked up automatically rather than requiring anyone to re-copy anything. @@ -102,8 +102,8 @@ registered address. A bare `--env` against an environment with nothing deployed — or a control plane too old to report what is deployed — fails before launching, naming the -environment and how to fix it: `spinloop remote deploy --env -` to deploy something, or `spinloop remote bootstrap` to update the +environment and how to fix it: `spinloop cloud deploy --env +` to deploy something, or `spinloop cloud bootstrap` to update the control plane. ### Flags @@ -112,7 +112,7 @@ control plane. | ---- | ------- | | `-H`, `--harness` | Which harness to launch (or set `SPINLOOP_HARNESS`) | | `-O`, `--spinloop` | Apply this Spinloop before launching (bare: `./Spinloop`) | -| `-e`, `--env` | The registered [environment](remote.md) to launch against; with no Spinloop applied, configures the harness from what is deployed there — mutually exclusive with fleet routing, since each names where the model is served from | +| `-e`, `--env` | The registered [environment](cloud.md) to launch against; with no Spinloop applied, configures the harness from what is deployed there — mutually exclusive with fleet routing, since each names where the model is served from | | `--providers` | Path to a custom catalogue, for the applied Spinloop | | `-f`, `--fleet` | Route through this fleet file (default: `./fleet.yaml`, when the Spinloop is not named) | | `--node` | Pin the launch to one fleet node | @@ -138,7 +138,7 @@ spinloop harness open --prefer active -f fleet.yaml spinloop queries the fleet, prefers a node already serving the Spinloop's model, and points the launched agent at that node's engine — the same injection that -carries a [remote environment](remote.md)'s endpoint address and key, with a +carries a [cloud environment](cloud.md)'s endpoint address and key, with a selection step in front. It reports which node it chose, and why, before the agent starts. @@ -303,7 +303,7 @@ opencode. After applying, just launch your agent — or do both at once with | Flag | Meaning | | ---- | ------- | -| `-e`, `--env` | The registered [environment](remote.md) the Spinloop points at: names the harness provider and, with no `BASEURL`, supplies the endpoint's address | +| `-e`, `--env` | The registered [environment](cloud.md) the Spinloop points at: names the harness provider and, with no `BASEURL`, supplies the endpoint's address | | `-o`, `--output` | Max output tokens — overrides the Spinloop's `OUTPUT` | | `-H`, `--harness` | Which harness to configure (or set `SPINLOOP_HARNESS`) | | `--providers` | Path to a custom catalogue (a Spinloop never names one) | @@ -318,13 +318,13 @@ Notes: - A Spinloop's `PRESET` line is for [`spinloop serve`](serve.md); `apply` ignores it — never fetched, even when it's a URL. - With `--env `, the Spinloop points at a registered - [environment](remote.md): the harness provider is keyed on the environment + [environment](cloud.md): the harness provider is keyed on the environment name (so several environments built from the same engine keep their own entries), and a Spinloop with no `BASEURL` takes the endpoint's address from - the environment's `remote.json` `base_url`, which its deployment writes. A + the environment's `cloud.json` `base_url`, which its deployment writes. A `BASEURL` in the Spinloop wins over it. An unregistered name fails, naming - the `spinloop remote deploy --env ` that would create it. Without - `--env`, apply reads no remote config at all. + the `spinloop cloud deploy --env ` that would create it. Without + `--env`, apply reads no cloud config at all. ## spinloop harness unapply diff --git a/docs/commands/index.md b/docs/commands/index.md index 336b7127..37072034 100644 --- a/docs/commands/index.md +++ b/docs/commands/index.md @@ -19,7 +19,7 @@ help` the usage summary. | [`spinloop gateway`](gateway.md) | Serve the fleet under one OpenAI-compatible endpoint | | [`spinloop orchestrator`](orchestrator.md) | Work a backlog of items against the fleet, at the fleet's declared pace | | [`spinloop work`](work.md) | Drive the orchestrator's work list from the shell: add, list, logs, abort, remove — or watch it live on `board` | -| [`spinloop remote`](remote.md) | Run the model on a cloud GPU that stops when you do | +| [`spinloop cloud`](cloud.md) | Run the model on a cloud GPU that stops when you do | | [`spinloop hf`](hf.md) | Write a `Spinloop` for a Hugging Face model, from its page reference | | [`spinloop alias`](alias.md) | Name a `Spinloop` so the name works anywhere a path does | | [`spinloop unalias`](unalias.md) | Drop a registered name | diff --git a/docs/commands/serve.md b/docs/commands/serve.md index fdc73296..bd1aa66f 100644 --- a/docs/commands/serve.md +++ b/docs/commands/serve.md @@ -68,7 +68,7 @@ to stderr there. `--dry-run` never opens the view. ## The engine comes from `PROVIDER` `PROVIDER` already names the engine, so `serve` needs no keyword of its own — -the same way [`spinloop remote deploy`](remote.md) picks the engine for a cloud +the same way [`spinloop cloud deploy`](cloud.md) picks the engine for a cloud GPU: | `PROVIDER` | `serve` runs | diff --git a/docs/env-vars.md b/docs/env-vars.md index 7cfdd4ce..5f9fd85b 100644 --- a/docs/env-vars.md +++ b/docs/env-vars.md @@ -8,39 +8,39 @@ from the environment or a `.env` beside the Spinloop — never written into an | Variable | Used by | Meaning | | --- | --- | --- | -| `SPINLOOP_CONFIG_DIR` | everything | spinloop's config directory, used **verbatim** (no `spinloop` segment appended). Overrides `XDG_CONFIG_HOME` and `~/.config`. Everything spinloop owns lives here: `config.json` (default-harness preference + alias registry), `remote.json`, the `remotes//` environment registry, the `keystore/` file credential store, the daemon state dir, and the CDK source cache. Set it when there is no usable `$HOME` — e.g. a systemd service. See [config resolution](#config-directory-resolution). | +| `SPINLOOP_CONFIG_DIR` | everything | spinloop's config directory, used **verbatim** (no `spinloop` segment appended). Overrides `XDG_CONFIG_HOME` and `~/.config`. Everything spinloop owns lives here: `config.json` (default-harness preference + alias registry), `cloud.json`, the `clouds//` environment registry, the `keystore/` file credential store, the daemon state dir, and the CDK source cache. Set it when there is no usable `$HOME` — e.g. a systemd service. See [config resolution](#config-directory-resolution). | | `SPINLOOP_HARNESS` | all harness commands | Which harness to configure/launch (`opencode`, `pi` or `lucinate`). Precedence: `--harness`/`-H` flag > `SPINLOOP_HARNESS` > stored preference > `opencode`. | | `SPINLOOP_ALIAS` | every command that takes a Spinloop path | A name registered with [`spinloop alias`](commands/alias.md), used when the command is given no path. Precedence: the path or alias argument > `SPINLOOP_ALIAS` > `./Spinloop`. It holds a registry name, never a path, and a same-named file in the working directory does not shadow it. It decides *which* Spinloop is the default, not *whether* one is applied — a bare `spinloop harness open` still applies nothing, and `spinloop alias` ignores it. | | `SPINLOOP_PROVIDERS` | `spinloop provider list`, `spinloop harness add`, `spinloop harness apply`, … | Path to a `providers.yaml` that overrides the built-in catalogue. Precedence: `--providers` flag > `SPINLOOP_PROVIDERS` > embedded. | | `SPINLOOP_BASE_URL` | `spinloop harness add`, `spinloop harness apply` | Base-URL override for the provider being configured. Precedence: `--base-url`/`-u` > `SPINLOOP_BASE_URL` > the provider's own option var > the catalogue default. | | `SPINLOOP_API_TOKEN` | `spinloop daemon`, `spinloop serve --api`, `spinloop gateway` | Bearer token for the daemon control API — and the token a [gateway](commands/gateway.md)'s callers must present. One of three peer sources, alongside `--api-token-file` and `--api-token`; two at once is an error. From a service manager prefer the file form — see [serve](commands/serve.md). A non-loopback listen without any of them refuses to start. | -| `SPINLOOP_REMOTE_KEYSTORE` | `spinloop remote auth` | Set to `file` to keep the stored control-plane credential in the owner-only file under the config directory, even where an OS keystore is reachable — the opt-out for a machine whose keystore is locked or unreachable. Unset, the OS keystore is used where available. See [credentials](commands/remote.md#credentials). | +| `SPINLOOP_CLOUD_KEYSTORE` | `spinloop cloud auth` | Set to `file` to keep the stored control-plane credential in the owner-only file under the config directory, even where an OS keystore is reachable — the opt-out for a machine whose keystore is locked or unreachable. Unset, the OS keystore is used where available. See [credentials](commands/cloud.md#credentials). | | `SPINLOOP_LOG_LEVEL` | `spinloop daemon`, `spinloop serve` | How much spinloop records about the control API and the supervised engine: `debug`, `info` (default), `warn` or `error`. Precedence: `--log-level` flag > `SPINLOOP_LOG_LEVEL` > `info`. An unrecognised value refuses to start rather than falling back to the default. Under `spinloop serve` the `.env` beside the Spinloop can set it; the daemon reads no Spinloop, so there it comes from the environment its service manager gives it. Records go to stderr; see [what gets logged](commands/serve.md#what-gets-logged). | | *(per-node, named by `tokenEnv`)* | `spinloop fleet` | A fleet node's bearer token. `fleet.yaml` names the variable rather than holding the value; it resolves from the environment, then the `.env` beside the fleet file. See [the fleet file](fleet-file.md#tokens). | | *(per-node, named by `engineTokenEnv`)* | `spinloop fleet`, `spinloop harness open` | The key a fleet node's **engine** is gated with. Resolved the same way, and supplied by the client when it starts that engine — so the node holds no key of its own and the two ends cannot disagree. See [the fleet file](fleet-file.md#tokens). | -| *(fleet-wide, named by `apiKeyEnv`)* | `spinloop fleet`, `spinloop harness open` | The default key for a `kind: remote` environment's engine, for every remote node that does not name its own `engineTokenEnv`. Resolved the same way. See [the fleet file](fleet-file.md#tokens). | +| *(fleet-wide, named by `apiKeyEnv`)* | `spinloop fleet`, `spinloop harness open` | The default key for a `kind: cloud` environment's engine, for every cloud node that does not name its own `engineTokenEnv`. Resolved the same way. See [the fleet file](fleet-file.md#tokens). | -## Remote (`spinloop remote`) +## Cloud (`spinloop cloud`) | Variable | Meaning | | --- | --- | -| `SPINLOOP_REMOTE_START_URL` | Override the start Lambda Function URL from the remote config. | -| `SPINLOOP_REMOTE_STOP_URL` | Override the stop Lambda Function URL. | -| `SPINLOOP_REMOTE_DEPLOY_URL` | Override the deploy Lambda Function URL. | -| `SPINLOOP_REMOTE_STATS_URL` | Override the stats Lambda Function URL. | -| `SPINLOOP_REMOTE_ENV_URL` | Override the env Lambda Function URL. | -| `SPINLOOP_REMOTE_UPDATE_URL` | Override the update Lambda Function URL (drives `keep`). | -| `SPINLOOP_REMOTE_REGION` | Override the AWS region (else `AWS_REGION`, else the region in the Function URL host). | -| `SPINLOOP_REMOTE_PACKAGE_MANAGER` | Pin the package manager (`pnpm`/`npm`) `spinloop remote bootstrap` and `bake` use. | +| `SPINLOOP_CLOUD_START_URL` | Override the start Lambda Function URL from the cloud config. | +| `SPINLOOP_CLOUD_STOP_URL` | Override the stop Lambda Function URL. | +| `SPINLOOP_CLOUD_DEPLOY_URL` | Override the deploy Lambda Function URL. | +| `SPINLOOP_CLOUD_STATS_URL` | Override the stats Lambda Function URL. | +| `SPINLOOP_CLOUD_ENV_URL` | Override the env Lambda Function URL. | +| `SPINLOOP_CLOUD_UPDATE_URL` | Override the update Lambda Function URL (drives `keep`). | +| `SPINLOOP_CLOUD_REGION` | Override the AWS region (else `AWS_REGION`, else the region in the Function URL host). | +| `SPINLOOP_CLOUD_PACKAGE_MANAGER` | Pin the package manager (`pnpm`/`npm`) `spinloop cloud bootstrap` and `bake` use. | -These let the remote commands run without a `remote.json` on disk — the config +These let the cloud commands run without a `cloud.json` on disk — the config can come entirely from the environment. `--env ` is still required, and on this path the name you give *is* the environment identifier the control plane acts on, since there is no file to take one from: ```sh -SPINLOOP_REMOTE_START_URL=... SPINLOOP_REMOTE_STOP_URL=... SPINLOOP_REMOTE_REGION=... \ - spinloop remote start --env ci +SPINLOOP_CLOUD_START_URL=... SPINLOOP_CLOUD_STOP_URL=... SPINLOOP_CLOUD_REGION=... \ + spinloop cloud start --env ci ``` ## Standard variables spinloop honours @@ -48,8 +48,8 @@ SPINLOOP_REMOTE_START_URL=... SPINLOOP_REMOTE_STOP_URL=... SPINLOOP_REMOTE_REGIO | Variable | Meaning | | --- | --- | | `XDG_CONFIG_HOME` | Base for spinloop's config dir (`$XDG_CONFIG_HOME/spinloop`) when `SPINLOOP_CONFIG_DIR` is unset. | -| `AWS_REGION` | AWS region for the remote control calls when the remote config names none. | -| `HF_TOKEN` | Hugging Face token. Read by `spinloop hf` (sent as a bearer for gated or private repos) and used to seed gated model weights during `spinloop remote deploy`. Precedence in `hf`: `HF_TOKEN` > `HUGGING_FACE_HUB_TOKEN` > the token file. | +| `AWS_REGION` | AWS region for the cloud control calls when the cloud config names none. | +| `HF_TOKEN` | Hugging Face token. Read by `spinloop hf` (sent as a bearer for gated or private repos) and used to seed gated model weights during `spinloop cloud deploy`. Precedence in `hf`: `HF_TOKEN` > `HUGGING_FACE_HUB_TOKEN` > the token file. | | `HUGGING_FACE_HUB_TOKEN` | Hugging Face token, the second of the two `spinloop hf` reads, after `HF_TOKEN`. | | `HF_HOME` | Base of the Hugging Face home: its `token` file (third in `hf`'s token order) and, when `HF_HUB_CACHE` is unset, the hub cache lives at `$HF_HOME/hub`. | | `HF_HUB_CACHE` | The Hugging Face hub cache `spinloop hf` checks for a copy already on disk, before the hub. Default `$HF_HOME/hub`, else `~/.cache/huggingface/hub`. | diff --git a/docs/faq.md b/docs/faq.md index 63d98fc2..398fd0ed 100644 --- a/docs/faq.md +++ b/docs/faq.md @@ -39,7 +39,7 @@ is for: one per project, applied per project. Register each under a short name with `spinloop alias` and the names work anywhere a path does; `SPINLOOP_ALIAS` names one for the whole shell. -**What does `spinloop remote` cost?** It runs in your own AWS account. The +**What does `spinloop cloud` cost?** It runs in your own AWS account. The control plane and baked AMIs are one-time; a GPU instance bills only while it is running, and an idle endpoint is stopped (no GPU billing) and later terminated by its retention window. You can also read diff --git a/docs/fleet-file.md b/docs/fleet-file.md index 4c91c6ff..bd023965 100644 --- a/docs/fleet-file.md +++ b/docs/fleet-file.md @@ -13,7 +13,7 @@ written in the file. # fleet.yaml — the machines, and how to reach each prefer: idle # optional; idle (default) or active wake: on # optional; on (default) or off -apiKeyEnv: REMOTE_KEY # optional; the default engine key for kind: remote nodes +apiKeyEnv: REMOTE_KEY # optional; the default engine key for kind: cloud nodes nodes: - name: studio # required, unique; what you type at `fleet start ` @@ -27,8 +27,8 @@ nodes: port: 18080 path: /v1 - - name: qwen # for a kind: remote node, the registered environment's name - kind: remote + - name: qwen # for a kind: cloud node, the registered environment's name + kind: cloud instance-type: g6e.2xlarge # optional; the EC2 type its environment launches as gateway: # optional; where a spinloop gateway serves this fleet @@ -53,22 +53,22 @@ and how to create one. | `prefer` | no | `idle` (default) or `active` — how routing ranks several nodes that could all serve; see [Spreading or consolidating](#spreading-or-consolidating) | | `wake` | no | `on` (default) or `off` — whether routing may start an engine on a node that is not running one; see [Waking](#waking) | | `concurrency` | no | The most work the fleet may have in flight at once, for the [`spinloop orchestrator`](commands/orchestrator.md); see [Concurrency](#concurrency) | -| `apiKeyEnv` | no | The variable holding the key the fleet's `kind: remote` nodes share; see [Tokens](#tokens) | +| `apiKeyEnv` | no | The variable holding the key the fleet's `kind: cloud` nodes share; see [Tokens](#tokens) | | `gateway` | no | The address a [`spinloop gateway`](commands/gateway.md) serves the fleet under; see [Gateway](#gateway) | ### A node | Field | Required? | Meaning | | ---------------- | -------------------- | ------------------------------------------------------------------------------------------------------------------------------ | -| `name` | yes | Unique in the file; what you type at `fleet start `. For a `kind: remote` node, the registered environment it drives | +| `name` | yes | Unique in the file; what you type at `fleet start `. For a `kind: cloud` node, the registered environment it drives | | `host` | for `kind: daemon` | Where the daemon answers — a LAN name, a tailscale name, or an address | | `port` | no | The daemon's control API port; 4242 when omitted | -| `kind` | no | `daemon` (default) or `remote`; see [Remote environments](#remote-environments) | +| `kind` | no | `daemon` (default) or `cloud`; see [Cloud environments](#cloud-environments) | | `tokenEnv` | no | The variable holding this daemon's bearer token; none means no authentication (a loopback-only daemon) | | `engineTokenEnv` | no | The variable holding the key this node's *engine* is gated with; see [Tokens](#tokens) | | `engine` | no | An override of where the engine serves — `host`, `port`, `path`, each optional; see [Where a node's engine answers](#where-a-nodes-engine-answers) | | `file` | no | The [Spinloop](spinloop-file.md) that describes what this node runs; see [A node's Spinloop source](#a-nodes-spinloop-source) | -| `instance-type` | no, `kind: remote` only | The EC2 instance type the environment launches as, e.g. `g6e.xlarge`; see [Remote environments](#remote-environments) | +| `instance-type` | no, `kind: cloud` only | The EC2 instance type the environment launches as, e.g. `g6e.xlarge`; see [Cloud environments](#cloud-environments) | | `tags` | no | Key/value pairs naming the kind of work the node takes on; only the [`spinloop orchestrator`](commands/orchestrator.md) reads them; see [Tags](#tags) | ### The `gateway` section @@ -108,10 +108,10 @@ elsewhere fails with that explanation rather than a bare connection refused — bind the engine to a reachable address (llama.cpp's `--host 0.0.0.0`), or declare an `engine` block, which is you taking responsibility for reachability. -## Remote environments +## Cloud environments `kind` (defaulted to `daemon`) says how the fleet reaches a node. A node can -also be an [`spinloop remote`](commands/remote.md) environment rather than a +also be an [`spinloop cloud`](commands/cloud.md) environment rather than a machine: its `name` is the registered environment it drives — no `host` needed — and it is reached through its control plane, which signs each call with your AWS credentials, so it needs no bearer token: @@ -119,27 +119,27 @@ AWS credentials, so it needs no bearer token: ```yaml nodes: - name: qwen # the registered environment, and what you type at `fleet start ` - kind: remote + kind: cloud ``` -The environment's control URLs live in its `remote.json` (under -`~/.config/spinloop/remotes//`), written by `spinloop remote deploy` — or -by [`spinloop fleet deploy`](commands/fleet.md#deploying-remote-nodes), which +The environment's control URLs live in its `cloud.json` (under +`~/.config/spinloop/clouds//`), written by `spinloop cloud deploy` — or +by [`spinloop fleet deploy`](commands/fleet.md#deploying-cloud-nodes), which creates it from the fleet file itself — and never stored in the fleet file. So a daemon and an environment sit side by side as the same kind of row, and an environment that has not been deployed yet shows as `config-error` on its row rather than blanking the fleet. See -[`examples/fleet-remote`](https://github.com/spinloop-ai/spinloop/blob/main/examples/fleet-remote/README.md) and +[`examples/fleet-cloud`](https://github.com/spinloop-ai/spinloop/blob/main/examples/fleet-cloud/README.md) and [`examples/fleet-mixed`](https://github.com/spinloop-ai/spinloop/blob/main/examples/fleet-mixed/README.md). -A `kind: remote` node may also name the EC2 instance type its environment +A `kind: cloud` node may also name the EC2 instance type its environment launches as, with `instance-type` (a family and size separated by a dot, e.g. `g6e.xlarge`): ```yaml nodes: - name: qwen - kind: remote + kind: cloud instance-type: g6e.2xlarge ``` @@ -148,14 +148,14 @@ It is a property of the cloud environment, not of the fleet's view of it: **fresh** launch uses it. A re-wake of a stopped instance keeps the type it was launched with — EC2 cannot resize a running or stopped box — so a changed value takes effect only after the instance is terminated (an idle sweep or -`spinloop remote stop`) and launched again. Omitted, the environment launches as +`spinloop cloud stop`) and launched again. Omitted, the environment launches as its control plane's default type. Naming `instance-type` on a `kind: daemon` node is a configuration error: a daemon's hardware is the operator's to choose, not the fleet file's. ## A node's Spinloop source -Both `fleet deploy` (for a `kind: remote` node's environment) and `fleet start` +Both `fleet deploy` (for a `kind: cloud` node's environment) and `fleet start` (for a `kind: daemon` node's engine) need to know what Spinloop file describes what a node runs. A node names it with `file`, resolved relative to the fleet file: @@ -163,7 +163,7 @@ file: ```yaml nodes: - name: qwen - kind: remote + kind: cloud file: ./envs/qwen.Spinloop ``` @@ -171,7 +171,7 @@ nodes: key. When it is absent, resolution tries, in order: 1. `name` registered as a `spinloop alias` (`spinloop alias add qwen - ./envs/qwen.Spinloop`) — the same lookup `spinloop remote deploy` performs + ./envs/qwen.Spinloop`) — the same lookup `spinloop cloud deploy` performs for a Spinloop argument; 2. a subdirectory named after the node, beside the fleet file — `qwen/Spinloop` next to `fleet.yaml` for a node named `qwen`, no fields needed on either side. @@ -189,7 +189,7 @@ Nothing resolving is a per-node error naming all three ways a source could have been given. For `fleet deploy` that always fails the node (there is nothing to create an environment from); for `fleet start` on a `kind: daemon` node it likewise fails that node's start — there is no fallback to a plain, config-less -start once this field exists. A `kind: remote` node's `start` is unaffected by +start once this field exists. A `kind: cloud` node's `start` is unaffected by any of this: what it serves is fixed at deploy time, not pushed at start time. This does not apply to `spinloop dashboard`'s `s` key, which still starts @@ -248,13 +248,13 @@ nodes: - name: gpu-box host: 198.51.100.7 - name: prod - kind: remote + kind: cloud wake: off # this one node stays asleep even though the fleet wakes ``` -This matters most for a `kind: remote` node, whose wake boots a billed cloud +This matters most for a `kind: cloud` node, whose wake boots a billed cloud instance rather than starting a process on a machine you already run — so you -can leave the fleet's daemons on `wake: on` while deciding a given remote +can leave the fleet's daemons on `wake: on` while deciding a given cloud environment's waking separately, in either direction: `wake: off` on one node under a fleet that otherwise wakes, or `wake: on` on one node under a fleet that otherwise does not. @@ -335,7 +335,7 @@ When a launch through the gateway has no model of its own to route by (see way to label the provider it configures — otherwise every gateway a fleet might name would collide under the same generic id. `name` supplies that label directly; with none given, the section's address's host stands in (e.g. -`localhost:4000`). Either way opencode and Pi show it the way a remote +`localhost:4000`). Either way opencode and Pi show it the way a cloud environment is shown — `Gateway (remote-llms)` rather than a bare `OpenAI-compatible`, the same pattern as `llama.cpp (dev-2)`. @@ -362,10 +362,10 @@ There are three references, all resolved the same way: authorises using its engine — and a node may need either, both, or neither. The daemon never hands its engine's key out: it says only that one is required. -- **`apiKeyEnv`** (top level) — the default key for every `kind: remote` node. - A remote environment is always keyed, so a node that names no +- **`apiKeyEnv`** (top level) — the default key for every `kind: cloud` node. + A cloud environment is always keyed, so a node that names no `engineTokenEnv` of its own takes the fleet-wide reference. A node's own - `engineTokenEnv` overrides it, so one remote may carry a distinct key while + `engineTokenEnv` overrides it, so one cloud node may carry a distinct key while the rest of the fleet shares one. A `kind: daemon` node never takes the fleet-wide reference: it is gated only by its own `engineTokenEnv`. @@ -376,15 +376,15 @@ There are three references, all resolved the same way: engineTokenEnv: GATED_ENGINE_KEY # to talk to its engine ``` -A `kind: remote` environment is always keyed, so it needs an engine key too — +A `kind: cloud` environment is always keyed, so it needs an engine key too — its `engineTokenEnv` works as above, and a fleet-wide `apiKeyEnv` is the -default for every remote node that does not name one of its own: +default for every cloud node that does not name one of its own: ```yaml -apiKeyEnv: REMOTE_ENGINE_KEY # the default for every kind: remote node +apiKeyEnv: REMOTE_ENGINE_KEY # the default for every kind: cloud node nodes: - name: qwen - kind: remote + kind: cloud engineTokenEnv: OTHER_KEY # overrides it for this node ``` @@ -402,8 +402,8 @@ the reason named: - no `nodes` — list at least one; - a node with no `name`, or two nodes with the same name; - a `kind: daemon` node with no `host`; -- an unknown `kind` — only `daemon` and `remote` are supported; -- a `kind: remote` node whose `name` is not shaped like a registered +- an unknown `kind` — only `daemon` and `cloud` are supported; +- a `kind: cloud` node whose `name` is not shaped like a registered environment name (no `/`, no trailing `.json`) — the name is the key of the environment it drives; - an `instance-type` that is not shaped like an EC2 instance type (a family and @@ -424,7 +424,7 @@ Fleet files, each with a walkthrough: - [`examples/fleet-local/`](https://github.com/spinloop-ai/spinloop/tree/main/examples/fleet-local) — a fleet of one, on your own machine - [`examples/fleet/`](https://github.com/spinloop-ai/spinloop/tree/main/examples/fleet) — a small LAN fleet, all defaults - [`examples/fleet-docker/`](https://github.com/spinloop-ai/spinloop/tree/main/examples/fleet-docker) — a runnable multi-node fleet in containers -- [`examples/fleet-remote/`](https://github.com/spinloop-ai/spinloop/tree/main/examples/fleet-remote) — a fleet of cloud environments +- [`examples/fleet-cloud/`](https://github.com/spinloop-ai/spinloop/tree/main/examples/fleet-cloud) — a fleet of cloud environments - [`examples/fleet-mixed/`](https://github.com/spinloop-ai/spinloop/tree/main/examples/fleet-mixed) — daemons and environments side by side - [`examples/gateway-docker/`](https://github.com/spinloop-ai/spinloop/tree/main/examples/gateway-docker) — a fleet behind its gateway diff --git a/docs/getting-started.md b/docs/getting-started.md index 1a461351..516f318f 100644 --- a/docs/getting-started.md +++ b/docs/getting-started.md @@ -113,7 +113,7 @@ no Spinloop uses it — see [`spinloop alias`](commands/alias.md#naming-one-for- - [From a Hugging Face model](guides/hugging-face.md) — a model page's reference, into a `Spinloop` - [Run a daemon node](guides/daemon.md) — the engine under the control API -- [Deploy to a cloud GPU](guides/remote.md) — the same `Spinloop`, on a +- [Deploy to a cloud GPU](guides/cloud.md) — the same `Spinloop`, on a machine that stops when you do - [Run a fleet](guides/fleet.md) — every machine you run, from one place - [The `Spinloop` file](spinloop-file.md) — full syntax diff --git a/docs/guides/remote.md b/docs/guides/cloud.md similarity index 71% rename from docs/guides/remote.md rename to docs/guides/cloud.md index 9a4ca5f7..27a20c76 100644 --- a/docs/guides/remote.md +++ b/docs/guides/cloud.md @@ -12,8 +12,8 @@ difference between minutes and hours of GPU time. Two steps, per AWS account: ```sh -spinloop remote bootstrap # deploy the control plane (like `cdk bootstrap`) -spinloop remote bake # bake the engine AMIs; waits (~20-40 min) +spinloop cloud bootstrap # deploy the control plane (like `cdk bootstrap`) +spinloop cloud bake # bake the engine AMIs; waits (~20-40 min) ``` `bootstrap` prints a plan — account, region, resources, cost — and confirms @@ -23,7 +23,7 @@ waits until they are available, so `deploy` can go when it returns. Re-bake only when an engine version or the driver changes. Both run on the administrator's ambient AWS credentials and need Node 22 plus -a package manager. `spinloop remote auth --store` then keeps a day-to-day +a package manager. `spinloop cloud auth --store` then keeps a day-to-day credential in this machine's OS keystore so routine commands outlive SSO log-ins. @@ -33,13 +33,13 @@ The Spinloop says what the environment serves; `--env` names the environment — a machine-local choice that stays out of the file: ```sh -spinloop remote deploy --env qwen3.6-27b +spinloop cloud deploy --env qwen3.6-27b ``` Deploy reads the Spinloop and its preset — `PROVIDER` picks the engine, so the file that runs a model locally under [`spinloop serve`](local-serving.md) deploys the same model remotely — provisions the environment (Elastic IP, API -key, ingress), registers it under `~/.config/spinloop/remotes//`, and +key, ingress), registers it under `~/.config/spinloop/clouds//`, and stores what to serve. A redeploy over a registered or live environment needs `--overwrite`; it never silently clobbers a running instance. By default only your own IP may reach the instance (`--allowed-cidr` changes that). @@ -47,13 +47,13 @@ your own IP may reach the instance (`--allowed-cidr` changes that). ## The usual flow ```sh -spinloop remote start --env qwen3.6-27b # boots it (~10 min from cold) -eval "$(spinloop remote start --env qwen3.6-27b --print-env)" # ...and exports +spinloop cloud start --env qwen3.6-27b # boots it (~10 min from cold) +eval "$(spinloop cloud start --env qwen3.6-27b --print-env)" # ...and exports # OPENAI_BASE_URL + OPENAI_API_KEY spinloop harness apply --env qwen3.6-27b # point your agent at it spinloop harness open --env qwen3.6-27b # work -spinloop remote pause --env qwen3.6-27b # done for now: stopped, re-wakeable -spinloop remote stop --env qwen3.6-27b # done for good: terminated +spinloop cloud pause --env qwen3.6-27b # done for now: stopped, re-wakeable +spinloop cloud stop --env qwen3.6-27b # done for good: terminated ``` Every command that acts on an endpoint selects it with `--env ` — the @@ -66,16 +66,16 @@ for `eval` only with `--print-env`, so a plain `start` leaves stdout empty. Useful along the way: ```sh -spinloop remote status --env qwen3.6-27b # up? healthy? where? how long idle? -spinloop remote logs --env qwen3.6-27b # readable even after it's gone -spinloop remote keep 4h --env qwen3.6-27b # the idle sweep won't touch it for 4h -spinloop remote restart --env qwen3.6-27b # fresh engine, same address +spinloop cloud status --env qwen3.6-27b # up? healthy? where? how long idle? +spinloop cloud logs --env qwen3.6-27b # readable even after it's gone +spinloop cloud keep 4h --env qwen3.6-27b # the idle sweep won't touch it for 4h +spinloop cloud restart --env qwen3.6-27b # fresh engine, same address ``` ## From another machine The environment is registered per user and per machine, so two machines that -share one (each with its `remote.json` in the registry) can work the same +share one (each with its `cloud.json` in the registry) can work the same endpoint. Deploy from one, then on the other — with **no Spinloop at all**: ```sh @@ -88,7 +88,7 @@ automatically rather than re-copied by anyone. ## Where next -- [`spinloop remote`](../commands/remote.md) — every subcommand, in full -- [Run a fleet](fleet.md) — remote environments as fleet nodes, alongside daemons +- [`spinloop cloud`](../commands/cloud.md) — every subcommand, in full +- [Run a fleet](fleet.md) — cloud environments as fleet nodes, alongside daemons - [Deploying your own control plane](https://github.com/spinloop-ai/spinloop/tree/main/remote) — the AWS project behind the command, for the curious diff --git a/docs/guides/fleet.md b/docs/guides/fleet.md index 86742148..eb261d2d 100644 --- a/docs/guides/fleet.md +++ b/docs/guides/fleet.md @@ -3,7 +3,7 @@ Every machine you run serves engines from a [`spinloop daemon`](daemon.md); a `fleet.yaml` names those machines, and `spinloop` observes and drives all of them from one place — status, metrics, an interactive dashboard, starts and -stops, and logs. A fleet can also hold [remote environments](remote.md) as +stops, and logs. A fleet can also hold [cloud environments](cloud.md) as nodes, beside daemons, in the same rows. ```sh @@ -38,10 +38,10 @@ nodes: tokenEnv: GPU_BOX_TOKEN # the *name* of the variable, never the token - name: qwen # a cloud environment, beside the daemons - kind: remote + kind: cloud ``` -A `kind: remote` node is a registered [remote environment](remote.md): its +A `kind: cloud` node is a registered [cloud environment](cloud.md): its `name` is the registered one, no `host` is needed, and it is reached through its control plane. The file is found the way a `Spinloop` is — `./fleet.yaml` in the working directory, or `--fleet `. @@ -85,7 +85,7 @@ Two settings shape the choice, in the fleet file: - **[`wake`](../fleet-file.md#waking)** (`on`, the default) decides whether routing may start an engine on a node that is not running one. Set it `off` where machines are not to be started on demand; a node may declare its own - `wake` to override the file — most useful for a remote node, whose wake + `wake` to override the file — most useful for a cloud node, whose wake boots a billed cloud instance. - **[`prefer`](../fleet-file.md#spreading-or-consolidating)** ranks nodes that could all serve you: `idle` (the default) takes the machine quietest diff --git a/docs/guides/gateway.md b/docs/guides/gateway.md index ce971385..15b798b5 100644 --- a/docs/guides/gateway.md +++ b/docs/guides/gateway.md @@ -1,6 +1,6 @@ # Serve the fleet as a gateway -A fleet of [daemons](daemon.md) and [environments](remote.md) is many +A fleet of [daemons](daemon.md) and [environments](cloud.md) is many addresses. `spinloop gateway` puts one in front of them: an OpenAI-compatible endpoint where each request is answered by the fleet's own selector, and a stopped node is started — or the request refused — the way the diff --git a/docs/guides/harness.md b/docs/guides/harness.md index d40d7679..da913f5a 100644 --- a/docs/guides/harness.md +++ b/docs/guides/harness.md @@ -88,7 +88,7 @@ spinloop code -- agent-args # -- stops the parsing: everything after is the ag ``` Where the model is served is a launch concern, learned one of three ways: a -Spinloop you apply, `--env` for a [registered remote environment](remote.md), +Spinloop you apply, `--env` for a [registered cloud environment](cloud.md), or a [fleet file](fleet.md) that routes the launch to a node that serves — or is woken to serve — the wanted model. diff --git a/docs/guides/hugging-face.md b/docs/guides/hugging-face.md index c91beaea..68fd15d0 100644 --- a/docs/guides/hugging-face.md +++ b/docs/guides/hugging-face.md @@ -69,11 +69,11 @@ spinloop harness apply # point the agent at i Or skip the middle: `spinloop hf --apply` configures the active harness straight from the result. To serve the model from a cloud GPU instead of this machine, give the written Spinloop to -[`spinloop remote deploy`](remote.md). +[`spinloop cloud deploy`](cloud.md). ## Where next - [`spinloop hf`](../commands/hf.md) — the full reference, including gated repos and mirrors - [Serve a model locally](local-serving.md) — running the engine -- [Deploy to a cloud GPU](remote.md) — the same Spinloop, in the cloud +- [Deploy to a cloud GPU](cloud.md) — the same Spinloop, in the cloud diff --git a/docs/http-api.md b/docs/http-api.md index 59e6aaa6..ea78292f 100644 --- a/docs/http-api.md +++ b/docs/http-api.md @@ -115,7 +115,7 @@ Returns `400 Bad Request` if `offset` or `limit` is not a whole number, or if ### PUT `/v1/deploy-config` Updates the configuration for the *next* engine start. -- Request body: `remote.DeployConfig` JSON +- Request body: `inference.DeployConfig` JSON - Returns `200 OK` with a message indicating if the change is active now or will take effect on the next start. - Returns `400 Bad Request` if the configuration is invalid or fails to push. diff --git a/docs/index.md b/docs/index.md index 7114d0dc..20d0b8d8 100644 --- a/docs/index.md +++ b/docs/index.md @@ -33,7 +33,7 @@ Four words carry the whole tool: | [Open your coding agent](guides/harness.md) | Configure the agent for a provider and model, and launch it | | [From a Hugging Face model](guides/hugging-face.md) | Turn a model page's reference into a `Spinloop` that serves it | | [Run a daemon node](guides/daemon.md) | Keep an engine supervised over the HTTP control API, so anything can start, stop, and watch it | -| [Deploy to a cloud GPU](guides/remote.md) | The same `Spinloop`, on a machine that stops when you do | +| [Deploy to a cloud GPU](guides/cloud.md) | The same `Spinloop`, on a machine that stops when you do | | [Run a fleet](guides/fleet.md) | One spinloop observing and driving every machine you run | | [Serve the fleet as a gateway](guides/gateway.md) | The whole fleet under one OpenAI-compatible endpoint | | [Work a backlog](guides/work-items.md) | A file of work items, worked by one-shot agents at the pace the fleet allows | @@ -44,7 +44,7 @@ Four words carry the whole tool: ## Environment variables The ones you will meet first — **[Environment variables](env-vars.md) is the -full list**, including the `SPINLOOP_REMOTE_*` overrides: +full list**, including the `SPINLOOP_CLOUD_*` overrides: | Variable | Effect | | -------- | ------ | diff --git a/docs/maintainer/internals.md b/docs/maintainer/internals.md index c244f9a5..72ef2609 100644 --- a/docs/maintainer/internals.md +++ b/docs/maintainer/internals.md @@ -23,11 +23,11 @@ These are mistakes already made here; each was silent rather than loud, which is **A busy engine does not answer its own metrics endpoint.** llama.cpp serves `/metrics` from the same queue it serves inference from, so a scrape taken while a prompt is being processed waits for that prompt to finish — tens of seconds on a long context. `/v1/metrics` used to scrape inline, so the handler blocked for as long as the engine had work, and `spinloop metrics` (5s client timeout, against a handler whose scrape timeout was also 5s) could not win that race: the view went blank exactly when there was something worth watching, reporting the node as `unreachable`. The counters now come from the background sampler's last reading (`engineSample`), which the daemon takes every 15s regardless of who is asking — so the handler never waits on the engine, and staleness is bounded by the sample interval. Three things to preserve if you touch this: the *cached scrape error* is still reported, because silently omitting the token block is what once hid a scraper pointed at the wrong port; the sample is forgotten on start, so one engine's counters are never reported against the next; and the sampler retries at `catchUpInterval` (1s) until a reading lands, dropping to the full interval only afterwards — without that, a freshly started engine reports no counters for up to 15s, which is exactly the window someone watching a node they just started is looking at. -**An exported-but-empty variable is a gap, not a choice.** `setEnvIfAbsent` keys on the variable being *present*, so an `OPENAI_BASE_URL=` in the environment counts as set and suppresses the value routing meant to supply — leaving the agent pointed at nothing. `setEnvIfBlank` is the one to use for an address or a key, matching what `harnessEnv` already does for the remote endpoint's values; the distinction only shows up when something exports an empty string, which shells do more often than you would think. +**An exported-but-empty variable is a gap, not a choice.** `setEnvIfAbsent` keys on the variable being *present*, so an `OPENAI_BASE_URL=` in the environment counts as set and suppresses the value routing meant to supply — leaving the agent pointed at nothing. `setEnvIfBlank` is the one to use for an address or a key, matching what `harnessEnv` already does for the cloud endpoint's values; the distinction only shows up when something exports an empty string, which shells do more often than you would think. **A pre-warm that races the engine's faults cannot win.** The cloud's model load is I/O-bound and the engine arrives first: it maps its weights and faults pages in as it copies them to the GPU, so from the first second of a load the volume serves *its* per-page faults, and a provisioned gp3 root hands out at most 4,000 IOPS × ~68 KB ≈ 260 MB/s whatever else wants the disk. A sequential pre-warm reader that starts with the engine spends its first seconds ahead of the faults, then goes flat — the daemon's `read_bytes` did exactly that on a live instance (~1.5 GB of real EBS reads in the first seconds, then nothing for the whole load), because the engine's faults consume the budget and its readahead turns every later "read" into a cache hit. And the shape's 32 GB of RAM cannot hold a ~30 GB model in the page cache at all, so even a finished pre-warm would only shift the faults to a different second. Live-checked 2026-08-23: pre-warm on cost double-reads and saved no time (the ~30 GB model loaded in ~115 s either way). The feature was removed after that check; the provisioned gp3 throughput and IOPS stay, because the S3 sync is the one reader whose limit is the volume's. -**`contextsize.Parse` is decimal.** `128k` is 128000, not 131072 — a `CONTEXT` written that way is not the power-of-two window it looks like. It also *overrides* a preset's `ctx-size` (both in `serve` and in `remote deploy`), so the Spinloop, not the preset, decides the window whenever it states one. +**`contextsize.Parse` is decimal.** `128k` is 128000, not 131072 — a `CONTEXT` written that way is not the power-of-two window it looks like. It also *overrides* a preset's `ctx-size` (both in `serve` and in `cloud deploy`), so the Spinloop, not the preset, decides the window whenever it states one. **`up` dispatches by directory and reuses both branches.** `cmd/spinloop/up.go` routes a working-directory `fleet.yaml` to the fleet start path — `runFleetDrive` over the named nodes, or over every node when none are given, since a bare `fleet start` lists and does nothing — and everything else to `runServe`'s own body, so `up` and `serve` resolve and word things identically by construction. The completion slot is the only CWD-dependent one: `upSlot` offers the fleet's node names where `./fleet.yaml` parses, the Spinloop slot elsewhere, and nothing where a fleet file is present but unreadable — `__complete` never errors, whatever the directory holds. @@ -35,13 +35,13 @@ These are mistakes already made here; each was silent rather than loud, which is **An IAM user's inline policies are capped at 2,048 characters in aggregate.** The control plane's seven functions each take a `grantInvokeUrl` pair — two actions, the auth-type conditions, the function's ARN — and with the log-reading, stack-discovery, pricing and self-service statements the document far exceeds that; the first deploy of the `RemoteCliUser` inline policy failed with `ServiceLimitExceeded`, which CDK does not warn about ahead of time. It is now a stack-owned `AWS::IAM::ManagedPolicy` (`RemoteCliPolicy`, 6,144 cap; the deployed document measures ~2 KB, so the grant list has room to grow). Keep it managed rather than re-inlining it, and keep the iam self-service ARN built from the `AWS::Partition`/`AWS::AccountId` pseudo parameters instead of the user's `Arn`: the policy attaches to that user, so referencing the user from inside it is a dependency cycle. -**The file credential store's index is non-secret by design.** OS keystores offer no way to list entries, so the file store — used where no keystore is reachable or `SPINLOOP_REMOTE_KEYSTORE=file` — keeps a plain-text index of the stored regions beside the `0600` per-region files under `/keystore/`. The report (`spinloop remote auth`) reads the index, so a file added, removed or renamed by hand is reported wrong until the index matches; and a corrupt index is reported, not silently reset, because a report that misleads about what is stored misleads about which access keys exist on the AWS side. +**The file credential store's index is non-secret by design.** OS keystores offer no way to list entries, so the file store — used where no keystore is reachable or `SPINLOOP_CLOUD_KEYSTORE=file` — keeps a plain-text index of the stored regions beside the `0600` per-region files under `/keystore/`. The report (`spinloop cloud auth`) reads the index, so a file added, removed or renamed by hand is reported wrong until the index matches; and a corrupt index is reported, not silently reset, because a report that misleads about what is stored misleads about which access keys exist on the AWS side. **The two local model caches are separate.** A model `llama-server` downloaded sits in llama.cpp's cache (`$LLAMA_CACHE`, else the platform's user cache directory) as flat filenames; one fetched with the hub's tools sits in the Hugging Face hub cache (`$HF_HUB_CACHE`, else `$HF_HOME/hub`, else `~/.cache/huggingface/hub`) as `models----/snapshots//` of symlinks into a content-addressed blob store. Neither tool looks in the other's, so a model already on the machine is "already on the machine" in only one of the two senses. `spinloop hf` therefore resolves both roots up front (`hf.ResolveRoots`) and checks both before touching the network; a cache-aware lookup that consults only one side re-downloads what is already there. The hub-cache shape has two more traps: a snapshot entry whose symlink dangles is an interrupted or half-finished download and must count as absent, and `refs/` holds a commit sha, so a revision name is only a snapshot once it has been resolved through `refs/`. **opencode run takes its working directory from the PWD variable, not its own cwd.** `opencode run` with no directory flag sets the session's directory — and with it the directory every tool the agent runs — from `process.env.PWD` where that variable is set, falling back to the process's cwd only where it is absent (as of v1.18.30). The dispatcher gave the child the item's directory with `cmd.Dir`, but the environment it inherited carried the PWD of wherever the orchestrator was started, so the session and the agent's tools ran in the orchestrator's directory while the process's cwd sat in the item's: an item's `dir:` was silently ignored, and nothing failed. `startChild` therefore rewrites the child's PWD to the item's directory (resolved with `filepath.Abs`) before the exec, on top of `cmd.Dir` carrying it. The variable then holds the unresolved path — the `/var/...` form on macOS — which is why the stub agent's record takes the physical directory with `pwd -P` and the variable's own value with `printenv PWD`. -**`internal/daemon` must not depend on a cloud package.** What an engine should serve is `inference.DeployConfig`, in `internal/inference` — a leaf that imports only the standard library. Both the daemon and the cloud control plane speak it, so neither has to import the other to be described: a daemon is handed one over its control API, and `internal/remote` persists one against an environment. The dependency used to run the other way, `internal/daemon` importing `internal/remote` for the type and for a config-directory helper that only forwarded to `internal/config.Dir`, which made the package that knows nothing about AWS depend on the package that is nothing but AWS. Keep new shared vocabulary in `internal/inference` only where it describes an engine's workload and needs nothing of ours to express; anything cloud-shaped — `IsInstanceType` and the rest of the EC2 vocabulary — stays in `internal/remote`. +**`internal/daemon` must not depend on a cloud package.** What an engine should serve is `inference.DeployConfig`, in `internal/inference` — a leaf that imports only the standard library. Both the daemon and the cloud control plane speak it, so neither has to import the other to be described: a daemon is handed one over its control API, and `internal/cloud` persists one against an environment. The dependency used to run the other way, `internal/daemon` importing `internal/cloud` for the type and for a config-directory helper that only forwarded to `internal/config.Dir`, which made the package that knows nothing about AWS depend on the package that is nothing but AWS. Keep new shared vocabulary in `internal/inference` only where it describes an engine's workload and needs nothing of ours to express; anything cloud-shaped — `IsInstanceType` and the rest of the EC2 vocabulary — stays in `internal/cloud`. **A captured engine's stdout is a pseudo-terminal, and the normaliser above it is a line model, not a screen emulator.** llama.cpp's download bar prints nothing when its stdout is not a terminal, and the download happens in-process before the HTTP listener comes up — there is no status API to scrape — so the capture (the daemon and the serve view) presents the engine's stdout as a PTY: the only way the progress reaches the engine log at all. `ptylog.go` turns the terminal stream into the log's lines as a column, and three of its rules are easy to break by "simplifying". The engine's up-down cursor dance is *part of each redraw* — the bar's state is drawn between a cursor-up and a cursor-back-down, the cursor parked on an anchor line — so a cursor move must move the drawing between the column's lines and never end a line: committing on a move records every state as a "final" one and resets the dedup, which is how the first design put the whole download, state for state, in the log. A state's replacement is lazy — a carriage return settles the state drawn before it, and the line's content is cleared only by the next text byte — so `S\r\r\n`, the PTY's ONLCR turning a lone `\n` into a CRLF, commits `S` intact. And the tick at the interval's half-rate runs on *every* line of the column, because a bar's last state outlives the engine's move off its line: the download finishes, the engine goes quiet on stdout, and the log still owes it. Because the engine's stdout is now a terminal, the capture branch sets `NO_COLOR=1` in the engine's environment: llama.cpp routes its log lines to stderr and colours them by whether it sees a terminal — *stdout* among them — and the normaliser never sees the stderr path, so without it escapes would reach the log file. The bar draws no colour, so nothing is lost; the forwarding paths (piped `serve`, off-terminal `serve --api`) set nothing and must stay byte-for-byte what the engine wrote — a test pins the raw `\r` surviving. Where no pseudo-terminal can be opened — Windows, a restricted environment — the fallback writes stdout to the log file directly, no worse than the capture did before, and the engine is told `NO_COLOR` the same way; a test stands the opening in for with a failure to keep that branch alive. The pump always drains — the throttle delays the log's writes, never the reads, so the engine can never wedge on a full terminal — and the log file closes only after the pump's final record, never under it. **Calling `prog.Send` from inside `Update` deadlocks the program.** The event loop reads its own channel on the goroutine that calls `Update`, so a command that hands a message back through `prog.Send` while `Update` is on the stack waits for a reader that is waiting for it. The board hung on exactly this: pressing Enter on the form's last field, or `a`/`x` on a card, and the whole TUI went still — no spinner, no keystrokes, Ctrl-C ignored. Every message from work that starts inside `Update` must *return* through the loop instead: `tea.Batch(workCmd, spinCmd)` delivers each command's messages normally, and a repaint chain re-arms through its own returned `tea.Tick`. The tests guard this by landing commands through the real program loop, not by calling `Update` and discarding the command. @@ -50,10 +50,10 @@ These are mistakes already made here; each was silent rather than loud, which is A few Bubble Tea/lipgloss specifics that are easy to break by "simplifying": -- The `tea.Program` holds the model by **pointer**: Bubble Tea never reads a value model's `Init` back, so the first round's mutations (its deadline spend — real cloud calls, for remote environments) would be silently discarded. +- The `tea.Program` holds the model by **pointer**: Bubble Tea never reads a value model's `Init` back, so the first round's mutations (its deadline spend — real cloud calls, for cloud environments) would be silently discarded. - The tick reschedules itself whenever it fires; a one-shot `tea.Tick` without the reschedule leaves the board still after the second round. A second, faster chain (`dashSpinTickMsg`) runs only while an action is in flight, so the spinner and the elapsed time beside a verb advance; it stops on the first tick that finds nothing in flight. - Every reading carries the time its own call returned (`fleet.NodeResult.At`, set in the fan-out), and the board draws a reading only when it was taken later than the one already on screen. Reads run concurrently and take differing times, so a reading can land after one taken later than it — including a round issued before an action finished and landing after it, which would otherwise repaint the node's pre-action state. -- A start reports a `fleet.StartPhase` — what it is doing, when that began, when the next attempt is due — rather than a line of text, and `fleet.RenderPhase(phase, now)` is the only place it becomes text. A wait therefore counts down and a boot counts up on every repaint, and a situation the start has moved on from cannot be left on the tile: each phase replaces the one before it. `spinloop remote start` renders the same phases to stderr, so the tile and the CLI cannot word one situation differently. +- A start reports a `fleet.StartPhase` — what it is doing, when that began, when the next attempt is due — rather than a line of text, and `fleet.RenderPhase(phase, now)` is the only place it becomes text. A wait therefore counts down and a boot counts up on every repaint, and a situation the start has moved on from cannot be left on the tile: each phase replaces the one before it. `spinloop cloud start` renders the same phases to stderr, so the tile and the CLI cannot word one situation differently. - One function, `dashNodeView`, produces both a panel's lines and its health tier, from the reading, the action, the current time, and how old a reading of that node may be. Nothing in it reads a clock, so every pairing of a start's phase against a reading can be enumerated in a test. - A tile's first line is a header bar drawn in raw ANSI — the body is one plain string under a single lipgloss style, so per-character colour cannot be lipgloss's. The board's own title bar (`dashTitleBar`) uses lipgloss instead, and the two share one surface index (`barSurface`) because they are set through different mechanisms and would otherwise drift. - A grid row joins the *corresponding lines* of the tiles it places, not the tile blocks — joining whole blocks glues the second tile's top border to the first tile's bottom border and shifts its body down a line. diff --git a/docs/openapi.yaml b/docs/openapi.yaml index e9b019d5..ebd3d287 100644 --- a/docs/openapi.yaml +++ b/docs/openapi.yaml @@ -351,7 +351,7 @@ components: useless to anyone else, and it cannot know the name a client reaches this host by — a LAN name, a tailscale name, a published container port. The caller composes these against the host it already has. A node - that does know that name — a remote environment, whose control plane + that does know that name — a cloud environment, whose control plane publishes the instance's address — reports it in `host`. required: [port] properties: @@ -359,7 +359,7 @@ components: type: string description: | The name or address a client reaches the engine by, when the node - knows it. A daemon leaves it absent; a remote environment's status + knows it. A daemon leaves it absent; a cloud environment's status fills it with the instance's published address. port: type: integer @@ -461,7 +461,7 @@ components: The environment's retention deadline, RFC 3339: the idle sweep will not terminate the instance before it. A property of the cloud instance, not the engine, so it is absent for local daemon nodes - and for remote environments without an update URL. The stats reply + and for cloud environments without an update URL. The stats reply carries it only while it is a time in the future — a passed deadline keeps nothing, so it is dropped there and on this read. @@ -569,7 +569,7 @@ components: DeployConfig: type: object description: | - What to serve, in the same shape `spinloop remote deploy` derives from an + What to serve, in the same shape `spinloop cloud deploy` derives from an Spinloop. Mirrors Go's remote.DeployConfig. There is no default runner: an unset or invalid one fails loudly rather than guessing. required: [runner] diff --git a/docs/spinloop-file.md b/docs/spinloop-file.md index b7a4132b..d7ade33e 100644 --- a/docs/spinloop-file.md +++ b/docs/spinloop-file.md @@ -72,7 +72,7 @@ because the Spinloop itself was read. ## Running the model on a cloud GPU -For a model too big for your machine, `spinloop remote` runs it on a +For a model too big for your machine, `spinloop cloud` runs it on a scale-to-zero GPU endpoint — one that runs only while you're using it. The Spinloop says what the endpoint serves: @@ -89,24 +89,24 @@ The Spinloop says *what*; it no longer says *where*. Where is an with a `--env ` flag on the commands that act on it: ```sh -spinloop remote deploy --env qwen3.6-27b-prod # from the directory holding the Spinloop +spinloop cloud deploy --env qwen3.6-27b-prod # from the directory holding the Spinloop spinloop harness apply --env qwen3.6-27b-prod # point opencode at it spinloop harness open --env qwen3.6-27b-prod # work -spinloop remote stop --env qwen3.6-27b-prod # done +spinloop cloud stop --env qwen3.6-27b-prod # done ``` Each `--env ` reads the environment's registered config at -`${XDG_CONFIG_HOME:-~/.config}/spinloop/remotes//remote.json` — the file -[`spinloop remote deploy`](commands/remote.md) writes when it creates the +`${XDG_CONFIG_HOME:-~/.config}/spinloop/clouds//cloud.json` — the file +[`spinloop cloud deploy`](commands/cloud.md) writes when it creates the environment — so deployment state stays per-user and per-machine while the Spinloop itself stays clean enough to commit. A command given no `--env` uses the `default` environment, and an unregistered name fails, naming the `deploy --env` that would create it. A name is a plain identifier: `--env ./x.json` -fails, saying so. See [`spinloop remote`](commands/remote.md) for the full +fails, saying so. See [`spinloop cloud`](commands/cloud.md) for the full lifecycle. Note the missing `BASEURL`: the endpoint's address belongs to the deployment, -which records it in the environment's `remote.json` as `base_url`, and +which records it in the environment's `cloud.json` as `base_url`, and [`spinloop harness apply`](commands/harness.md#spinloop-harness-apply) reads it from there. Write a `BASEURL` only to override that. @@ -116,7 +116,7 @@ with the model reading as `qwen3.6-27b-prod/qwen3.6-27b`. `PROVIDER` still supplies the engine's settings; only the name changes, so several environments built from the same engine each keep their own entry instead of overwriting one. The provider's display name is qualified by the environment too — `llama.cpp -(qwen3.6-27b-prod)` rather than a bare `llama.cpp` — so a remote environment reads +(qwen3.6-27b-prod)` rather than a bare `llama.cpp` — so a cloud environment reads distinctly from a local engine of the same kind in a harness model picker. A launch may not state both `--env` and a fleet (the `--fleet` flag, or the @@ -173,7 +173,7 @@ harness`](commands/fleet.md#launching-the-harness) needs no Spinloop at all when the fleet file names one: with none given, it configures a generic OpenAI-compatible provider at the gateway's address itself, its model list populated from the gateway's own listing, and no `MODEL`/`ALIAS` to pick. -That provider reads the way a remote environment's does — labelled by the +That provider reads the way a cloud environment's does — labelled by the gateway's own `name`, or its address when the section names none — the same "llama.cpp (dev-2)" pattern described above. @@ -195,10 +195,10 @@ One instruction per line: a keyword followed by a single value. | `ALIAS` | one of `MODEL`/`ALIAS` | `--alias` | `ALIAS deepseek` | | `CONTEXT` | no | `--context` | `CONTEXT 128k` | | `OUTPUT` | no | `--output` | `OUTPUT 32k` | -| `PARALLEL` | no | `spinloop serve`, `spinloop remote deploy` | `PARALLEL 2` | +| `PARALLEL` | no | `spinloop serve`, `spinloop cloud deploy` | `PARALLEL 2` | | `BASEURL` | no | `--base-url` | `BASEURL https://gateway/v1` | | `PRESET` | no | `spinloop serve` | `PRESET ./preset.ini` | -| `ENV` | no (repeatable) | `spinloop remote`, `spinloop harness` | `ENV AWS_PROFILE=prod` | +| `ENV` | no (repeatable) | `spinloop cloud`, `spinloop harness` | `ENV AWS_PROFILE=prod` | Rules: @@ -220,7 +220,7 @@ Rules: out, `spinloop` records a quarter of the context. It cannot exceed the context window. - `PARALLEL` sets the number of concurrent request slots for `spinloop serve` - and `spinloop remote deploy` — a plain integer, not a size. It has no meaning + and `spinloop cloud deploy` — a plain integer, not a size. It has no meaning for a hosted provider selection, only for a served engine, so unlike `CONTEXT`/`OUTPUT` it has no `add`/`remove` CLI flag. Since `CONTEXT` always means "context per request", and llama.cpp's own `--ctx-size` is a total @@ -240,15 +240,15 @@ Rules: [`spinloop serve`](commands/serve.md); `apply` ignores it. A relative path resolves against the Spinloop's own directory, or against its URL when the Spinloop itself was fetched from one; `PRESET` may also be an absolute URL of - its own, fetched only when `serve` (or `spinloop remote deploy`) builds the + its own, fetched only when `serve` (or `spinloop cloud deploy`) builds the launch command — never merely because the Spinloop was read. The file is read in the flag vocabulary of the engine `PROVIDER` names, so a preset written for llama.cpp is not portable to oMLX and vice versa. - `ENV` sets an environment variable on the machine running `spinloop` and is the one keyword that **may repeat**. Its value is a single `KEY=VALUE` token (no - spaces). The `spinloop remote` commands read it — along with a `.env` beside the + spaces). The `spinloop cloud` commands read it — along with a `.env` beside the Spinloop — before they sign their AWS calls, so credentials, region and - `SPINLOOP_REMOTE_*` overrides can travel with the Spinloop. `spinloop harness open` reads + `SPINLOOP_CLOUD_*` overrides can travel with the Spinloop. `spinloop harness open` reads it too, passing the whole `.env` and the `ENV` lines to the agent it launches. Precedence, highest to lowest: an `ENV` line, then a variable already set in your shell, then the `.env` — the same rule everywhere spinloop resolves local @@ -256,10 +256,10 @@ Rules: to a deployed instance, and on the harness path it shapes only the launched agent, never spinloop's own environment. - The `REMOTE` keyword was removed: a `REMOTE` line fails, naming the line and - the replacement — `spinloop remote deploy --env ` at deploy time and - `--env ` on `apply`, `unapply`, `harness`, and the `remote` + the replacement — `spinloop cloud deploy --env ` at deploy time and + `--env ` on `apply`, `unapply`, `harness`, and the `cloud` subcommands. Where a `REMOTE` line pointed at a path or a URL, register the - environment's config with `spinloop remote deploy --env ` instead, and + environment's config with `spinloop cloud deploy --env ` instead, and name it from the flags. - Keywords are **case-insensitive** — `provider`, `Provider`, and `PROVIDER` are all accepted — but **UPPERCASE is canonical** and is what `spinloop harness export` diff --git a/docs/troubleshooting.md b/docs/troubleshooting.md index 43734c08..51abba2b 100644 --- a/docs/troubleshooting.md +++ b/docs/troubleshooting.md @@ -70,17 +70,17 @@ scaled; see [`spinloop serve`](commands/serve.md#parallelism). directory](env-vars.md#config-directory-resolution), then start the engine again through the API. -## A remote endpoint won't come up +## A cloud endpoint won't come up -- **The control plane is not deployed.** Every `remote` command that acts on - an endpoint needs `spinloop remote bootstrap` to have run once per account; +- **The control plane is not deployed.** Every `cloud` command that acts on + an endpoint needs `spinloop cloud bootstrap` to have run once per account; a missing control plane says so. An older control plane that lacks a feature says to re-run `bootstrap` to add it. - **A warning says the control plane is at a different version.** The `spinloop` you are running differs from the one that deployed the control - plane. Commands still run, but re-run `spinloop remote bootstrap` to update + plane. Commands still run, but re-run `spinloop cloud bootstrap` to update it, or install the matching `spinloop`. -- **The AMI is not baked.** `spinloop remote bake` once per engine, and it +- **The AMI is not baked.** `spinloop cloud bake` once per engine, and it waits until the AMI is available. - **A cold start takes about ten minutes.** `start` prints its progress on stderr; `--timeout` (default 15m) bounds the wait. `status` and `logs` @@ -91,9 +91,9 @@ scaled; see [`spinloop serve`](commands/serve.md#parallelism). ## A fleet row is wrong, or a node won't answer -- **`config-error` on a `kind: remote` row** means the environment is not - registered on this machine — `spinloop remote deploy --env ` (or - `spinloop fleet deploy`) writes its `remote.json`. +- **`config-error` on a `kind: cloud` row** means the environment is not + registered on this machine — `spinloop cloud deploy --env ` (or + `spinloop fleet deploy`) writes its `cloud.json`. - **A node's token is not in *your* shell.** The fleet file names the *variable* (`tokenEnv`), never the value. Set the variable the row names, from the `.env` beside the fleet file or your environment. @@ -112,8 +112,8 @@ scaled; see [`spinloop serve`](commands/serve.md#parallelism). be woken (its own `wake`, or the file's, is off). The failure names the node that would have woken and the `spinloop fleet start ` command that would start it. -- **An undeployed remote environment is never a candidate** — it has nothing - to serve yet; choosing what to deploy is `spinloop remote deploy`'s call. +- **An undeployed cloud environment is never a candidate** — it has nothing + to serve yet; choosing what to deploy is `spinloop cloud deploy`'s call. - **404 naming other paths** — the gateway serves `/health`, `/v1/models`, `/v1/chat/completions`, `/v1/completions`, and `/v1/fleet`, and says so. diff --git a/examples/fleet-remote/README.md b/examples/fleet-cloud/README.md similarity index 81% rename from examples/fleet-remote/README.md rename to examples/fleet-cloud/README.md index e78d4f8e..6b9914d2 100644 --- a/examples/fleet-remote/README.md +++ b/examples/fleet-cloud/README.md @@ -1,6 +1,6 @@ -# A fleet of remote environments +# A fleet of cloud environments -`spinloop fleet` observes [`spinloop remote`](../../docs/commands/remote.md) +`spinloop fleet` observes [`spinloop cloud`](../../docs/commands/cloud.md) environments the same way it observes machines running `spinloop daemon`: each environment becomes a node. There are no bearer tokens here — the control plane signs each call with your AWS credentials — and the file names environments, @@ -11,14 +11,14 @@ never an account, so it is safe to keep under version control. ### 1. Have some environments An environment is created and registered when you deploy into it: `spinloop -remote deploy --env `, run in a directory whose `Spinloop` describes what +cloud deploy --env `, run in a directory whose `Spinloop` describes what the environment serves, writes that environment's control URLs to -`~/.config/spinloop/remotes//remote.json`. +`~/.config/spinloop/clouds//cloud.json`. Deploying needs the shared control plane once before it — see [`spinloop -remote`](../../docs/commands/remote.md). List what you already have: +cloud`](../../docs/commands/cloud.md). List what you already have: ```sh -spinloop remote ls +spinloop cloud ls ``` Or create both of this fleet's environments straight from the file: @@ -39,10 +39,10 @@ its own Spinloop file. ```yaml nodes: - name: qwen # the registered environment, and what you type at `fleet start qwen` - kind: remote # resolved from qwen/Spinloop — see fleet.yaml + kind: cloud # resolved from qwen/Spinloop — see fleet.yaml - name: llama - kind: remote # resolved from llama/Spinloop the same way + kind: cloud # resolved from llama/Spinloop the same way ``` Neither node declares a `file` field: each one's own name is already a @@ -74,4 +74,4 @@ fleet still renders. - [`spinloop fleet`](../../docs/commands/fleet.md) — the commands over these nodes - [The fleet file](../../docs/fleet-file.md) — the file's format, node kinds, and routing -- [`spinloop remote`](../../docs/commands/remote.md) — the environments these nodes drive +- [`spinloop cloud`](../../docs/commands/cloud.md) — the environments these nodes drive diff --git a/examples/fleet-remote/fleet.yaml b/examples/fleet-cloud/fleet.yaml similarity index 74% rename from examples/fleet-remote/fleet.yaml rename to examples/fleet-cloud/fleet.yaml index b69e1035..0909ff1c 100644 --- a/examples/fleet-remote/fleet.yaml +++ b/examples/fleet-cloud/fleet.yaml @@ -1,18 +1,18 @@ -# A fleet of `spinloop remote` environments, observed like machines. +# A fleet of `spinloop cloud` environments, observed like machines. # # spinloop status # one row per environment: state and what it serves # spinloop metrics -w # a live dashboard # spinloop fleet deploy --all # create both environments from this file # # Each node's `name` is the registered environment it drives — one per -# `spinloop remote deploy`, or per `spinloop fleet deploy` — and what you type -# at `fleet start `. `kind: remote` marks it as a cloud environment +# `spinloop cloud deploy`, or per `spinloop fleet deploy` — and what you type +# at `fleet start `. `kind: cloud` marks it as a cloud environment # rather than a host. No bearer tokens here: the control plane signs each -# call, and the environment's URLs live in its remote.json, not in this file. +# call, and the environment's URLs live in its cloud.json, not in this file. # So, like every fleet file, this one names environments, never an account. # # `fleet deploy` needs to know what each environment serves — the same thing -# `spinloop remote deploy ` needs, just resolved from the node rather +# `spinloop cloud deploy ` needs, just resolved from the node rather # than typed on the command line. Neither node here declares a `file` field: # each one's own name is a subdirectory beside this file (qwen/Spinloop, # llama/Spinloop) — the convention for a fleet laid out as one subdirectory @@ -23,8 +23,8 @@ # somewhere that does *not* match the node's own name. nodes: - name: qwen - kind: remote + kind: cloud instance-type: g6e.2xlarge # EC2 type this environment launches as (default: the control plane's) - name: llama - kind: remote + kind: cloud diff --git a/examples/fleet-remote/llama/Spinloop b/examples/fleet-cloud/llama/Spinloop similarity index 100% rename from examples/fleet-remote/llama/Spinloop rename to examples/fleet-cloud/llama/Spinloop diff --git a/examples/fleet-remote/qwen/Spinloop b/examples/fleet-cloud/qwen/Spinloop similarity index 55% rename from examples/fleet-remote/qwen/Spinloop rename to examples/fleet-cloud/qwen/Spinloop index 1f5e1261..24f5ab92 100644 --- a/examples/fleet-remote/qwen/Spinloop +++ b/examples/fleet-cloud/qwen/Spinloop @@ -1,6 +1,6 @@ # What the "qwen" environment serves. `spinloop fleet deploy qwen` reads this -# the same way `spinloop remote deploy --env qwen` does when run in this directory. -# See docs/commands/remote.md for what each instruction does. +# the same way `spinloop cloud deploy --env qwen` does when run in this directory. +# See docs/commands/cloud.md for what each instruction does. PROVIDER llamacpp ALIAS qwen3.6-27b MODEL unsloth/Qwen3.6-27B-MTP-GGUF:UD-Q6_K_XL diff --git a/examples/fleet-local/README.md b/examples/fleet-local/README.md index 0de20f56..13675489 100644 --- a/examples/fleet-local/README.md +++ b/examples/fleet-local/README.md @@ -59,7 +59,7 @@ real network: llama.cpp's default is what the preset uses. You need a block when the daemon cannot know: a container publishing the engine elsewhere, or a proxy. - **Loopback engine is fine here.** llama.cpp binds `127.0.0.1` unless told - otherwise, and this node *is* loopback. On a remote node that same engine + otherwise, and this node *is* loopback. On a cloud node that same engine would be unreachable, and routing says so rather than handing you an address that refuses connections. diff --git a/examples/fleet-mixed/.env.example b/examples/fleet-mixed/.env.example index 01ce4334..1e3fa3a1 100644 --- a/examples/fleet-mixed/.env.example +++ b/examples/fleet-mixed/.env.example @@ -1,7 +1,7 @@ # Copy to .env beside fleet.yaml and fill in. Never commit the real one — # fleet.yaml references these by name so the secrets stay here. # -# Only the daemon node needs a token here. The remote environments sign each +# Only the daemon node needs a token here. The cloud environments sign each # call with your AWS credentials, so they need nothing in this file. # # GPU_BOX_TOKEN is the SPINLOOP_API_TOKEN of that machine's daemon. diff --git a/examples/fleet-mixed/README.md b/examples/fleet-mixed/README.md index 2100bbf1..9a1e514f 100644 --- a/examples/fleet-mixed/README.md +++ b/examples/fleet-mixed/README.md @@ -1,7 +1,7 @@ # A mixed fleet One `fleet.yaml`, one set of commands, two kinds of node: machines running -`spinloop daemon` and [`spinloop remote`](../../docs/commands/remote.md) +`spinloop daemon` and [`spinloop cloud`](../../docs/commands/cloud.md) environments. The same fan-out reaches every node, so the fleet reads as a single table. @@ -40,10 +40,10 @@ spinloop fleet start gpu-box # tell the daemon what to run, and start it spinloop fleet deploy --all # create both environments from this file ``` -Creating the environments this way is the same as running `spinloop remote +Creating the environments this way is the same as running `spinloop cloud deploy --env ` once per node in this file, from the directory holding that node's `Spinloop` — see [`spinloop -remote`](../../docs/commands/remote.md) — just one command for both. +cloud`](../../docs/commands/cloud.md) — just one command for both. ### 3. Observe the whole fleet @@ -61,6 +61,6 @@ and the rest of the fleet still shows. ## See also - [`examples/fleet`](../fleet/README.md) — a fleet of daemons only -- [`examples/fleet-remote`](../fleet-remote/README.md) — a fleet of remote environments only +- [`examples/fleet-cloud`](../fleet-cloud/README.md) — a fleet of cloud environments only - [`spinloop fleet`](../../docs/commands/fleet.md) — the commands over these nodes - [The fleet file](../../docs/fleet-file.md) — the file's format, node kinds, and routing diff --git a/examples/fleet-mixed/fleet.yaml b/examples/fleet-mixed/fleet.yaml index c4256d2b..cbc38e50 100644 --- a/examples/fleet-mixed/fleet.yaml +++ b/examples/fleet-mixed/fleet.yaml @@ -1,13 +1,13 @@ -# A mixed fleet: a machine running `spinloop daemon` and two `spinloop remote` +# A mixed fleet: a machine running `spinloop daemon` and two `spinloop cloud` # environments, observed side by side. The same command reaches every node, # whatever kind it is. # # spinloop status, or spinloop metrics -w # # A daemon node names the machine (host) and the variable holding its bearer -# token (tokenEnv). A remote node names its environment — its node name *is* the +# token (tokenEnv). A cloud node names its environment — its node name *is* the # env. The token and the control URLs stay out of this file — in the .env and -# in each environment's remote.json — so it names machines and environments, +# in each environment's cloud.json — so it names machines and environments, # never secrets or accounts. # # None of the three nodes below declares a `file` field: each one's own name @@ -23,7 +23,7 @@ nodes: # Two cloud environments, driven through their control plane. - name: qwen - kind: remote + kind: cloud - name: llama - kind: remote + kind: cloud diff --git a/examples/fleet/studio/Spinloop b/examples/fleet/studio/Spinloop index 4b565820..d2e93a07 100644 --- a/examples/fleet/studio/Spinloop +++ b/examples/fleet/studio/Spinloop @@ -1,7 +1,7 @@ # What studio runs — read by `spinloop fleet start studio`, which derives a # deploy config from it and pushes it to the daemon. PROVIDER llamacpp here, # not mlx: pushing a config to a node — from `fleet start` or from a routed -# wake — only supports the runners remote deploy also supports, llamacpp or +# wake — only supports the runners cloud deploy also supports, llamacpp or # vllm. That is a pre-existing limit of the derivation this reuses, not # something new here. PROVIDER llamacpp diff --git a/examples/llamacpp/muse-glimmer-30b/README.md b/examples/llamacpp/muse-glimmer-30b/README.md index b5d9c6cf..b4872d2f 100644 --- a/examples/llamacpp/muse-glimmer-30b/README.md +++ b/examples/llamacpp/muse-glimmer-30b/README.md @@ -374,7 +374,7 @@ which this text-only example avoids anyway. **The drafter needs one extra line for the cloud.** Locally the preset relies on `--spec-type draft-dflash` pulling the `dflash-` sibling off `--hf-repo`; -`spinloop remote deploy` does not go through that path. It reads +`spinloop cloud deploy` does not go through that path. It reads `spec-draft-model` from the preset, takes its **basename** and asks the seed for that file from the model's own repo, so the local path is never sent and the instance loads its own synced copy. Add: @@ -410,7 +410,7 @@ MODEL meta-models/Muse-Glimmer-30B-GGUF:kquant-dynamic …and register it under a name of your choosing: ```sh -spinloop remote deploy --env muse-glimmer-30b +spinloop cloud deploy --env muse-glimmer-30b ``` Be precise with that suffix. The seed downloads everything matching @@ -429,14 +429,14 @@ Adding `MODEL` breaks the local `spinloop serve` path above, since it becomes doesn't resolve. Keep separate Spinloops if you want both. ```sh -spinloop remote deploy --env muse-glimmer-30b --dry-run -spinloop remote deploy --env muse-glimmer-30b -eval "$(spinloop remote start --env muse-glimmer-30b --print-env)" # --print-env +spinloop cloud deploy --env muse-glimmer-30b --dry-run +spinloop cloud deploy --env muse-glimmer-30b +eval "$(spinloop cloud start --env muse-glimmer-30b --print-env)" # --print-env # is what prints the OPENAI_BASE_URL/OPENAI_API_KEY # export lines; without it, start's output is # progress text on stderr and there is # nothing on stdout for eval to run -spinloop remote stop --env muse-glimmer-30b +spinloop cloud stop --env muse-glimmer-30b ``` Costs and the idle/max-runtime bounds are in @@ -446,5 +446,5 @@ Costs and the idle/max-runtime bounds are in - [`examples/llamacpp/qwen3.6-27b`](../qwen3.6-27b/README.md) — the example this one is modelled on. -- [`docs/commands/remote.md`](../../../docs/commands/remote.md) — full - `spinloop remote` reference. +- [`docs/commands/cloud.md`](../../../docs/commands/cloud.md) — full + `spinloop cloud` reference. diff --git a/examples/llamacpp/muse-glimmer-30b/Spinloop b/examples/llamacpp/muse-glimmer-30b/Spinloop index d5595e1b..a942be99 100644 --- a/examples/llamacpp/muse-glimmer-30b/Spinloop +++ b/examples/llamacpp/muse-glimmer-30b/Spinloop @@ -24,10 +24,10 @@ PRESET ./preset.ini # `spinloop serve` builds the llama-server command f # the dynamic one — so the repo and file are set separately in the preset # (hf + hff). See README.md. # -# For `spinloop remote deploy` you DO need a MODEL line, because the cloud seed +# For `spinloop cloud deploy` you DO need a MODEL line, because the cloud seed # globs filenames rather than resolving a tag. The glob is case-insensitive, so # this still matches Muse-Glimmer-30B-KQuant-Dynamic-Q4_K_XL.gguf: # MODEL meta-models/Muse-Glimmer-30B-GGUF:kquant-dynamic -# (then `spinloop remote deploy --env muse-glimmer-30b` from this directory) +# (then `spinloop cloud deploy --env muse-glimmer-30b` from this directory) # Adding MODEL breaks the local `spinloop serve` path above — read the "Deploying # to the cloud" section of README.md before you do. diff --git a/examples/llamacpp/muse-glimmer-30b/preset.ini b/examples/llamacpp/muse-glimmer-30b/preset.ini index 7dcfae68..d6b275b0 100644 --- a/examples/llamacpp/muse-glimmer-30b/preset.ini +++ b/examples/llamacpp/muse-glimmer-30b/preset.ini @@ -116,7 +116,7 @@ spec-draft-n-max = 15 # dflash.block_size is 16, and in-place denoising yields # Muse-Glimmer-30B-KQuant-17GB-Q4_K_M.gguf (1.0% degradation instead of 0.2%, # 2.7 GB smaller) or halve the KV with `ctk = q8_0` and `ctv = q8_0`. # -# `spinloop remote deploy` does not read --hf-repo's auto-download, so the +# `spinloop cloud deploy` does not read --hf-repo's auto-download, so the # cloud path needs the drafter named explicitly: # # spec-draft-model = ./Muse-Glimmer-30B-GGUF/dflash-Muse-Glimmer-30B-Q4_K_M.gguf diff --git a/examples/llamacpp/qwen3.6-27b/README.md b/examples/llamacpp/qwen3.6-27b/README.md index a49b66da..e90af208 100644 --- a/examples/llamacpp/qwen3.6-27b/README.md +++ b/examples/llamacpp/qwen3.6-27b/README.md @@ -146,5 +146,5 @@ Now start `opencode` and select `llamacpp/qwen3.6-27b`. - The bigger mixture-of-experts sibling: [`examples/llamacpp/qwen3.6-35b-a3b`](../qwen3.6-35b-a3b/README.md) -- The next generation of this dense size, with `spinloop remote` deployment to +- The next generation of this dense size, with `spinloop cloud` deployment to AWS: [`examples/llamacpp/qwen3.8-27b`](../qwen3.8-27b/README.md) diff --git a/examples/llamacpp/qwen3.8-27b/README.md b/examples/llamacpp/qwen3.8-27b/README.md index 5326ef80..87a46921 100644 --- a/examples/llamacpp/qwen3.8-27b/README.md +++ b/examples/llamacpp/qwen3.8-27b/README.md @@ -2,7 +2,7 @@ Run Unsloth's GGUF build of Qwen3.8-27B locally with `llama-server`, then point opencode at it with the [`Spinloop`](Spinloop) in this directory. The same file also -deploys it to a GPU in AWS with [`spinloop remote`](#running-it-on-aws) — no +deploys it to a GPU in AWS with [`spinloop cloud`](#running-it-on-aws) — no infrastructure to hand-write, just this Spinloop and one extra line. Qwen3.8-27B is a dense 27B model built on Qwen's hybrid attention architecture @@ -185,12 +185,12 @@ over to a local/`llama.cpp` deployment, and what doesn't: GGUF quants here come from the community (Unsloth, bartowski, ggml-org). It's a good fit for a single-GPU box with opencode, which is what this Spinloop is for; for serious throughput, `spinloop`'s `vllm` provider and - `spinloop remote`'s vLLM runner are the closer match to Qwen's guidance. + `spinloop cloud`'s vLLM runner are the closer match to Qwen's guidance. ## Running it on AWS The same Spinloop and preset run this model on a GPU in the cloud — provisioned -by [`spinloop remote`](../../../docs/commands/remote.md) — rather than the +by [`spinloop cloud`](../../../docs/commands/cloud.md) — rather than the machine in front of you, and terminate themselves once you stop using them. This is real, billed AWS infrastructure (an EC2 GPU instance, an Elastic IP, image-builder pipelines), so each step below shows you a plan and asks for @@ -199,8 +199,8 @@ confirmation before it creates anything. ### Once per AWS account: bootstrap the control plane ```sh -spinloop remote bootstrap # shows a plan, then deploys -spinloop remote bootstrap --dry-run # see the plan without deploying +spinloop cloud bootstrap # shows a plan, then deploys +spinloop cloud bootstrap --dry-run # see the plan without deploying ``` This deploys the shared control plane — the AMI-baking pipelines for @@ -211,30 +211,30 @@ vCPU quota in the target region. It creates no instance and no Elastic IP. ### Deploy this Spinloop as an environment ```sh -spinloop remote deploy --env qwen3.8-27b # from this directory -spinloop remote deploy --env qwen3.8-27b --dry-run # see what would be sent first +spinloop cloud deploy --env qwen3.8-27b # from this directory +spinloop cloud deploy --env qwen3.8-27b --dry-run # see what would be sent first ``` -The `--env` name is the environment `spinloop remote` creates and registers. +The `--env` name is the environment `spinloop cloud` creates and registers. `deploy` reads `PROVIDER`, `ALIAS`, `CONTEXT` and `PRESET` from the Spinloop — the same values [`spinloop serve`](../../../docs/commands/serve.md) uses locally — provisions the environment's Elastic IP, API key, ingress rule (defaulting to your own public IP) and state, and registers it at -`~/.config/spinloop/remotes/qwen3.8-27b/remote.json`. If the shared bucket +`~/.config/spinloop/clouds/qwen3.8-27b/cloud.json`. If the shared bucket doesn't have these weights cached yet, deploy fetches them in the background (15–20 minutes) — wait for that before your first `start`. ### Start it, use it, stop it ```sh -eval "$(spinloop remote start --env qwen3.8-27b)" # boots the instance +eval "$(spinloop cloud start --env qwen3.8-27b)" # boots the instance # (~10 min cold), exports # OPENAI_BASE_URL / OPENAI_API_KEY spinloop status --env --env qwen3.8-27b # is it up, is it healthy spinloop harness apply --env qwen3.8-27b # point opencode at the running endpoint spinloop harness open --env qwen3.8-27b # work -spinloop remote stop --env qwen3.8-27b # done — shut it down now rather +spinloop cloud stop --env qwen3.8-27b # done — shut it down now rather # than waiting for the idle timer ``` @@ -242,7 +242,7 @@ Once deployed, this box has the memory to run past the 32768-token default — raise `CONTEXT`/`ctx-size` in the Spinloop and preset together (up to the model's native 262144) before your next `deploy`. -See [`spinloop remote`](../../../docs/commands/remote.md) for `logs`, `metrics`, +See [`spinloop cloud`](../../../docs/commands/cloud.md) for `logs`, `metrics`, and how to name and switch between multiple deployed environments. ## Vision input (optional) diff --git a/examples/llamacpp/qwen3.8-27b/Spinloop b/examples/llamacpp/qwen3.8-27b/Spinloop index aa3a55d8..c5cd2d1a 100644 --- a/examples/llamacpp/qwen3.8-27b/Spinloop +++ b/examples/llamacpp/qwen3.8-27b/Spinloop @@ -1,4 +1,4 @@ -# Qwen3.8-27B served locally by llama.cpp, or on a cloud GPU via spinloop remote. +# Qwen3.8-27B served locally by llama.cpp, or on a cloud GPU via spinloop cloud. # ALIAS names the model in opencode and selects the preset section below. PROVIDER llamacpp ALIAS qwen3.8-27b diff --git a/examples/llamacpp/qwen3.8-27b/preset.ini b/examples/llamacpp/qwen3.8-27b/preset.ini index 1700e934..fdaed0e4 100644 --- a/examples/llamacpp/qwen3.8-27b/preset.ini +++ b/examples/llamacpp/qwen3.8-27b/preset.ini @@ -1,6 +1,6 @@ # A llama.cpp preset for Qwen3.8-27B, mirroring the llama-server command in # this directory's README. `spinloop serve` turns it into that command and runs it, -# whether that's on your own machine or on a `spinloop remote` GPU. +# whether that's on your own machine or on a `spinloop cloud` GPU. # See https://github.com/ggml-org/llama.cpp/blob/master/docs/preset.md [*] diff --git a/examples/mtplx/qwen3.8-27b/README.md b/examples/mtplx/qwen3.8-27b/README.md index c49e18ac..c24c792d 100644 --- a/examples/mtplx/qwen3.8-27b/README.md +++ b/examples/mtplx/qwen3.8-27b/README.md @@ -14,7 +14,7 @@ unified memory. - An **Apple Silicon** Mac (M1 or later). MTPLX is Apple Silicon only — there is no Intel or Linux build, and **no machine image**, so - [`spinloop remote`](../../../docs/commands/remote.md) cannot deploy it. It + [`spinloop cloud`](../../../docs/commands/cloud.md) cannot deploy it. It serves locally, or on a [fleet node](../../../docs/commands/fleet.md) you run yourself. - [MTPLX](https://mtplx.com), installed so that `mtplx` is on your `PATH`. diff --git a/examples/mtplx/qwen3.8-27b/Spinloop b/examples/mtplx/qwen3.8-27b/Spinloop index 1668afa0..207ff467 100644 --- a/examples/mtplx/qwen3.8-27b/Spinloop +++ b/examples/mtplx/qwen3.8-27b/Spinloop @@ -6,5 +6,5 @@ CONTEXT 32768 # --context-window; raise it on a bigger box PARALLEL 4 # --max-active-requests; the scheduler mode is in the preset PRESET ./preset.ini # `spinloop serve` launches `mtplx serve` from this # BASEURL http://127.0.0.1:8000/v1 # uncomment for a non-default host/port -# `spinloop remote` cannot deploy this: mtplx is Apple-Silicon-only and has no +# `spinloop cloud` cannot deploy this: mtplx is Apple-Silicon-only and has no # machine image to run it on a cloud GPU diff --git a/examples/omlx/gemma-4-e2b/README.md b/examples/omlx/gemma-4-e2b/README.md index 290ac9af..15800cbd 100644 --- a/examples/omlx/gemma-4-e2b/README.md +++ b/examples/omlx/gemma-4-e2b/README.md @@ -12,7 +12,7 @@ to prove the setup end to end before reaching for something larger (see the ## Prerequisites - An **Apple Silicon** Mac (M1 or later). oMLX is Apple Silicon only, so - [`spinloop remote`](../../../docs/commands/remote.md) cannot deploy it. + [`spinloop cloud`](../../../docs/commands/cloud.md) cannot deploy it. - [oMLX](https://omlx.ai), installed from its DMG or from source. - The 6-bit build is only a few GB, so memory is not a concern here. diff --git a/examples/omlx/qwen3.6/README.md b/examples/omlx/qwen3.6/README.md index ff70e415..18d66875 100644 --- a/examples/omlx/qwen3.6/README.md +++ b/examples/omlx/qwen3.6/README.md @@ -18,7 +18,7 @@ every turn, that is the difference between a usable and an unusable setup. - An **Apple Silicon** Mac (M1 or later). oMLX is Apple Silicon only — there is no Intel or Linux build, and no cloud equivalent, so - [`spinloop remote`](../../../docs/commands/remote.md) cannot deploy it. + [`spinloop cloud`](../../../docs/commands/cloud.md) cannot deploy it. - [oMLX](https://omlx.ai), installed from its DMG or from source. - Enough unified memory for the weights. The 4-bit build is roughly 20 GB, so plan for a 32 GB machine or better. diff --git a/examples/remote-spinloop/README.md b/examples/remote-spinloop/README.md index 8e0fe3c8..5e0e748c 100644 --- a/examples/remote-spinloop/README.md +++ b/examples/remote-spinloop/README.md @@ -57,7 +57,7 @@ works from any directory, on this machine. ## 4. Only `serve` fetches the preset `spinloop harness apply` never reads `PRESET` — a local one or a URL, it's `spinloop -serve`'s business alone. Only running `spinloop serve` (or `spinloop remote +serve`'s business alone. Only running `spinloop serve` (or `spinloop cloud deploy`) fetches `http://localhost:8000/preset.ini`: ```sh diff --git a/internal/catalog/catalog.go b/internal/catalog/catalog.go index b547b4e1..142e7134 100644 --- a/internal/catalog/catalog.go +++ b/internal/catalog/catalog.go @@ -242,12 +242,12 @@ func BuildProviderBlock(id string, p *Provider, modelOverride, baseURLOverride s return block, defaultModel, nil } -// RemoteProviderLabel is the display name for a harness provider that a remote +// CloudProviderLabel is the display name for a harness provider that a remote // environment has renamed, so it reads distinctly from a local engine of the // same kind: the engine's display name qualified by the environment, e.g. // "llama.cpp (dev-2)". With no engine name it is the environment alone, which is // still unique — no local provider shares it. -func RemoteProviderLabel(engine, env string) string { +func CloudProviderLabel(engine, env string) string { if engine == "" { return env } @@ -341,7 +341,7 @@ func BuildPiProvider(id string, p *Provider, modelOverride, baseURLOverride stri // remote one, and for the local server a reference to a variable set // nowhere would hide the models. That exception is deliberately conditioned // on the endpoint rather than only on the key, because a placeholder - // written for a remote endpoint could not be repaired by exporting the key + // written for a cloud endpoint could not be repaired by exporting the key // afterwards — Pi would keep sending the placeholder. switch { case p.APIKeyEnv != "" && p.APIKeyOptional && resolve(p.APIKeyEnv) == "" && IsLocalEndpoint(prov.BaseURL): diff --git a/internal/catalog/catalog_test.go b/internal/catalog/catalog_test.go index 1b6bb7f8..cde2623a 100644 --- a/internal/catalog/catalog_test.go +++ b/internal/catalog/catalog_test.go @@ -440,7 +440,7 @@ func TestBuildProviderBlock_NoModelIsNoDefault(t *testing.T) { } } -func TestRemoteProviderLabel(t *testing.T) { +func TestCloudProviderLabel(t *testing.T) { cases := []struct { name, engine, env, want string }{ @@ -450,8 +450,8 @@ func TestRemoteProviderLabel(t *testing.T) { } for _, tc := range cases { t.Run(tc.name, func(t *testing.T) { - if got := RemoteProviderLabel(tc.engine, tc.env); got != tc.want { - t.Errorf("RemoteProviderLabel(%q, %q) = %q, want %q", tc.engine, tc.env, got, tc.want) + if got := CloudProviderLabel(tc.engine, tc.env); got != tc.want { + t.Errorf("CloudProviderLabel(%q, %q) = %q, want %q", tc.engine, tc.env, got, tc.want) } }) } @@ -611,7 +611,7 @@ func TestBuildPiProvider_RequiredKeyKeepsReference(t *testing.T) { // The opencode block injects the resolved key, so a remote llama.cpp endpoint // is authenticated rather than 401ing. -func TestBuildProviderBlock_LlamacppReferencesKeyForRemote(t *testing.T) { +func TestBuildProviderBlock_LlamacppReferencesKeyForCloud(t *testing.T) { cat, _ := Load() resolve := func(name string) string { if name == "OPENAI_API_KEY" { @@ -637,14 +637,14 @@ func TestBuildProviderBlock_LlamacppReferencesKeyForRemote(t *testing.T) { t.Error("a local keyless server should get no apiKey at all") } - // A remote endpoint keeps the reference even with the key unset, so setting + // A cloud endpoint keeps the reference even with the key unset, so setting // it before the agent runs is enough. block, _, err = BuildProviderBlock("llamacpp", cat.Providers["llamacpp"], "local-model", "http://198.51.100.1:8000/v1", noEnv) if err != nil { t.Fatalf("BuildProviderBlock: %v", err) } if got := block["options"].(map[string]any)["apiKey"]; got != "{env:OPENAI_API_KEY}" { - t.Errorf("apiKey = %v, want the env reference for a remote endpoint", got) + t.Errorf("apiKey = %v, want the env reference for a cloud endpoint", got) } } @@ -733,7 +733,7 @@ func TestIsLocalEndpoint(t *testing.T) { } // Pi resolves its $VAR reference when it runs, so a placeholder written for a -// remote endpoint can never be repaired by exporting the key afterwards — Pi +// cloud endpoint can never be repaired by exporting the key afterwards — Pi // would keep sending the placeholder. The placeholder is therefore only right // for a local server. func TestBuildPiProvider_RemoteOptionalKeyKeepsReference(t *testing.T) { @@ -805,7 +805,7 @@ func TestBuildProviderBlock_OMLXReferencesKeyWhenSet(t *testing.T) { // TestBuildPiProvider_OMLXPlaceholderThenReference pins both sides of the Pi // rule for oMLX: a keyless local server gets the literal placeholder (Pi hides a -// provider's models until some auth is configured), while a remote endpoint gets +// provider's models until some auth is configured), while a cloud endpoint gets // the $VAR reference Pi resolves at run time. func TestBuildPiProvider_OMLXPlaceholderThenReference(t *testing.T) { cat, _ := Load() @@ -822,12 +822,12 @@ func TestBuildPiProvider_OMLXPlaceholderThenReference(t *testing.T) { t.Errorf("api = %q, want openai-completions", local.API) } - remote, _, err := BuildPiProvider("omlx", p, "my-model", "http://mac-studio.local:8000/v1", noEnv) + cloud, _, err := BuildPiProvider("omlx", p, "my-model", "http://mac-studio.local:8000/v1", noEnv) if err != nil { t.Fatalf("BuildPiProvider: %v", err) } - if remote.APIKey != "$OPENAI_API_KEY" { - t.Errorf("remote apiKey = %q, want $OPENAI_API_KEY", remote.APIKey) + if cloud.APIKey != "$OPENAI_API_KEY" { + t.Errorf("remote apiKey = %q, want $OPENAI_API_KEY", cloud.APIKey) } } @@ -863,7 +863,7 @@ func TestBuildPiProvider_HonoursPerProviderBaseURLEnv(t *testing.T) { // A provider with a key variable must get the reference Pi resolves at // run time, not the placeholder meant for keyless local servers. if p.APIKeyEnv != "" && prov.APIKey != "$"+p.APIKeyEnv { - t.Errorf("apiKey = %q, want $%s for a remote endpoint", prov.APIKey, p.APIKeyEnv) + t.Errorf("apiKey = %q, want $%s for a cloud endpoint", prov.APIKey, p.APIKeyEnv) } }) } diff --git a/internal/catalog/providers.yaml b/internal/catalog/providers.yaml index 41e5f5dc..b17f9ff6 100644 --- a/internal/catalog/providers.yaml +++ b/internal/catalog/providers.yaml @@ -112,8 +112,8 @@ providers: name: llama.cpp npm: "@ai-sdk/openai-compatible" # A local llama-server usually needs no key, and none is injected when the - # var is unset. But the same provider also names a remote deployment (see - # `spinloop remote deploy`), which is started with --api-key, so the key is + # var is unset. But the same provider also names a cloud deployment (see + # `spinloop cloud deploy`), which is started with --api-key, so the key is # picked up when there is one. apiKeyEnv: OPENAI_API_KEY apiKeyOptional: true diff --git a/internal/remote/aws.go b/internal/cloud/aws.go similarity index 96% rename from internal/remote/aws.go rename to internal/cloud/aws.go index 94103629..400e0005 100644 --- a/internal/remote/aws.go +++ b/internal/cloud/aws.go @@ -1,4 +1,4 @@ -package remote +package cloud import ( "context" @@ -22,7 +22,7 @@ import ( ) // LoadAWSConfig resolves the AWS config for a region, applying the credential -// precedence the remote commands sign with: explicit AWS environment +// precedence the cloud commands sign with: explicit AWS environment // credentials or an explicit profile selection win (the default chain, as // before); then a stored control-plane credential for the region, if one is // in the keystore; then the rest of the standard chain (shared config, SSO @@ -87,7 +87,7 @@ func CallerIdentity(ctx context.Context, cfg aws.Config) (string, error) { } // ControlPlaneUserName is the IAM user the control-plane stack creates for the -// CLI's long-lived credential: `spinloop remote auth --store` creates access +// CLI's long-lived credential: `spinloop cloud auth --store` creates access // keys for this user and stores one in this machine's keystore. The name is // fixed, so the CLI addresses the user without reading a stack output; the // stack and its tests use the same literal. @@ -155,7 +155,7 @@ func IAMDeleteAccessKey(ctx context.Context, cfg aws.Config, userName, accessKey } // ControlPlaneStackDeployed reports whether the named CloudFormation stack exists in -// the account and region — i.e. whether `spinloop remote bootstrap` has already +// the account and region — i.e. whether `spinloop cloud bootstrap` has already // run. A stack that does not exist is reported as false, not an error. func ControlPlaneStackDeployed(ctx context.Context, cfg aws.Config, stackName string) (bool, error) { _, err := cloudformation.NewFromConfig(cfg).DescribeStacks(ctx, &cloudformation.DescribeStacksInput{ @@ -170,7 +170,7 @@ func ControlPlaneStackDeployed(ctx context.Context, cfg aws.Config, stackName st return true, nil } -// ControlPlane is what `spinloop remote bootstrap` deployed once for the account: +// ControlPlane is what `spinloop cloud bootstrap` deployed once for the account: // the control URLs every environment shares, plus the weights bucket. It is // discovered from the control-plane stack's CloudFormation outputs, so it reflects // what is actually deployed and works from any machine with account access. @@ -188,7 +188,7 @@ func DiscoverControlPlane(ctx context.Context, cfg aws.Config, stackName string) if err != nil { if strings.Contains(err.Error(), "does not exist") { return ControlPlane{}, fmt.Errorf( - "the control plane (stack %q) is not deployed in this account and region — run `spinloop remote bootstrap` first", + "the control plane (stack %q) is not deployed in this account and region — run `spinloop cloud bootstrap` first", stackName) } return ControlPlane{}, err @@ -205,7 +205,7 @@ func DiscoverControlPlane(ctx context.Context, cfg aws.Config, stackName string) // controlPlaneFromOutputs maps a control-plane stack's CloudFormation outputs // onto the config that drives its Lambdas. Pure, so the mapping is testable // without a network: a stack output added to the template but not here would -// otherwise be dropped from every registered environment's remote.json. +// otherwise be dropped from every registered environment's cloud.json. func controlPlaneFromOutputs(stackName string, outputs map[string]string) (ControlPlane, error) { layer := ControlPlane{ Config: Config{ @@ -225,7 +225,7 @@ func controlPlaneFromOutputs(stackName string, outputs map[string]string) (Contr // existed; the subcommands that need them say so themselves. if layer.Config.StartURL == "" || layer.Config.StopURL == "" || layer.Config.DeployURL == "" { return ControlPlane{}, fmt.Errorf( - "stack %q is missing its control-URL outputs — re-run `spinloop remote bootstrap` to update it", + "stack %q is missing its control-URL outputs — re-run `spinloop cloud bootstrap` to update it", stackName) } return layer, nil diff --git a/internal/remote/aws_test.go b/internal/cloud/aws_test.go similarity index 98% rename from internal/remote/aws_test.go rename to internal/cloud/aws_test.go index 65dc9574..9f2738dd 100644 --- a/internal/remote/aws_test.go +++ b/internal/cloud/aws_test.go @@ -1,4 +1,4 @@ -package remote +package cloud import ( "context" @@ -14,8 +14,8 @@ func TestControlPlaneFromOutputs_MapsEveryStackOutput(t *testing.T) { // publishes for the config. If the template gains a new control-URL output, // the mapping in controlPlaneFromOutputs must take it on: an output landed // here but not mapped is dropped from every registered environment's - // remote.json without error — update_url was exactly that, which left - // `spinloop remote keep` unusable on freshly registered environments. + // cloud.json without error — update_url was exactly that, which left + // `spinloop cloud keep` unusable on freshly registered environments. outputs := map[string]string{ "StartUrl": "https://start.example.aws/", "StopUrl": "https://stop.example.aws/", diff --git a/internal/remote/bake.go b/internal/cloud/bake.go similarity index 93% rename from internal/remote/bake.go rename to internal/cloud/bake.go index 1e5c7a73..35e81045 100644 --- a/internal/remote/bake.go +++ b/internal/cloud/bake.go @@ -1,4 +1,4 @@ -package remote +package cloud import ( "context" @@ -11,7 +11,7 @@ import ( // BakedRunners reports, for the account's own images, which runners already // have a runtime AMI baked — read from the tags the Image Builder distribution // applies (`cloud-vm-llm:role=runtime-ami`, `cloud-vm-llm:runner=`). It -// lets `spinloop remote bake` tell when a bake has finished. +// lets `spinloop cloud bake` tell when a bake has finished. func BakedRunners(ctx context.Context, cfg aws.Config) (map[string]bool, error) { out, err := ec2.NewFromConfig(cfg).DescribeImages(ctx, &ec2.DescribeImagesInput{ Owners: []string{"self"}, diff --git a/internal/remote/remote.go b/internal/cloud/cloud.go similarity index 95% rename from internal/remote/remote.go rename to internal/cloud/cloud.go index 93dc9a49..c449c42e 100644 --- a/internal/remote/remote.go +++ b/internal/cloud/cloud.go @@ -1,8 +1,8 @@ -// Package remote controls the scale-to-zero GPU inference instance defined by +// Package cloud controls the scale-to-zero GPU inference instance defined by // this repository's remote/ subproject, by calling its Lambdas through their // Function URLs. The URLs use IAM auth, so every request is SigV4-signed // (service "lambda") with the caller's AWS credentials. -package remote +package cloud import ( "bytes" @@ -33,7 +33,7 @@ import ( // model into VRAM, which takes minutes. var httpClient = &http.Client{Timeout: 10 * time.Minute} -// Config holds the connection details for the remote instance's control +// Config holds the connection details for the cloud instance's control // Lambdas: deploying remote/ prints it as the SpinloopRemoteConfig output, ready // to paste into the config file. type Config struct { @@ -68,13 +68,13 @@ type Config struct { Environment string `json:"environment"` } -// LoadEnvironment reads a named environment's configuration: the remote.json -// in its registry directory, with the SPINLOOP_REMOTE_* overrides applied on +// LoadEnvironment reads a named environment's configuration: the cloud.json +// in its registry directory, with the SPINLOOP_CLOUD_* overrides applied on // top. It is the only way an environment is resolved — there is no path // outside the registry, and no name that resolves by a different rule. // // A name with no file is not a failure on its own: the overrides may carry a -// complete configuration, which is how the remote commands run on a machine +// complete configuration, which is how the cloud commands run on a machine // with nothing on disk. In that case the name is the environment identifier // the control calls carry, since a configuration assembled from variables has // no file to take one from. Where the overrides are incomplete too, @@ -104,33 +104,33 @@ func LoadEnvironment(name string, getenv func(string) string) (Config, error) { // finishConfig applies env overrides and validates. source names the config // file for error messages. func finishConfig(cfg Config, getenv func(string) string, source string) (Config, error) { - if v := getenv("SPINLOOP_REMOTE_START_URL"); v != "" { + if v := getenv("SPINLOOP_CLOUD_START_URL"); v != "" { cfg.StartURL = v } - if v := getenv("SPINLOOP_REMOTE_STOP_URL"); v != "" { + if v := getenv("SPINLOOP_CLOUD_STOP_URL"); v != "" { cfg.StopURL = v } - if v := getenv("SPINLOOP_REMOTE_DEPLOY_URL"); v != "" { + if v := getenv("SPINLOOP_CLOUD_DEPLOY_URL"); v != "" { cfg.DeployURL = v } - if v := getenv("SPINLOOP_REMOTE_STATS_URL"); v != "" { + if v := getenv("SPINLOOP_CLOUD_STATS_URL"); v != "" { cfg.StatsURL = v } - if v := getenv("SPINLOOP_REMOTE_ENV_URL"); v != "" { + if v := getenv("SPINLOOP_CLOUD_ENV_URL"); v != "" { cfg.EnvURL = v } - if v := getenv("SPINLOOP_REMOTE_SEED_URL"); v != "" { + if v := getenv("SPINLOOP_CLOUD_SEED_URL"); v != "" { cfg.SeedURL = v } - if v := getenv("SPINLOOP_REMOTE_UPDATE_URL"); v != "" { + if v := getenv("SPINLOOP_CLOUD_UPDATE_URL"); v != "" { cfg.UpdateURL = v } - if v := getenv("SPINLOOP_REMOTE_REGION"); v != "" { + if v := getenv("SPINLOOP_CLOUD_REGION"); v != "" { cfg.Region = v } if cfg.StartURL == "" || cfg.StopURL == "" { return Config{}, fmt.Errorf( - "remote is not configured: paste the SpinloopRemoteConfig output of the remote/ deployment into %s", + "cloud is not configured: paste the SpinloopRemoteConfig output of the remote/ deployment into %s", source) } if cfg.Region == "" { @@ -141,7 +141,7 @@ func finishConfig(cfg Config, getenv func(string) string, source string) (Config } if cfg.Region == "" { return Config{}, fmt.Errorf( - "cannot determine the AWS region: set \"region\" in %s or SPINLOOP_REMOTE_REGION", + "cannot determine the AWS region: set \"region\" in %s or SPINLOOP_CLOUD_REGION", source) } return cfg, nil @@ -187,7 +187,7 @@ type Response struct { Deployed bool `json:"deployed"` Seeding bool `json:"seeding"` // SeedID identifies the seed a deploy started, so it can be followed with - // `spinloop remote seed status`. The instance id it replaces was an + // `spinloop cloud seed status`. The instance id it replaces was an // implementation detail that changes if the seed is relaunched. SeedID string `json:"seedId"` // ServedName is the name the engine answers to beside the model id — the @@ -247,7 +247,7 @@ func IsInstanceType(value string) bool { func Deploy(ctx context.Context, cfg Config, dc inference.DeployConfig, allowedCidr string, reseed bool, apiKey string) (*Response, error) { if cfg.DeployURL == "" { return nil, fmt.Errorf( - "no deploy_url configured: add the remote/ deployment's DeployUrl output to the remote config (or set SPINLOOP_REMOTE_DEPLOY_URL)") + "no deploy_url configured: add the remote/ deployment's DeployUrl output to the cloud config (or set SPINLOOP_CLOUD_DEPLOY_URL)") } body, err := json.Marshal(struct { inference.DeployConfig @@ -301,7 +301,7 @@ const stateSeeding = "seeding" // following the seed and resuming the wait, so the error carries both. func giveUpWaiting(ctx context.Context, state, seedID string) error { if state == stateSeeding && seedID != "" { - return fmt.Errorf("gave up waiting for the endpoint: the weights are still seeding (seed %s) — follow it with `spinloop remote seed status %s`, and re-run start with a longer --timeout: %w", + return fmt.Errorf("gave up waiting for the endpoint: the weights are still seeding (seed %s) — follow it with `spinloop cloud seed status %s`, and re-run start with a longer --timeout: %w", seedID, seedID, ctx.Err()) } return fmt.Errorf("gave up waiting for the endpoint: %w", ctx.Err()) @@ -456,7 +456,7 @@ func Restart(ctx context.Context, cfg Config, force bool, progress func(string), progress("stopped; waking it") resp, err := Start(ctx, cfg, progress, onState, nil) if err != nil { - return nil, fmt.Errorf("%w — the instance is stopped; `spinloop remote start` will bring it back", err) + return nil, fmt.Errorf("%w — the instance is stopped; `spinloop cloud start` will bring it back", err) } return resp, nil } @@ -487,7 +487,7 @@ func pauseURL(stopURL string, force bool) string { func Keep(ctx context.Context, cfg Config, retainUntil time.Time) (*Response, error) { if cfg.UpdateURL == "" { return nil, fmt.Errorf( - "no update_url configured: the remote deployment needs to be updated for keep support") + "no update_url configured: the cloud deployment needs to be updated for keep support") } u, err := url.Parse(cfg.UpdateURL) if err != nil { @@ -514,7 +514,7 @@ func Keep(ctx context.Context, cfg Config, retainUntil time.Time) (*Response, er func Env(ctx context.Context, cfg Config) (*Response, error) { if cfg.EnvURL == "" { return nil, fmt.Errorf( - "no env_url configured: the remote deployment needs to be updated for env support") + "no env_url configured: the cloud deployment needs to be updated for env support") } resp, err := call(ctx, cfg, http.MethodGet, cfg.EnvURL, nil) if err != nil { @@ -636,12 +636,12 @@ const refreshCredsHint = "refresh your env credentials, profile, or SSO session" // storedCredsHint is the refresh guidance when the stored control-plane key // was the credential in use: refreshing the ambient credentials would not // change what signs the request, so the fix is to store a new key. -const storedCredsHint = "run `spinloop remote auth --store` to create a new stored key" +const storedCredsHint = "run `spinloop cloud auth --store` to create a new stored key" // credsRefreshHint picks the refresh guidance for a rejected request by the // source of the credential that signed it: explicit ambient credentials win // over the stored key in resolution, so they name themselves; a stored key in -// use names `spinloop remote auth --store`; anything else is ambient. +// use names `spinloop cloud auth --store`; anything else is ambient. func credsRefreshHint(region string) string { if _, ok := LookupStoredCredential(region); ok && !explicitAmbientCreds() { return storedCredsHint @@ -788,7 +788,7 @@ type ( func Stats(ctx context.Context, cfg Config) (*StatsResponse, error) { if cfg.StatsURL == "" { return nil, fmt.Errorf( - "no stats_url configured: the control plane needs re-deploying with `pnpm run deploy` (or set SPINLOOP_REMOTE_STATS_URL)") + "no stats_url configured: the control plane needs re-deploying with `pnpm run deploy` (or set SPINLOOP_CLOUD_STATS_URL)") } out, err := callStats(ctx, cfg) if err != nil { diff --git a/internal/remote/remote_test.go b/internal/cloud/cloud_test.go similarity index 99% rename from internal/remote/remote_test.go rename to internal/cloud/cloud_test.go index 5f7ee050..aba06091 100644 --- a/internal/remote/remote_test.go +++ b/internal/cloud/cloud_test.go @@ -1,4 +1,4 @@ -package remote +package cloud import ( "context" @@ -85,8 +85,8 @@ func TestLoadConfig_EnvOverrides(t *testing.T) { isolateConfig(t) writeConfig(t, Config{StartURL: "https://old/", StopURL: "https://old-stop/", Region: "us-east-1"}) cfg, err := LoadEnvironment("default", envMap(map[string]string{ - "SPINLOOP_REMOTE_START_URL": "https://new/", - "SPINLOOP_REMOTE_REGION": "eu-west-2", + "SPINLOOP_CLOUD_START_URL": "https://new/", + "SPINLOOP_CLOUD_REGION": "eu-west-2", })) if err != nil { t.Fatal(err) @@ -245,7 +245,7 @@ func TestStart_GiveUpDuringSeedingNamesTheSeed(t *testing.T) { if err == nil { t.Fatal("expected an error when the deadline expires") } - for _, want := range []string{"seeding", "llamacpp--org-model--Q4_K_M", "spinloop remote seed status", "--timeout"} { + for _, want := range []string{"seeding", "llamacpp--org-model--Q4_K_M", "spinloop cloud seed status", "--timeout"} { if !strings.Contains(err.Error(), want) { t.Errorf("give-up error does not name %q: %v", want, err) } @@ -285,7 +285,7 @@ func TestStart_GiveUpAfterADroppedConnectionNamesTheSeed(t *testing.T) { if err == nil { t.Fatal("expected an error once the deadline passed") } - for _, want := range []string{"seeding", "llamacpp--org-model--Q4_K_M", "spinloop remote seed status"} { + for _, want := range []string{"seeding", "llamacpp--org-model--Q4_K_M", "spinloop cloud seed status"} { if !strings.Contains(err.Error(), want) { t.Errorf("give-up error does not name %q: %v", want, err) } @@ -754,7 +754,7 @@ func TestRestart_WakeFailureNamesRecovery(t *testing.T) { t.Fatal("expected a wake failure") } if !strings.Contains(err.Error(), "stopped") || - !strings.Contains(err.Error(), "spinloop remote start") { + !strings.Contains(err.Error(), "spinloop cloud start") { t.Errorf("expected the recovery hint in the error, got %v", err) } if len(calls) != 2 || calls[0] != "stop" || calls[1] != "wake" { @@ -1197,7 +1197,7 @@ func TestLoadEnvironment(t *testing.T) { t.Errorf("unexpected config: %+v", cfg) } - cfg, err = LoadEnvironment("prod", envMap(map[string]string{"SPINLOOP_REMOTE_REGION": "eu-west-2"})) + cfg, err = LoadEnvironment("prod", envMap(map[string]string{"SPINLOOP_CLOUD_REGION": "eu-west-2"})) if err != nil { t.Fatal(err) } @@ -1214,7 +1214,7 @@ func TestLoadEnvironment_Missing(t *testing.T) { if err == nil { t.Fatal("expected a not-configured error") } - if !strings.Contains(err.Error(), "remotes/nosuchenv/remote.json") { + if !strings.Contains(err.Error(), "clouds/nosuchenv/cloud.json") { t.Errorf("the failure should name the environment's registry path, got %v", err) } } @@ -1392,7 +1392,7 @@ func TestLoadConfig_EnvURLOverride(t *testing.T) { isolateConfig(t) writeConfig(t, Config{StartURL: "https://start/", StopURL: "https://stop/", EnvURL: "https://old-env/", Region: "eu-west-1"}) cfg, err := LoadEnvironment("default", envMap(map[string]string{ - "SPINLOOP_REMOTE_ENV_URL": "https://new-env/", + "SPINLOOP_CLOUD_ENV_URL": "https://new-env/", })) if err != nil { t.Fatal(err) diff --git a/internal/remote/environments.go b/internal/cloud/environments.go similarity index 79% rename from internal/remote/environments.go rename to internal/cloud/environments.go index a7df9435..ab6d1bf5 100644 --- a/internal/remote/environments.go +++ b/internal/cloud/environments.go @@ -1,4 +1,4 @@ -package remote +package cloud import ( "encoding/json" @@ -10,7 +10,7 @@ import ( ) // ConfigHome returns spinloop's own config directory, where both the legacy -// remote.json and the environments registry live. It delegates to +// cloud.json and the environments registry live. It delegates to // internal/config.Dir, so the SPINLOOP_CONFIG_DIR override and the fallback // rules are resolved in one place; it fails when the directory cannot be // determined (see config.Dir). @@ -19,17 +19,17 @@ func ConfigHome() (string, error) { } // remotesRoot is the environments registry directory: one subdirectory per -// named environment, each holding a remote.json. +// named environment, each holding a cloud.json. func remotesRoot() (string, error) { home, err := ConfigHome() if err != nil { return "", err } - return filepath.Join(home, "remotes"), nil + return filepath.Join(home, "clouds"), nil } -// EnvDir returns an environment's directory, /remotes/. A -// remote deployment's state (currently just remote.json) lives here, keyed by +// EnvDir returns an environment's directory, /clouds/. A +// cloud deployment's state (currently just cloud.json) lives here, keyed by // name so several instances never share a file. func EnvDir(name string) (string, error) { root, err := remotesRoot() @@ -39,13 +39,13 @@ func EnvDir(name string) (string, error) { return filepath.Join(root, name), nil } -// EnvConfigPath returns the remote.json inside an environment's directory. +// EnvConfigPath returns the cloud.json inside an environment's directory. func EnvConfigPath(name string) (string, error) { dir, err := EnvDir(name) if err != nil { return "", err } - return filepath.Join(dir, "remote.json"), nil + return filepath.Join(dir, "cloud.json"), nil } // IsEnvName reports whether a value is a plain environment name: non-empty, @@ -66,7 +66,7 @@ func IsEnvName(value string) bool { } // EnvInfo describes one registered environment for listing. OK is false when -// the environment's remote.json is missing or unreadable. +// the environment's cloud.json is missing or unreadable. type EnvInfo struct { Name string BaseURL string @@ -76,7 +76,7 @@ type EnvInfo struct { // ListEnvironments returns the registered environments, sorted by name. An // absent registry is not an error — it yields no environments. Each entry's -// remote.json is read best-effort: a directory without a readable one is still +// cloud.json is read best-effort: a directory without a readable one is still // listed, with OK false, rather than failing the whole listing. func ListEnvironments() ([]EnvInfo, error) { root, err := remotesRoot() @@ -111,7 +111,7 @@ func ListEnvironments() ([]EnvInfo, error) { return envs, nil } -// SaveEnvironment registers a deployed environment: its remote.json (the +// SaveEnvironment registers a deployed environment: its cloud.json (the // shared control URLs, region, base URL, and the environment identifier) is // written under the registry, owner-only, since it names a deployment's URLs // and address. Registering a second environment never touches the first. @@ -127,5 +127,5 @@ func SaveEnvironment(name string, cfg Config) error { if err := os.MkdirAll(dir, 0o700); err != nil { return err } - return os.WriteFile(filepath.Join(dir, "remote.json"), append(data, '\n'), 0o600) + return os.WriteFile(filepath.Join(dir, "cloud.json"), append(data, '\n'), 0o600) } diff --git a/internal/remote/environments_test.go b/internal/cloud/environments_test.go similarity index 88% rename from internal/remote/environments_test.go rename to internal/cloud/environments_test.go index 9e177687..65f3370e 100644 --- a/internal/remote/environments_test.go +++ b/internal/cloud/environments_test.go @@ -1,4 +1,4 @@ -package remote +package cloud import ( "os" @@ -12,10 +12,10 @@ func TestIsEnvName(t *testing.T) { "qwen3.6-27b-prod": true, "default": true, "qwen3.6": true, // a dot that is not a .json suffix - "./remote.json": false, - "/abs/remote.json": false, - "remotes/x": false, - "remote.json": false, + "./cloud.json": false, + "/abs/cloud.json": false, + "clouds/x": false, + "cloud.json": false, `win\path`: false, "": false, } @@ -29,13 +29,13 @@ func TestIsEnvName(t *testing.T) { func TestEnvConfigPath(t *testing.T) { home := t.TempDir() t.Setenv("XDG_CONFIG_HOME", home) - want := filepath.Join(home, "spinloop", "remotes", "prod", "remote.json") + want := filepath.Join(home, "spinloop", "clouds", "prod", "cloud.json") if got := must1(EnvConfigPath("prod")); got != want { t.Errorf("EnvConfigPath = %q, want %q", got, want) } } -// writeEnv registers an environment's remote.json for a test. +// writeEnv registers an environment's cloud.json for a test. func writeEnv(t *testing.T, name, body string) { t.Helper() if err := os.MkdirAll(must1(EnvDir(name)), 0o700); err != nil { @@ -55,7 +55,7 @@ func TestListEnvironments(t *testing.T) { } writeEnv(t, "prod", `{"start_url":"https://s","stop_url":"https://x","region":"eu-west-1","base_url":"http://1.2.3.4:8000/v1"}`) - // A directory with an unreadable/invalid remote.json is listed, not fatal. + // A directory with an unreadable/invalid cloud.json is listed, not fatal. if err := os.MkdirAll(must1(EnvDir("broken")), 0o700); err != nil { t.Fatal(err) } @@ -94,7 +94,7 @@ func TestSaveEnvironment(t *testing.T) { } // Owner-only: the file names a deployment's URLs and address. if fi.Mode().Perm() != 0o600 { - t.Errorf("remote.json mode = %v, want 0600", fi.Mode().Perm()) + t.Errorf("cloud.json mode = %v, want 0600", fi.Mode().Perm()) } // Round-trips through the loader, environment identifier included. got, err := LoadEnvironment("prod", func(string) string { return "" }) @@ -132,7 +132,7 @@ func TestLoadEnvironment_ByName(t *testing.T) { t.Run("a file at the superseded path is not read", func(t *testing.T) { home := t.TempDir() t.Setenv("XDG_CONFIG_HOME", home) - legacy := filepath.Join(home, "spinloop", "remote.json") + legacy := filepath.Join(home, "spinloop", "cloud.json") if err := os.MkdirAll(filepath.Dir(legacy), 0o700); err != nil { t.Fatal(err) } @@ -150,7 +150,7 @@ func TestLoadEnvironment_ByName(t *testing.T) { if err == nil { t.Fatal("expected an error naming the environment's path") } - if !strings.Contains(err.Error(), "remotes/default/remote.json") { + if !strings.Contains(err.Error(), "clouds/default/cloud.json") { t.Errorf("the failure should name the registry path, got %v", err) } }) @@ -160,9 +160,9 @@ func TestLoadEnvironment_ByName(t *testing.T) { t.Run("overrides configure a named environment with no file", func(t *testing.T) { t.Setenv("XDG_CONFIG_HOME", t.TempDir()) got, err := LoadEnvironment("ci", envMap(map[string]string{ - "SPINLOOP_REMOTE_START_URL": "https://s", - "SPINLOOP_REMOTE_STOP_URL": "https://x", - "SPINLOOP_REMOTE_REGION": "eu-west-1", + "SPINLOOP_CLOUD_START_URL": "https://s", + "SPINLOOP_CLOUD_STOP_URL": "https://x", + "SPINLOOP_CLOUD_REGION": "eu-west-1", })) if err != nil { t.Fatalf("overrides should configure it: %v", err) diff --git a/internal/remote/follow.go b/internal/cloud/follow.go similarity index 94% rename from internal/remote/follow.go rename to internal/cloud/follow.go index 86571ca5..2581b828 100644 --- a/internal/remote/follow.go +++ b/internal/cloud/follow.go @@ -1,4 +1,4 @@ -package remote +package cloud import ( "sync" @@ -9,7 +9,7 @@ import ( // re-asks from by default. The shipping agent's delivery lag means an event // can land with a timestamp slightly behind one already returned, so a poll // deliberately re-reads a little; FollowCursor suppresses what it has -// already returned, by event id. Both `spinloop remote logs -f` and a fleet +// already returned, by event id. Both `spinloop cloud logs -f` and a fleet // node's own log follow share this constant, so a fleet-dashboard poll and a // standalone follow tolerate the same shipping lag. const FollowOverlap = 10 * time.Second @@ -20,7 +20,7 @@ const FollowOverlap = 10 * time.Second // has to re-ask a little behind the newest event it has already seen, to // catch anything the shipping agent delivered late, and then suppress by // event id whatever that overlap re-reads. One cursor holds that state for -// the life of one follow, whether that follow is `spinloop remote logs -f` +// the life of one follow, whether that follow is `spinloop cloud logs -f` // polling in a loop or a fleet node answering repeated Logs calls. type FollowCursor struct { overlap time.Duration diff --git a/internal/remote/follow_test.go b/internal/cloud/follow_test.go similarity index 99% rename from internal/remote/follow_test.go rename to internal/cloud/follow_test.go index 3b76317d..2530ab70 100644 --- a/internal/remote/follow_test.go +++ b/internal/cloud/follow_test.go @@ -1,4 +1,4 @@ -package remote +package cloud import ( "testing" diff --git a/internal/remote/keystore.go b/internal/cloud/keystore.go similarity index 96% rename from internal/remote/keystore.go rename to internal/cloud/keystore.go index 936de62a..a943edf5 100644 --- a/internal/remote/keystore.go +++ b/internal/cloud/keystore.go @@ -1,4 +1,4 @@ -package remote +package cloud import ( "encoding/json" @@ -16,7 +16,7 @@ import ( ) // StoredCredential is one long-lived control-plane credential the operator -// stored with `spinloop remote auth --store`: an access key for the +// stored with `spinloop cloud auth --store`: an access key for the // control-plane user, held in the OS keystore (or an owner-only file where no // keystore exists) rather than in a shared AWS config. type StoredCredential struct { @@ -51,7 +51,7 @@ const keyringIndexUser = "index" // "file": the opt-out for a headless macOS session whose keychain is locked // or unreachable, and the way the test suite keeps stored credentials inside // a temp config directory. -const keyStoreEnvVar = "SPINLOOP_REMOTE_KEYSTORE" +const keyStoreEnvVar = "SPINLOOP_CLOUD_KEYSTORE" // keyringBackend is the slice of the OS keystore the store drives, so tests // substitute an in-memory fake without touching the machine's real keystore. @@ -95,7 +95,7 @@ func keyringAvailable() bool { // openCredStore opens the machine's credential store: the OS keystore where // one is available, otherwise the owner-only file store. A machine with no -// keystore at all is the fallback case; SPINLOOP_REMOTE_KEYSTORE=file chooses +// keystore at all is the fallback case; SPINLOOP_CLOUD_KEYSTORE=file chooses // the file store even where a keystore is reachable; a keystore that opened // and then fails is reported, not papered over by silently switching stores. func openCredStore() (credentialStore, error) { @@ -118,7 +118,7 @@ func openCredStore() (credentialStore, error) { var openCredStoreFn = openCredStore func (s credentialStore) fileFor(region string) string { - return filepath.Join(s.dir, "remote-"+region+".json") + return filepath.Join(s.dir, "cloud-"+region+".json") } func (s credentialStore) put(region string, cred StoredCredential) error { @@ -244,10 +244,10 @@ func (s credentialStore) list() ([]StoredCredential, error) { } for _, e := range entries { name := e.Name() - if e.IsDir() || filepath.Ext(name) != ".json" || !strings.HasPrefix(name, "remote-") { + if e.IsDir() || filepath.Ext(name) != ".json" || !strings.HasPrefix(name, "cloud-") { continue } - region := name[len("remote-") : len(name)-len(".json")] + region := name[len("cloud-") : len(name)-len(".json")] if region != "" { regions = append(regions, region) } diff --git a/internal/remote/keystore_test.go b/internal/cloud/keystore_test.go similarity index 99% rename from internal/remote/keystore_test.go rename to internal/cloud/keystore_test.go index 121a1aef..31cbf00c 100644 --- a/internal/remote/keystore_test.go +++ b/internal/cloud/keystore_test.go @@ -1,4 +1,4 @@ -package remote +package cloud import ( "os" @@ -233,7 +233,7 @@ func TestKeyStoreEnvVarForcesFileStore(t *testing.T) { if err := StoreCredential(testCred("us-east-1")); err != nil { t.Fatalf("StoreCredential: %v", err) } - if _, err := os.Stat(filepath.Join(home, ".config", "spinloop", "keystore", "remote-us-east-1.json")); err != nil { + if _, err := os.Stat(filepath.Join(home, ".config", "spinloop", "keystore", "cloud-us-east-1.json")); err != nil { t.Fatalf("file store not used under the forced config dir: %v", err) } got, ok := LookupStoredCredential("us-east-1") diff --git a/internal/remote/logs.go b/internal/cloud/logs.go similarity index 97% rename from internal/remote/logs.go rename to internal/cloud/logs.go index 4ede7998..cb4d61d5 100644 --- a/internal/remote/logs.go +++ b/internal/cloud/logs.go @@ -1,4 +1,4 @@ -package remote +package cloud import ( "context" @@ -139,8 +139,8 @@ func logGroupsFor(source string) ([]logGroup, error) { func FetchLogs(ctx context.Context, cfg Config, q LogQuery) (LogResult, error) { if q.Environment == "" { return LogResult{}, fmt.Errorf( - "this remote config names no environment, so its log streams cannot be identified: " + - "re-register it with `spinloop remote deploy` (which writes the environment name)") + "this cloud config names no environment, so its log streams cannot be identified: " + + "re-register it with `spinloop cloud deploy` (which writes the environment name)") } awsCfg, err := LoadAWSConfig(ctx, cfg.Region) if err != nil { @@ -180,7 +180,7 @@ func fetchLogs(ctx context.Context, api logsAPI, q LogQuery, region string) (Log if missing == len(groups) { return LogResult{}, fmt.Errorf( "no log group exists for this environment's %s logs (looked for %s): "+ - "the control plane was deployed before log shipping — re-deploy it with `spinloop remote bootstrap`", + "the control plane was deployed before log shipping — re-deploy it with `spinloop cloud bootstrap`", q.Source, strings.Join(groupNames(groups), ", ")) } sortEvents(events) diff --git a/internal/remote/logs_test.go b/internal/cloud/logs_test.go similarity index 98% rename from internal/remote/logs_test.go rename to internal/cloud/logs_test.go index 2873d578..5ac52b79 100644 --- a/internal/remote/logs_test.go +++ b/internal/cloud/logs_test.go @@ -1,4 +1,4 @@ -package remote +package cloud import ( "context" @@ -262,7 +262,7 @@ func TestFetchLogsReportsWhenNoGroupExistsAtAll(t *testing.T) { if err == nil { t.Fatal("every group missing should be an error, not an empty result") } - if !strings.Contains(err.Error(), "spinloop remote bootstrap") { + if !strings.Contains(err.Error(), "spinloop cloud bootstrap") { t.Errorf("error = %q, want it to name the fix", err) } } @@ -359,7 +359,7 @@ func TestFetchLogsRequiresAnEnvironmentName(t *testing.T) { if err == nil { t.Fatal("a config with no environment cannot identify its streams") } - if !strings.Contains(err.Error(), "spinloop remote deploy") { + if !strings.Contains(err.Error(), "spinloop cloud deploy") { t.Errorf("error = %q, want it to say how to re-register", err) } } diff --git a/internal/remote/pathhelper_test.go b/internal/cloud/pathhelper_test.go similarity index 94% rename from internal/remote/pathhelper_test.go rename to internal/cloud/pathhelper_test.go index a0f15b0c..8752bd37 100644 --- a/internal/remote/pathhelper_test.go +++ b/internal/cloud/pathhelper_test.go @@ -1,4 +1,4 @@ -package remote +package cloud // must1 unwraps a (value, error) path resolver in tests, failing on error. // The config-dir resolvers now return an error; a test that only needs the diff --git a/internal/remote/seed.go b/internal/cloud/seed.go similarity index 98% rename from internal/remote/seed.go rename to internal/cloud/seed.go index 0d5ab10c..87dd3cd5 100644 --- a/internal/remote/seed.go +++ b/internal/cloud/seed.go @@ -1,4 +1,4 @@ -package remote +package cloud import ( "context" @@ -96,7 +96,7 @@ type SeedStopped struct { func seedURL(cfg Config) (string, error) { if cfg.SeedURL == "" { return "", fmt.Errorf( - "no seed_url configured: add the remote/ deployment's SeedUrl output to the remote config (or set SPINLOOP_REMOTE_SEED_URL)") + "no seed_url configured: add the remote/ deployment's SeedUrl output to the cloud config (or set SPINLOOP_CLOUD_SEED_URL)") } return cfg.SeedURL, nil } diff --git a/internal/remote/seed_test.go b/internal/cloud/seed_test.go similarity index 98% rename from internal/remote/seed_test.go rename to internal/cloud/seed_test.go index fe71b0f2..b6180367 100644 --- a/internal/remote/seed_test.go +++ b/internal/cloud/seed_test.go @@ -1,4 +1,4 @@ -package remote +package cloud import ( "context" @@ -228,7 +228,7 @@ func TestSeedStop_NothingRunningIsNotAnError(t *testing.T) { func TestSeedCalls_NameTheMissingSeedURL(t *testing.T) { stubAWSEnv(t) - // A remote config written before the seed Lambda existed. + // A cloud config written before the seed Lambda existed. cfg := Config{StartURL: "http://x", StopURL: "http://x", Region: "eu-west-1"} ctx := context.Background() @@ -245,7 +245,7 @@ func TestSeedCalls_NameTheMissingSeedURL(t *testing.T) { continue } // The message must say what to add and where it comes from. - for _, want := range []string{"seed_url", "SeedUrl", "SPINLOOP_REMOTE_SEED_URL"} { + for _, want := range []string{"seed_url", "SeedUrl", "SPINLOOP_CLOUD_SEED_URL"} { if !strings.Contains(err.Error(), want) { t.Errorf("%s: the error should name %q, got: %v", name, want, err) } @@ -309,7 +309,7 @@ func TestConfig_SeedURLOverride(t *testing.T) { isolateConfig(t) writeConfig(t, Config{StartURL: "http://start", StopURL: "http://stop", Region: "eu-west-1"}) cfg, err := LoadEnvironment("default", func(k string) string { - if k == "SPINLOOP_REMOTE_SEED_URL" { + if k == "SPINLOOP_CLOUD_SEED_URL" { return "http://override" } return "" @@ -324,7 +324,7 @@ func TestConfig_SeedURLOverride(t *testing.T) { // The regression this guards: SeedUrl was added as a stack output and to // Config, but the discovery that maps outputs onto Config did not read it — so -// `spinloop remote seed` reported "no seed_url configured" against every real +// `spinloop cloud seed` reported "no seed_url configured" against every real // deployment while the stubbed CLI tests passed. func TestControlPlaneFromOutputs_CarriesEveryURL(t *testing.T) { layer, err := controlPlaneFromOutputs("cloud-vm-llm", map[string]string{ diff --git a/internal/remote/source.go b/internal/cloud/source.go similarity index 98% rename from internal/remote/source.go rename to internal/cloud/source.go index f433105f..737e0b43 100644 --- a/internal/remote/source.go +++ b/internal/cloud/source.go @@ -1,4 +1,4 @@ -package remote +package cloud import ( "archive/tar" @@ -50,7 +50,7 @@ func ResolveRef(version, override string) string { } // SourceRoot is the parent of the ref-keyed CDK source caches, -// /cdk. It is named cdk/ to avoid confusion with the remotes/ +// /cdk. It is named cdk/ to avoid confusion with the clouds/ // environment registry. func SourceRoot() (string, error) { home, err := ConfigHome() @@ -82,7 +82,7 @@ func skipSource(rel string) bool { return true } switch rel { - case ".env", "remote.json", "cdk-outputs.json": + case ".env", "remote.json", "cloud.json", "cdk-outputs.json": return true } return false diff --git a/internal/remote/source_test.go b/internal/cloud/source_test.go similarity index 99% rename from internal/remote/source_test.go rename to internal/cloud/source_test.go index b11a1f81..2a89bc3d 100644 --- a/internal/remote/source_test.go +++ b/internal/cloud/source_test.go @@ -1,4 +1,4 @@ -package remote +package cloud import ( "archive/tar" diff --git a/internal/remote/versioncheck.go b/internal/cloud/versioncheck.go similarity index 95% rename from internal/remote/versioncheck.go rename to internal/cloud/versioncheck.go index 4636a901..9a06131b 100644 --- a/internal/remote/versioncheck.go +++ b/internal/cloud/versioncheck.go @@ -1,4 +1,4 @@ -package remote +package cloud import ( "fmt" @@ -52,7 +52,7 @@ func checkControlPlaneVersion(h http.Header) { } versionWarnOnce.Do(func() { fmt.Fprintf(versionWarnWriter, - "Warning: the control plane is at version %s but this spinloop is %s. Run `spinloop remote bootstrap` to bring it up to date.\n", + "Warning: the control plane is at version %s but this spinloop is %s. Run `spinloop cloud bootstrap` to bring it up to date.\n", strings.TrimPrefix(got, "v"), strings.TrimPrefix(cliVersion, "v")) }) } diff --git a/internal/remote/versioncheck_test.go b/internal/cloud/versioncheck_test.go similarity index 97% rename from internal/remote/versioncheck_test.go rename to internal/cloud/versioncheck_test.go index 9ab52693..ca6f91c5 100644 --- a/internal/remote/versioncheck_test.go +++ b/internal/cloud/versioncheck_test.go @@ -1,4 +1,4 @@ -package remote +package cloud import ( "bytes" @@ -65,7 +65,7 @@ func TestCheckControlPlaneVersion_WarnsOnceOnMismatch(t *testing.T) { if strings.Count(out, "Warning:") != 1 { t.Errorf("expected exactly one warning, got %q", out) } - for _, want := range []string{"1.28.0", "1.30.0", "spinloop remote bootstrap"} { + for _, want := range []string{"1.28.0", "1.30.0", "spinloop cloud bootstrap"} { if !strings.Contains(out, want) { t.Errorf("warning %q does not mention %q", out, want) } diff --git a/internal/config/config.go b/internal/config/config.go index 8f0ad06f..bb6ddb34 100644 --- a/internal/config/config.go +++ b/internal/config/config.go @@ -35,7 +35,7 @@ const DirEnvVar = "SPINLOOP_CONFIG_DIR" // Dir returns spinloop's config directory — the single root every file spinloop // owns resolves under (this package's config.json, and, via -// internal/remote.ConfigHome, remote.json, the environment registry, the +// internal/cloud.ConfigHome, cloud.json, the environment registry, the // daemon state dir and the CDK source cache). Resolution order: // // 1. SPINLOOP_CONFIG_DIR, used verbatim; diff --git a/internal/daemon/daemon.go b/internal/daemon/daemon.go index a37e74c9..0bf11c1c 100644 --- a/internal/daemon/daemon.go +++ b/internal/daemon/daemon.go @@ -320,13 +320,13 @@ type StatusResponse struct { // binds 127.0.0.1:8080, which is useless to anyone else, and it cannot know // the name a client reaches this host by — a LAN name, a tailscale name, a // published container port. The caller composes these against the host it -// already has. A node that does know that name — a remote environment, whose +// already has. A node that does know that name — a cloud environment, whose // control plane publishes the instance's address — reports it in Host, and the // caller uses it in place of the host it would otherwise supply. type EngineEndpoint struct { // Host is the name or address a client reaches the engine by, when the // node knows it. A daemon leaves it empty — it cannot know a client-facing - // name — but a remote environment's status fills it with the instance's + // name — but a cloud environment's status fills it with the instance's // published address, which is all a caller needs. Host string `json:"host,omitempty"` // Port is the port the engine listens on — the engine's, never the diff --git a/internal/daemon/daemon_test.go b/internal/daemon/daemon_test.go index ee61f98c..87fd1526 100644 --- a/internal/daemon/daemon_test.go +++ b/internal/daemon/daemon_test.go @@ -434,7 +434,7 @@ while true; do sleep 0.05; done`) // The activity pair crosses the wire, not just the Go call: this is the // shape the stats Lambda curls, so a field that never serialised would - // leave `spinloop remote metrics` silently blank. + // leave `spinloop cloud metrics` silently blank. _, metricsBody := do("GET", "/v1/metrics", "sekrit", "") _, statusBody := do("GET", "/v1/status", "sekrit", "") if metricsBody["lastActiveAt"] == nil { diff --git a/internal/fleet/capabilities_test.go b/internal/fleet/capabilities_test.go index a3c0dacb..414ec2d4 100644 --- a/internal/fleet/capabilities_test.go +++ b/internal/fleet/capabilities_test.go @@ -25,20 +25,20 @@ func countingStatsServer(t *testing.T, body string, calls *int) string { } // registerStatsEnv registers an environment whose stats call is the given URL, -// which registerRemoteEnv does not set — the capabilities answer from a stats +// which registerCloudEnv does not set — the capabilities answer from a stats // reading, so a test of them needs one. func registerStatsEnv(t *testing.T, name, url string) { t.Helper() home := t.TempDir() t.Setenv("SPINLOOP_CONFIG_DIR", home) - dir := filepath.Join(home, "remotes", name) + dir := filepath.Join(home, "clouds", name) if err := os.MkdirAll(dir, 0o755); err != nil { t.Fatal(err) } body := fmt.Sprintf( `{"start_url":%q,"stop_url":%q,"stats_url":%q,"region":"us-east-1","environment":%q}`, url, url, url, name) - if err := os.WriteFile(filepath.Join(dir, "remote.json"), []byte(body), 0o600); err != nil { + if err := os.WriteFile(filepath.Join(dir, "cloud.json"), []byte(body), 0o600); err != nil { t.Fatal(err) } } @@ -64,10 +64,10 @@ func TestDaemonNodeImplementsNoCloudCapability(t *testing.T) { // A cloud node implements all three, so a caller reaches them by assertion // rather than by asking what kind it is. -func TestRemoteNodeImplementsTheCloudCapabilities(t *testing.T) { +func TestCloudNodeImplementsTheCloudCapabilities(t *testing.T) { stubAWSCreds(t) - up := remoteControlServer(t, `{"state":"running"}`, http.StatusOK) - registerRemoteEnv(t, "prod", up.URL, up.URL) + up := cloudControlServer(t, `{"state":"running"}`, http.StatusOK) + registerCloudEnv(t, "prod", up.URL, up.URL) cfg, err := ForEnvironment("prod") if err != nil { t.Fatal(err) @@ -88,7 +88,7 @@ func TestRemoteNodeImplementsTheCloudCapabilities(t *testing.T) { } // The instance facts come from the reading already taken, not a second call. -func TestRemoteNodeInstanceComesFromTheMetricsReading(t *testing.T) { +func TestCloudNodeInstanceComesFromTheMetricsReading(t *testing.T) { stubAWSCreds(t) calls := 0 srv := countingStatsServer(t, `{"state":"running","version":"1.40.0","instanceId":"i-0abc","instanceType":"g6e.xlarge","uptimeSeconds":7200}`, &calls) @@ -121,10 +121,10 @@ func TestRemoteNodeInstanceComesFromTheMetricsReading(t *testing.T) { // A node with nothing to price reports no cost rather than a zero one: "$0.00" // claims it cost nothing, which is a different statement. -func TestRemoteNodeCostIsUnreportedWithoutAReading(t *testing.T) { +func TestCloudNodeCostIsUnreportedWithoutAReading(t *testing.T) { stubAWSCreds(t) - up := remoteControlServer(t, `{"state":"stopped"}`, http.StatusOK) - registerRemoteEnv(t, "prod", up.URL, up.URL) + up := cloudControlServer(t, `{"state":"stopped"}`, http.StatusOK) + registerCloudEnv(t, "prod", up.URL, up.URL) cfg, err := ForEnvironment("prod") if err != nil { t.Fatal(err) diff --git a/internal/fleet/capability_calls_test.go b/internal/fleet/capability_calls_test.go index 4a72b871..9c6bb58c 100644 --- a/internal/fleet/capability_calls_test.go +++ b/internal/fleet/capability_calls_test.go @@ -5,10 +5,10 @@ import ( "testing" "time" + "github.com/spinloop-ai/spinloop/internal/cloud" "github.com/spinloop-ai/spinloop/internal/daemon" "github.com/spinloop-ai/spinloop/internal/inference" "github.com/spinloop-ai/spinloop/internal/metrics" - "github.com/spinloop-ai/spinloop/internal/remote" ) // stubNode is a node that answers whatever the test gives it and implements @@ -151,11 +151,11 @@ func TestQueriedLogsCall(t *testing.T) { } } -// A remote environment's queried read reaches its log store with the window +// A cloud environment's queried read reaches its log store with the window // the caller named. The store is CloudWatch, reached through the AWS SDK // rather than the control plane's HTTP endpoints, so this substitutes the // FetchLogsFn variable rather than an httptest server. -func TestRemoteNodeLogsMatching(t *testing.T) { +func TestCloudNodeLogsMatching(t *testing.T) { registerStatsEnv(t, "prod", "http://unused.invalid") cfg, err := ForEnvironment("prod") if err != nil { @@ -172,10 +172,10 @@ func TestRemoteNodeLogsMatching(t *testing.T) { restore := FetchLogsFn t.Cleanup(func() { FetchLogsFn = restore }) - var got remote.LogQuery - FetchLogsFn = func(_ context.Context, _ remote.Config, q remote.LogQuery) (remote.LogResult, error) { + var got cloud.LogQuery + FetchLogsFn = func(_ context.Context, _ cloud.Config, q cloud.LogQuery) (cloud.LogResult, error) { got = q - return remote.LogResult{Events: []remote.LogEvent{{Message: "boot line"}}}, nil + return cloud.LogResult{Events: []cloud.LogEvent{{Message: "boot line"}}}, nil } resp, err := s.LogsMatching(context.Background(), LogQuery{Source: "boot", Limit: 10, Instance: "i-1"}) diff --git a/internal/fleet/cloud_kind_test.go b/internal/fleet/cloud_kind_test.go new file mode 100644 index 00000000..eb9eac0e --- /dev/null +++ b/internal/fleet/cloud_kind_test.go @@ -0,0 +1,14 @@ +package fleet + +import ( + "strings" + "testing" +) + +// The old kind spelling is gone: `kind: remote` fails like any other unknown kind. +func TestLoadRemoteKindIsRejected(t *testing.T) { + _, err := Load(writeFleet(t, "nodes:\n - name: prod\n kind: remote\n", "")) + if err == nil || !strings.Contains(err.Error(), "remote") { + t.Fatalf("error = %v, want one naming the unsupported kind", err) + } +} diff --git a/internal/fleet/remote_node.go b/internal/fleet/cloud_node.go similarity index 77% rename from internal/fleet/remote_node.go rename to internal/fleet/cloud_node.go index 772bfd07..a2f65227 100644 --- a/internal/fleet/remote_node.go +++ b/internal/fleet/cloud_node.go @@ -9,31 +9,31 @@ import ( "sync" "time" + "github.com/spinloop-ai/spinloop/internal/cloud" "github.com/spinloop-ai/spinloop/internal/daemon" "github.com/spinloop-ai/spinloop/internal/inference" "github.com/spinloop-ai/spinloop/internal/metrics" - "github.com/spinloop-ai/spinloop/internal/remote" ) -// remoteNode is one member of the fleet as seen through a remote, scale-to-zero -// environment: a `remote.Config` reached over its cloud control plane rather than a +// cloudNode is one member of the fleet as seen through a remote, scale-to-zero +// environment: a `cloud.Config` reached over its cloud control plane rather than a // machine's daemon. It answers the same operations a local node answers, so the same // fan-out and rendering pass over both kinds — which is the whole point of routing a -// remote environment through the node contract. +// cloud environment through the node contract. // // The control-plane state (EC2) and a daemon's engine state are different // vocabularies. We do not translate one into the other: the status carries the -// control-plane state as the control plane reports it. A remote environment's +// control-plane state as the control plane reports it. A cloud environment's // "running" is the instance state, not a claim about the engine. -type remoteNode struct { +type cloudNode struct { name string - cfg remote.Config + cfg cloud.Config // logs holds the position a follow of this node's engine log has reached. - // It is the same cursor `spinloop remote logs -f` uses, and for the same + // It is the same cursor `spinloop cloud logs -f` uses, and for the same // reason: CloudWatch has no resumable read position of its own, so a poll // re-asks a little behind the newest event already seen and this // suppresses what the overlap re-reads, by event id. - logs *remote.FollowCursor + logs *cloud.FollowCursor // mu guards the facts Metrics retains for the capabilities to answer from. // A board refreshing several nodes reads and writes these from different @@ -47,23 +47,23 @@ type remoteNode struct { uptime int } -// NewRemoteNode builds the live node for a named remote environment. The config +// NewCloudNode builds the live node for a named cloud environment. The config // must be complete enough to send a signed control call — start, stop and region — // so a config that cannot be resolved is a configuration error against this node // rather than a call that fails part-way through. -func NewRemoteNode(name string, cfg remote.Config) (Node, error) { +func NewCloudNode(name string, cfg cloud.Config) (Node, error) { if cfg.StartURL == "" || cfg.StopURL == "" || cfg.Region == "" { return nil, fmt.Errorf( - "remote environment %q is not fully configured: start_url, stop_url and region are all required", + "cloud environment %q is not fully configured: start_url, stop_url and region are all required", name) } - return &remoteNode{name: name, cfg: cfg, logs: remote.NewFollowCursor(remote.FollowOverlap)}, nil + return &cloudNode{name: name, cfg: cfg, logs: cloud.NewFollowCursor(cloud.FollowOverlap)}, nil } -func (n *remoteNode) Name() string { return n.name } +func (n *cloudNode) Name() string { return n.name } -func (n *remoteNode) Status(ctx context.Context) (daemon.StatusResponse, error) { - resp, err := remote.Status(ctx, n.cfg) +func (n *cloudNode) Status(ctx context.Context) (daemon.StatusResponse, error) { + resp, err := cloud.Status(ctx, n.cfg) if err != nil { return daemon.StatusResponse{}, err } @@ -77,7 +77,7 @@ func (n *remoteNode) Status(ctx context.Context) (daemon.StatusResponse, error) // answer, and a failure to reach it leaves the version empty rather than // failing a status that otherwise succeeded. if resp.State == "running" || resp.State == "ready" { - if stats, err := remote.Stats(ctx, n.cfg); err == nil { + if stats, err := cloud.Stats(ctx, n.cfg); err == nil { n.mu.Lock() n.instance.Version = stats.Version n.instance.ID = stats.InstanceID @@ -85,11 +85,11 @@ func (n *remoteNode) Status(ctx context.Context) (daemon.StatusResponse, error) n.mu.Unlock() } } - return statusFromRemote(*resp), nil + return statusFromCloud(*resp), nil } -func (n *remoteNode) Metrics(ctx context.Context) (metrics.Stats, error) { - resp, err := remote.Stats(ctx, n.cfg) +func (n *cloudNode) Metrics(ctx context.Context) (metrics.Stats, error) { + resp, err := cloud.Stats(ctx, n.cfg) if err != nil { return metrics.Stats{}, err } @@ -102,7 +102,7 @@ func (n *remoteNode) Metrics(ctx context.Context) (metrics.Stats, error) { n.instance.Type = resp.InstanceType n.instance.Version = resp.Version n.mu.Unlock() - return statsFromRemote(*resp), nil + return statsFromCloud(*resp), nil } // Cost prices the session this environment has been running, from the instance @@ -113,14 +113,14 @@ func (n *remoteNode) Metrics(ctx context.Context) (metrics.Stats, error) { // A reading not yet taken, an instance that is not running, or a price the // Price List API does not return all yield the zero Cost and no error: a node // with no price to show reads the same however it came to have none. -func (n *remoteNode) Cost(ctx context.Context) (Cost, error) { +func (n *cloudNode) Cost(ctx context.Context) (Cost, error) { n.mu.Lock() instanceType, uptime := n.instance.Type, n.uptime n.mu.Unlock() if instanceType == "" || uptime <= 0 { return Cost{}, nil } - price, err := remote.GetOnDemandPrice(ctx, n.cfg.Region, instanceType) + price, err := cloud.GetOnDemandPrice(ctx, n.cfg.Region, instanceType) if err != nil || price <= 0 { return Cost{}, nil } @@ -131,7 +131,7 @@ func (n *remoteNode) Cost(ctx context.Context) (Cost, error) { // environment runs on. Empty fields are ones no reply carried — a stopped // environment has no instance id, and a control plane that could not reach the // daemon reports no version. -func (n *remoteNode) Instance() Instance { +func (n *cloudNode) Instance() Instance { n.mu.Lock() defer n.mu.Unlock() return n.instance @@ -140,16 +140,16 @@ func (n *remoteNode) Instance() Instance { // LogsMatching reads this environment's log store, narrowed by what the caller // asked for. Unlike Logs it is a query rather than a cursor: the caller states // the window it wants, so nothing is retained between calls. -func (n *remoteNode) LogsMatching(ctx context.Context, q LogQuery) (daemon.LogsResponse, error) { +func (n *cloudNode) LogsMatching(ctx context.Context, q LogQuery) (daemon.LogsResponse, error) { source := q.Source if source == "" { - source = remote.LogSourceEngine + source = cloud.LogSourceEngine } limit := q.Limit if limit <= 0 { - limit = remoteEngineTail + limit = cloudEngineTail } - rq := remote.LogQuery{ + rq := cloud.LogQuery{ Environment: n.cfg.Environment, Source: source, Limit: limit, @@ -163,10 +163,10 @@ func (n *remoteNode) LogsMatching(ctx context.Context, q LogQuery) (daemon.LogsR return daemon.LogsResponse{}, err } // A query is not a follow, so every event it returns is fresh to it. - return logsFromRemote(res.Events, len(res.Events) == 0), nil + return logsFromCloud(res.Events, len(res.Events) == 0), nil } -func (n *remoteNode) Start(ctx context.Context) (daemon.StatusResponse, error) { +func (n *cloudNode) Start(ctx context.Context) (daemon.StatusResponse, error) { return n.StartWithProgress(ctx, func(StartPhase) {}) } @@ -175,24 +175,24 @@ func (n *remoteNode) Start(ctx context.Context) (daemon.StatusResponse, error) { // and when the next attempt is due, the instance coming up, a dropped // connection. Each phase replaces the one before it. Callers may render or // discard them. -func (n *remoteNode) StartWithProgress(ctx context.Context, report func(StartPhase)) (daemon.StatusResponse, error) { +func (n *cloudNode) StartWithProgress(ctx context.Context, report func(StartPhase)) (daemon.StatusResponse, error) { if report == nil { - // remote.Start invokes its callbacks on every retry path, so nil is + // cloud.Start invokes its callbacks on every retry path, so nil is // substituted with a no-op here rather than left to whichever paths a // given start takes. report = func(StartPhase) {} } progress, onState := StartPhases(report) - resp, err := remote.Start(ctx, n.cfg, progress, onState, nil) + resp, err := cloud.Start(ctx, n.cfg, progress, onState, nil) if err != nil { return daemon.StatusResponse{}, err } - return statusFromRemote(*resp), nil + return statusFromCloud(*resp), nil } // StartWith is how a router wakes a node to serve something. A remote // environment's engine is not configured by a start: what it serves and the -// key that gates it are fixed by `spinloop remote deploy`, a heavier flow +// key that gates it are fixed by `spinloop cloud deploy`, a heavier flow // (provisioning, weight seeding, ingress) that a node start must not // conflate — so dc and engineKey are ignored, and this boots the instance // exactly as Start does. @@ -206,18 +206,18 @@ func (n *remoteNode) StartWithProgress(ctx context.Context, report func(StartPha // candidate ever reaches this call. A caller that skips that matching and // hands an undeployed environment straight to StartWith gets the boot // call's own answer instead, whatever that turns out to be. -func (n *remoteNode) StartWith(ctx context.Context, dc *inference.DeployConfig, engineKey string) (daemon.StatusResponse, error) { +func (n *cloudNode) StartWith(ctx context.Context, dc *inference.DeployConfig, engineKey string) (daemon.StatusResponse, error) { _ = dc _ = engineKey return n.StartWithProgress(ctx, func(StartPhase) {}) } -func (n *remoteNode) Stop(ctx context.Context) (daemon.StatusResponse, error) { - resp, err := remote.Stop(ctx, n.cfg) +func (n *cloudNode) Stop(ctx context.Context) (daemon.StatusResponse, error) { + resp, err := cloud.Stop(ctx, n.cfg) if err != nil { return daemon.StatusResponse{}, err } - return statusFromRemote(*resp), nil + return statusFromCloud(*resp), nil } // Keep pins this environment's instance so the idle sweep does not terminate it @@ -228,9 +228,9 @@ func (n *remoteNode) Stop(ctx context.Context) (daemon.StatusResponse, error) { // verbatim; when the reply omits it — a control plane that predates keep // echoing the value back — the requested deadline is returned instead, so a // caller always holds something to show. -func (n *remoteNode) Keep(ctx context.Context, d time.Duration) (string, error) { +func (n *cloudNode) Keep(ctx context.Context, d time.Duration) (string, error) { deadline := time.Now().Add(d) - resp, err := remote.Keep(ctx, n.cfg, deadline) + resp, err := cloud.Keep(ctx, n.cfg, deadline) if err != nil { return "", err } @@ -244,14 +244,14 @@ func (n *remoteNode) Keep(ctx context.Context, d time.Duration) (string, error) // substitute it. The read goes to the cloud's log store through the AWS SDK // rather than an HTTP Function URL, so there is no HTTP test server that can // stand in for it. -var FetchLogsFn = remote.FetchLogs +var FetchLogsFn = cloud.FetchLogs -// remoteEngineTail caps how many engine log events a node read pulls. Remote logs +// cloudEngineTail caps how many engine log events a node read pulls. Cloud logs // are a bounded tail pulled from the log store, not a byte cursor, so a follow of // a chatty engine must not page through an unbounded window. -const remoteEngineTail = 1000 +const cloudEngineTail = 1000 -func (n *remoteNode) Logs(ctx context.Context, offset int64, limit int) (daemon.LogsResponse, error) { +func (n *cloudNode) Logs(ctx context.Context, offset int64, limit int) (daemon.LogsResponse, error) { // daemon.TailLog means a fresh open of the view: start the cursor over, // so this open shows its own tail rather than having it suppressed as // already seen by whatever this node last followed. @@ -259,10 +259,10 @@ func (n *remoteNode) Logs(ctx context.Context, offset int64, limit int) (daemon. n.logs.Reset() } start := n.logs.Start() - res, err := FetchLogsFn(ctx, n.cfg, remote.LogQuery{ + res, err := FetchLogsFn(ctx, n.cfg, cloud.LogQuery{ Environment: n.cfg.Environment, - Source: remote.LogSourceEngine, - Limit: remoteEngineTail, + Source: cloud.LogSourceEngine, + Limit: cloudEngineTail, Start: start, }) if err != nil { @@ -273,10 +273,10 @@ func (n *remoteNode) Logs(ctx context.Context, offset int64, limit int) (daemon. // that is genuinely no log, ever. A later poll with nothing new is a // quiet log, not a missing one. missing := start.IsZero() && len(res.Events) == 0 - return logsFromRemote(fresh, missing), nil + return logsFromCloud(fresh, missing), nil } -// statusFromRemote maps the control plane's status reply onto the node's status. +// statusFromCloud maps the control plane's status reply onto the node's status. // It carries what a status reply can honestly be mapped across: the state, what // the engine is serving (runner, model, served name) and its last-active record. // The version is not in this reply — the stats reply carries it — so it is empty @@ -286,20 +286,20 @@ func (n *remoteNode) Logs(ctx context.Context, offset int64, limit int) (daemon. // A running environment's reply also names where its engine answers — the // instance's published address, which a daemon on the instance cannot know for // itself but the control plane can. It is carried as the engine's host, so -// routing resolves a remote node's address the way it resolves any node's. +// routing resolves a cloud node's address the way it resolves any node's. // Absent (a stopped or undeployed environment reports none) means no engine // address, exactly as the parts would be. // // Healthy is the control plane's own readiness reading — the same health // check (hitting the engine's /health, excluding the 503 it answers while -// still loading weights) a running remote view already carries — mapped +// still loading weights) a running cloud view already carries — mapped // onto Ready the way a local daemon's own reading is, so a router waiting -// for a remote engine to answer trusts this instead of falling back to +// for a cloud engine to answer trusts this instead of falling back to // whether its port merely accepts a connection, which it can do well before // the model has loaded. Absent (an older control plane, or the SSM agent // not yet reachable) leaves Ready empty, the same "no reading yet" a local // daemon reports before its own first check lands. -func statusFromRemote(resp remote.Response) daemon.StatusResponse { +func statusFromCloud(resp cloud.Response) daemon.StatusResponse { s := daemon.StatusResponse{ State: resp.State, Runner: resp.Runner, @@ -326,14 +326,14 @@ func statusFromRemote(resp remote.Response) daemon.StatusResponse { return s } -// statsFromRemote maps the stats Lambda's reply onto the shared stats shape. The +// statsFromCloud maps the stats Lambda's reply onto the shared stats shape. The // per-stat fields already alias the metrics types the collector produces, so the // mapping is a field-for-field copy. The reply's version and instance facts are // deliberately not among them: they describe a cloud environment and nothing // else, so putting them on the shape every node answers with would leave a // field every daemon reports empty. Metrics retains them on the node instead, // where Cost and Version answer from them. -func statsFromRemote(resp remote.StatsResponse) metrics.Stats { +func statsFromCloud(resp cloud.StatsResponse) metrics.Stats { return metrics.Stats{ State: resp.State, Runner: resp.Runner, @@ -352,7 +352,7 @@ func statsFromRemote(resp remote.StatsResponse) metrics.Stats { } } -// logsFromRemote maps one poll's fresh events — the ones remoteNode.Logs's +// logsFromCloud maps one poll's fresh events — the ones cloudNode.Logs's // FollowCursor has not already returned — onto the node's log reply. Events // arrive oldest first, so the content reads top to bottom. missing reports // that this was a from-the-beginning read that found nothing: the state a @@ -362,7 +362,7 @@ func statsFromRemote(resp remote.StatsResponse) metrics.Stats { // lives in the node's own cursor — so it carries the newest event shown for // whoever finds that useful to see, and 0 has the same effect when there is // none. -func logsFromRemote(fresh []remote.LogEvent, missing bool) daemon.LogsResponse { +func logsFromCloud(fresh []cloud.LogEvent, missing bool) daemon.LogsResponse { if len(fresh) == 0 { return daemon.LogsResponse{Missing: missing} } diff --git a/internal/fleet/remote_node_test.go b/internal/fleet/cloud_node_test.go similarity index 77% rename from internal/fleet/remote_node_test.go rename to internal/fleet/cloud_node_test.go index 04cb2b41..50530fa6 100644 --- a/internal/fleet/remote_node_test.go +++ b/internal/fleet/cloud_node_test.go @@ -14,10 +14,10 @@ import ( "testing" "time" + "github.com/spinloop-ai/spinloop/internal/cloud" "github.com/spinloop-ai/spinloop/internal/daemon" "github.com/spinloop-ai/spinloop/internal/inference" "github.com/spinloop-ai/spinloop/internal/metrics" - "github.com/spinloop-ai/spinloop/internal/remote" ) // stubAWSCreds pins the AWS credential chain to static environment credentials @@ -36,8 +36,8 @@ func stubAWSCreds(t *testing.T) { func boolPtr(b bool) *bool { return &b } -// remoteControlServer serves the shape a remote control endpoint answers. -func remoteControlServer(t *testing.T, body string, statusCode int) *httptest.Server { +// cloudControlServer serves the shape a remote control endpoint answers. +func cloudControlServer(t *testing.T, body string, statusCode int) *httptest.Server { t.Helper() srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) { w.Header().Set("Content-Type", "application/json") @@ -50,21 +50,21 @@ func remoteControlServer(t *testing.T, body string, statusCode int) *httptest.Se return srv } -// Fleet operations on remote environments go through the same signed control -// calls as `spinloop remote`, so they inherit the stored control-plane +// Fleet operations on cloud environments go through the same signed control +// calls as `spinloop cloud`, so they inherit the stored control-plane // credential: here the only credential the process can resolve is the one // stored for the region (the ambient chain is empty), and the status call -// still signs and gets its answer. The file store behind SPINLOOP_REMOTE_KEYSTORE +// still signs and gets its answer. The file store behind SPINLOOP_CLOUD_KEYSTORE // keeps the entry in a temp directory, never in the machine's keystore. -func TestRemoteNodeSignsWithStoredCredential(t *testing.T) { +func TestCloudNodeSignsWithStoredCredential(t *testing.T) { region := "ap-southeast-2" - srv := remoteControlServer(t, `{"state":"stopped"}`, http.StatusOK) + srv := cloudControlServer(t, `{"state":"stopped"}`, http.StatusOK) home := t.TempDir() t.Setenv("HOME", home) t.Setenv("XDG_CONFIG_HOME", filepath.Join(home, ".config")) t.Setenv("SPINLOOP_CONFIG_DIR", "") - t.Setenv("SPINLOOP_REMOTE_KEYSTORE", "file") + t.Setenv("SPINLOOP_CLOUD_KEYSTORE", "file") // No ambient credential at all: env, profile, config files, IMDS. t.Setenv("AWS_ACCESS_KEY_ID", "") t.Setenv("AWS_SECRET_ACCESS_KEY", "") @@ -74,7 +74,7 @@ func TestRemoteNodeSignsWithStoredCredential(t *testing.T) { t.Setenv("AWS_SHARED_CREDENTIALS_FILE", filepath.Join(home, "no-such-file")) t.Setenv("AWS_EC2_METADATA_DISABLED", "true") - cred := remote.StoredCredential{ + cred := cloud.StoredCredential{ AccessKeyID: "AKIATESTTESTTESTTEST", SecretAccessKey: "test-secret", Account: "0", @@ -82,12 +82,12 @@ func TestRemoteNodeSignsWithStoredCredential(t *testing.T) { Region: region, StoredAt: time.Now().UTC(), } - if err := remote.StoreCredential(cred); err != nil { + if err := cloud.StoreCredential(cred); err != nil { t.Fatalf("StoreCredential: %v", err) } - t.Cleanup(func() { remote.DeleteStoredCredential(region) }) + t.Cleanup(func() { cloud.DeleteStoredCredential(region) }) - node, err := NewRemoteNode("env", remote.Config{StartURL: srv.URL, StopURL: srv.URL, Region: region}) + node, err := NewCloudNode("env", cloud.Config{StartURL: srv.URL, StopURL: srv.URL, Region: region}) if err != nil { t.Fatal(err) } @@ -100,45 +100,45 @@ func TestRemoteNodeSignsWithStoredCredential(t *testing.T) { } } -func TestNewRemoteNodeRequiresACompleteConfig(t *testing.T) { - if _, err := NewRemoteNode("env", remote.Config{StopURL: "http://x", Region: "r"}); err == nil { +func TestNewCloudNodeRequiresACompleteConfig(t *testing.T) { + if _, err := NewCloudNode("env", cloud.Config{StopURL: "http://x", Region: "r"}); err == nil { t.Error("missing start_url should be a configuration error") } - if _, err := NewRemoteNode("env", remote.Config{StartURL: "http://x", StopURL: "http://x"}); err == nil { + if _, err := NewCloudNode("env", cloud.Config{StartURL: "http://x", StopURL: "http://x"}); err == nil { t.Error("missing region should be a configuration error") } - if _, err := NewRemoteNode("env", remote.Config{StartURL: "http://x", StopURL: "http://x", Region: "r"}); err != nil { + if _, err := NewCloudNode("env", cloud.Config{StartURL: "http://x", StopURL: "http://x", Region: "r"}); err != nil { t.Errorf("a complete config should build a node: %v", err) } } -func TestStatusFromRemote(t *testing.T) { - got := statusFromRemote(remote.Response{ +func TestStatusFromCloud(t *testing.T) { + got := statusFromCloud(cloud.Response{ State: "running", Healthy: boolPtr(true), BaseURL: "http://1.2.3.4:8000/v1", LastActiveAt: "2026-01-02T00:00:00Z", IdleSeconds: 30, }) if got.State != "running" || got.IdleSeconds != 30 || got.LastActiveAt != "2026-01-02T00:00:00Z" { - t.Errorf("statusFromRemote = %+v", got) + t.Errorf("statusFromCloud = %+v", got) } // A stopped endpoint reports no activity: nothing to measure from. - if got := statusFromRemote(remote.Response{State: "stopped"}); got.State != "stopped" || got.LastActiveAt != "" { + if got := statusFromCloud(cloud.Response{State: "stopped"}); got.State != "stopped" || got.LastActiveAt != "" { t.Errorf("stopped status = %+v", got) } } -func TestStatsFromRemote(t *testing.T) { +func TestStatsFromCloud(t *testing.T) { requests := 3 tokens := &metrics.TokenStats{Running: 2, PromptTokens: 5, GenerationTokens: 7, Requests: &requests} cpuPct := 30.0 history := []metrics.HistorySample{{Time: 1, CPU: &cpuPct, GPUs: []metrics.HistoryGPU{{Index: 0, Util: 61, Mem: &cpuPct}}}} - got := statsFromRemote(remote.StatsResponse{ + got := statsFromCloud(cloud.StatsResponse{ State: "running", Runner: "llamacpp", ModelID: "org/m", UptimeSeconds: 10, Tokens: tokens, LastActiveAt: "2026-01-02T00:00:00Z", IdleSeconds: 5, Version: "1.2.3", History: history, RetainUntil: "2026-01-02T04:00:00Z", }) if got.State != "running" || got.Runner != "llamacpp" || got.ModelID != "org/m" || got.UptimeSeconds != 10 { - t.Errorf("statsFromRemote = %+v", got) + t.Errorf("statsFromCloud = %+v", got) } if got.Tokens == nil || got.Tokens.Running != 2 || got.Tokens.Requests == nil || *got.Tokens.Requests != 3 { t.Errorf("token stats not carried over: %+v", got.Tokens) @@ -153,7 +153,7 @@ func TestStatsFromRemote(t *testing.T) { t.Errorf("history not carried over: %+v", got.History) } // And a reply without them stays without them. - if got := statsFromRemote(remote.StatsResponse{State: "running"}); got.History != nil { + if got := statsFromCloud(cloud.StatsResponse{State: "running"}); got.History != nil { t.Errorf("an absent history became present: %+v", got.History) } if got.RetainUntil != "2026-01-02T04:00:00Z" { @@ -161,25 +161,25 @@ func TestStatsFromRemote(t *testing.T) { } // An environment with no live retention carries nothing: the field is // absent, not a zero time a formatter would have to special-case. - if got := statsFromRemote(remote.StatsResponse{State: "running"}); got.RetainUntil != "" { + if got := statsFromCloud(cloud.StatsResponse{State: "running"}); got.RetainUntil != "" { t.Errorf("no retention should map to an empty retainUntil, got %q", got.RetainUntil) } } -func TestLogsFromRemote(t *testing.T) { - if got := logsFromRemote(nil, true); !got.Missing { +func TestLogsFromCloud(t *testing.T) { + if got := logsFromCloud(nil, true); !got.Missing { t.Errorf("no fresh events on a from-the-beginning read should be reported missing, got %+v", got) } // No fresh events on a later poll — the common steady state once a follow // is caught up — must not be reported missing: the log is not gone, it is // just quiet right now. - if got := logsFromRemote(nil, false); got.Missing { + if got := logsFromCloud(nil, false); got.Missing { t.Errorf("no fresh events on a later poll should not be reported missing, got %+v", got) } t1 := time.UnixMilli(1000) t2 := time.UnixMilli(2000) - got := logsFromRemote([]remote.LogEvent{ + got := logsFromCloud([]cloud.LogEvent{ {Message: "loading model", Timestamp: t1}, {Message: "server ready", Timestamp: t2}, }, false) want := "loading model\nserver ready\n" @@ -194,13 +194,13 @@ func TestLogsFromRemote(t *testing.T) { } } -// A remote node's log follow uses the exact same cursor as `spinloop remote -// logs -f` (remote.FollowCursor, seeded with remote.FollowOverlap) — not a +// A cloud node's log follow uses the exact same cursor as `spinloop cloud +// logs -f` (cloud.FollowCursor, seeded with cloud.FollowOverlap) — not a // lookalike reimplementation — so a second poll bounds its query behind the // newest event already shown by the same overlap window, and does not show // that event again. The bug this guards is the tail being replayed in full // on every poll regardless of what was already shown. -func TestRemoteNodeLogsSharesTheFollowCursorWithRemoteLogsCommand(t *testing.T) { +func TestCloudNodeLogsSharesTheFollowCursorWithCloudLogsCommand(t *testing.T) { stubAWSCreds(t) eventMs := time.Date(2026, 8, 9, 11, 30, 0, 0, time.UTC).UnixMilli() @@ -232,8 +232,8 @@ func TestRemoteNodeLogsSharesTheFollowCursorWithRemoteLogsCommand(t *testing.T) })) t.Cleanup(srv.Close) t.Setenv("AWS_ENDPOINT_URL_CLOUDWATCH_LOGS", srv.URL) - cfg := remote.Config{StartURL: "http://x", StopURL: "http://x", Environment: "env-1", Region: "us-east-1"} - node, err := NewRemoteNode("env", cfg) + cfg := cloud.Config{StartURL: "http://x", StopURL: "http://x", Environment: "env-1", Region: "us-east-1"} + node, err := NewCloudNode("env", cfg) if err != nil { t.Fatal(err) } @@ -259,7 +259,7 @@ func TestRemoteNodeLogsSharesTheFollowCursorWithRemoteLogsCommand(t *testing.T) if resp2.Content != "" || resp2.Missing { t.Errorf("the follow-up poll should show nothing new and not report missing, got %+v", resp2) } - wantStart := eventMs - remote.FollowOverlap.Milliseconds() + wantStart := eventMs - cloud.FollowOverlap.Milliseconds() for _, s := range starts { if s == nil { t.Fatal("the follow-up poll sent no start bound") @@ -271,10 +271,10 @@ func TestRemoteNodeLogsSharesTheFollowCursorWithRemoteLogsCommand(t *testing.T) } // StartWith boots a deployed-but-stopped environment exactly as Start does, -// ignoring the config and key it is handed: what a remote environment -// serves, and the key that gates it, are fixed by `spinloop remote deploy`, +// ignoring the config and key it is handed: what a cloud environment +// serves, and the key that gates it, are fixed by `spinloop cloud deploy`, // not by a wake call. -func TestRemoteNodeStartWithBootsADeployedEnvironment(t *testing.T) { +func TestCloudNodeStartWithBootsADeployedEnvironment(t *testing.T) { stubAWSCreds(t) mux := http.NewServeMux() mux.HandleFunc("GET /", func(w http.ResponseWriter, r *http.Request) { @@ -287,7 +287,7 @@ func TestRemoteNodeStartWithBootsADeployedEnvironment(t *testing.T) { }) srv := httptest.NewServer(mux) t.Cleanup(srv.Close) - node, err := NewRemoteNode("env", remote.Config{StartURL: srv.URL, StopURL: srv.URL, Region: "us-east-1"}) + node, err := NewCloudNode("env", cloud.Config{StartURL: srv.URL, StopURL: srv.URL, Region: "us-east-1"}) if err != nil { t.Fatal(err) } @@ -308,7 +308,7 @@ func TestRemoteNodeStartWithBootsADeployedEnvironment(t *testing.T) { // the gateway's own candidate matching confirms deployment from a stats read // before a candidate is ever chosen — so StartWith just boots, whatever a // status read would have said. -func TestRemoteNodeStartWithDoesNotItselfCheckDeployment(t *testing.T) { +func TestCloudNodeStartWithDoesNotItselfCheckDeployment(t *testing.T) { stubAWSCreds(t) mux := http.NewServeMux() mux.HandleFunc("GET /", func(w http.ResponseWriter, r *http.Request) { @@ -321,7 +321,7 @@ func TestRemoteNodeStartWithDoesNotItselfCheckDeployment(t *testing.T) { }) srv := httptest.NewServer(mux) t.Cleanup(srv.Close) - node, err := NewRemoteNode("env", remote.Config{StartURL: srv.URL, StopURL: srv.URL, Region: "us-east-1"}) + node, err := NewCloudNode("env", cloud.Config{StartURL: srv.URL, StopURL: srv.URL, Region: "us-east-1"}) if err != nil { t.Fatal(err) } @@ -330,12 +330,12 @@ func TestRemoteNodeStartWithDoesNotItselfCheckDeployment(t *testing.T) { } } -func TestRemoteNodeStatusOverTheControlPlane(t *testing.T) { +func TestCloudNodeStatusOverTheControlPlane(t *testing.T) { stubAWSCreds(t) - srv := remoteControlServer(t, + srv := cloudControlServer(t, `{"state":"running","healthy":true,"runner":"llamacpp","modelId":"org/m","servedName":"m",`+ `"base_url":"http://1.2.3.4:8000/v1","lastActiveAt":"2026-01-02T00:00:00Z","idleSeconds":30}`, http.StatusOK) - node, err := NewRemoteNode("env", remote.Config{StartURL: srv.URL, StopURL: srv.URL, Region: "us-east-1"}) + node, err := NewCloudNode("env", cloud.Config{StartURL: srv.URL, StopURL: srv.URL, Region: "us-east-1"}) if err != nil { t.Fatal(err) } @@ -358,16 +358,16 @@ func TestRemoteNodeStatusOverTheControlPlane(t *testing.T) { } } -// A remote node's keep pins the instance over its control plane: the deadline +// A cloud node's keep pins the instance over its control plane: the deadline // is computed from the duration and sent as an absolute time, and the value // returned to a caller is the control plane's own deadline, not the caller's // clock plus the duration. -func TestRemoteNodeKeepOverTheControlPlane(t *testing.T) { +func TestCloudNodeKeepOverTheControlPlane(t *testing.T) { stubAWSCreds(t) - srv := remoteControlServer(t, + srv := cloudControlServer(t, `{"retainUntil":"2026-01-02T04:00:00Z"}`, http.StatusOK) - cfg := remote.Config{StartURL: srv.URL, StopURL: srv.URL, UpdateURL: srv.URL, Region: "us-east-1"} - node, err := NewRemoteNode("env", cfg) + cfg := cloud.Config{StartURL: srv.URL, StopURL: srv.URL, UpdateURL: srv.URL, Region: "us-east-1"} + node, err := NewCloudNode("env", cfg) if err != nil { t.Fatal(err) } @@ -383,11 +383,11 @@ func TestRemoteNodeKeepOverTheControlPlane(t *testing.T) { // When the control plane's reply omits the deadline — one that predates keep // echoing the value back — the keep still succeeds and returns the requested // deadline, so a caller always has something to show. -func TestRemoteNodeKeepFallsBackToTheRequestedDeadline(t *testing.T) { +func TestCloudNodeKeepFallsBackToTheRequestedDeadline(t *testing.T) { stubAWSCreds(t) - srv := remoteControlServer(t, `{}`, http.StatusOK) - cfg := remote.Config{StartURL: srv.URL, StopURL: srv.URL, UpdateURL: srv.URL, Region: "us-east-1"} - node, _ := NewRemoteNode("env", cfg) + srv := cloudControlServer(t, `{}`, http.StatusOK) + cfg := cloud.Config{StartURL: srv.URL, StopURL: srv.URL, UpdateURL: srv.URL, Region: "us-east-1"} + node, _ := NewCloudNode("env", cfg) before := time.Now() got, err := node.(Keeper).Keep(context.Background(), 4*time.Hour) if err != nil { @@ -405,23 +405,23 @@ func TestRemoteNodeKeepFallsBackToTheRequestedDeadline(t *testing.T) { // A keep on an environment whose config has no update URL is a configuration // error before any call, naming the fix — it is not a call that fails part-way. -func TestRemoteNodeKeepWithoutAnUpdateURL(t *testing.T) { +func TestCloudNodeKeepWithoutAnUpdateURL(t *testing.T) { stubAWSCreds(t) - cfg := remote.Config{StartURL: "http://x", StopURL: "http://x", Region: "us-east-1"} - node, _ := NewRemoteNode("env", cfg) + cfg := cloud.Config{StartURL: "http://x", StopURL: "http://x", Region: "us-east-1"} + node, _ := NewCloudNode("env", cfg) if _, err := node.(Keeper).Keep(context.Background(), time.Hour); err == nil || !strings.Contains(err.Error(), "no update_url") { t.Errorf("expected a no-update-url error, got %v", err) } } -// The keep capability is exactly one of the remote node's: a local daemon node +// The keep capability is exactly one of the cloud node's: a local daemon node // has no retention tag to set, so it does not implement Keeper. A caller (the // dashboard) relies on this boundary to decide whether to offer a keep at all. -func TestKeeperIsRemoteOnly(t *testing.T) { - rn, _ := NewRemoteNode("env", remote.Config{StartURL: "http://x", StopURL: "http://x", Region: "r"}) +func TestKeeperIsCloudOnly(t *testing.T) { + rn, _ := NewCloudNode("env", cloud.Config{StartURL: "http://x", StopURL: "http://x", Region: "r"}) if _, ok := rn.(Keeper); !ok { - t.Error("a remote node should implement Keeper") + t.Error("a cloud node should implement Keeper") } dn := &daemonNode{name: "dev", client: &Client{BaseURL: "http://127.0.0.1:1", Token: "t"}} if _, ok := any(dn).(Keeper); ok { @@ -429,11 +429,11 @@ func TestKeeperIsRemoteOnly(t *testing.T) { } } -// statusFromRemote carries the serving facts when the daemon reports them, and +// statusFromCloud carries the serving facts when the daemon reports them, and // leaves them empty when it does not — a running-but-unreachable daemon, or an // engine that is not running a model, must not be invented into serving one. -func TestStatusFromRemoteServingFacts(t *testing.T) { - with := statusFromRemote(remote.Response{ +func TestStatusFromCloudServingFacts(t *testing.T) { + with := statusFromCloud(cloud.Response{ State: "running", Runner: "llamacpp", ModelID: "org/m", @@ -442,37 +442,37 @@ func TestStatusFromRemoteServingFacts(t *testing.T) { if with.Runner != "llamacpp" || with.Model != "org/m" || with.ServedName != "m" { t.Errorf("serving facts should map across, got %+v", with) } - without := statusFromRemote(remote.Response{State: "running"}) + without := statusFromCloud(cloud.Response{State: "running"}) if without.Runner != "" || without.Model != "" || without.ServedName != "" { t.Errorf("an absent serving fact must stay empty, got %+v", without) } } -// statusFromRemote maps the control plane's own health check onto Ready the +// statusFromCloud maps the control plane's own health check onto Ready the // way a local daemon's reading is reported, so a router waiting for a -// remote engine to answer trusts it instead of a raw TCP probe — which can +// cloud engine to answer trusts it instead of a raw TCP probe — which can // succeed well before the model has finished loading. -func TestStatusFromRemoteMapsHealthyOntoReady(t *testing.T) { - if got := statusFromRemote(remote.Response{State: "running", Healthy: boolPtr(true)}); got.Ready != "ready" { +func TestStatusFromCloudMapsHealthyOntoReady(t *testing.T) { + if got := statusFromCloud(cloud.Response{State: "running", Healthy: boolPtr(true)}); got.Ready != "ready" { t.Errorf("healthy=true should map to Ready=%q, got %q", "ready", got.Ready) } - if got := statusFromRemote(remote.Response{State: "running", Healthy: boolPtr(false)}); got.Ready != "not-ready" { + if got := statusFromCloud(cloud.Response{State: "running", Healthy: boolPtr(false)}); got.Ready != "not-ready" { t.Errorf("healthy=false should map to Ready=%q, got %q", "not-ready", got.Ready) } // No healthy reading at all — an older control plane, or the branch // where the SSM agent is not yet reachable — leaves Ready empty rather // than claiming either answer. - if got := statusFromRemote(remote.Response{State: "running"}); got.Ready != "" { + if got := statusFromCloud(cloud.Response{State: "running"}); got.Ready != "" { t.Errorf("an absent healthy reading should leave Ready empty, got %q", got.Ready) } } -// statusFromRemote carries a running environment's engine address — the +// statusFromCloud carries a running environment's engine address — the // control plane's published base url — as the engine's host, so routing can // reach it the way it reaches any node. A stopped or undeployed environment // reports none, so its status carries no engine address. -func TestStatusFromRemoteCarriesTheEngineAddress(t *testing.T) { - got := statusFromRemote(remote.Response{ +func TestStatusFromCloudCarriesTheEngineAddress(t *testing.T) { + got := statusFromCloud(cloud.Response{ State: "running", BaseURL: "http://1.2.3.4:8000/v1", }) @@ -483,14 +483,14 @@ func TestStatusFromRemoteCarriesTheEngineAddress(t *testing.T) { t.Errorf("engine endpoint = %+v, want host 1.2.3.4 port 8000 path /v1", got.Engine) } // No base url — a stopped or undeployed environment — means no address. - if got := statusFromRemote(remote.Response{State: "stopped"}); got.Engine != nil { + if got := statusFromCloud(cloud.Response{State: "stopped"}); got.Engine != nil { t.Errorf("a stopped environment should carry no engine address: %+v", got.Engine) } } -// A remote node drives start, stop and metrics over its control plane exactly +// A cloud node drives start, stop and metrics over its control plane exactly // like a node would, mapping each reply onto the node's types. -func TestRemoteNodeStartStopMetricsOverTheControlPlane(t *testing.T) { +func TestCloudNodeStartStopMetricsOverTheControlPlane(t *testing.T) { stubAWSCreds(t) mux := http.NewServeMux() mux.HandleFunc("POST /start", func(w http.ResponseWriter, r *http.Request) { @@ -507,8 +507,8 @@ func TestRemoteNodeStartStopMetricsOverTheControlPlane(t *testing.T) { }) srv := httptest.NewServer(mux) t.Cleanup(srv.Close) - cfg := remote.Config{StartURL: srv.URL + "/start", StopURL: srv.URL + "/stop", StatsURL: srv.URL + "/stats", Region: "us-east-1"} - node, err := NewRemoteNode("env", cfg) + cfg := cloud.Config{StartURL: srv.URL + "/start", StopURL: srv.URL + "/stop", StatsURL: srv.URL + "/stats", Region: "us-east-1"} + node, err := NewCloudNode("env", cfg) if err != nil { t.Fatal(err) } @@ -533,8 +533,8 @@ func TestRemoteNodeStartStopMetricsOverTheControlPlane(t *testing.T) { // A node whose config has no stats endpoint reports that, rather than a silent // empty result: a control plane predating stats is a readable state. -func TestRemoteNodeMetricsWithoutStatsURL(t *testing.T) { - node, err := NewRemoteNode("env", remote.Config{StartURL: "http://x", StopURL: "http://x", Region: "r"}) +func TestCloudNodeMetricsWithoutStatsURL(t *testing.T) { + node, err := NewCloudNode("env", cloud.Config{StartURL: "http://x", StopURL: "http://x", Region: "r"}) if err != nil { t.Fatal(err) } @@ -546,18 +546,18 @@ func TestRemoteNodeMetricsWithoutStatsURL(t *testing.T) { // A node for an environment with no name cannot be pointed at a log stream, so a // log read fails before it touches CloudWatch. -func TestRemoteNodeLogsWithoutAnEnvironmentFails(t *testing.T) { - node, _ := NewRemoteNode("env", remote.Config{StartURL: "http://x", StopURL: "http://x", Region: "r"}) +func TestCloudNodeLogsWithoutAnEnvironmentFails(t *testing.T) { + node, _ := NewCloudNode("env", cloud.Config{StartURL: "http://x", StopURL: "http://x", Region: "r"}) _, err := node.Logs(context.Background(), daemon.TailLog, 100) if err == nil || !strings.Contains(err.Error(), "environment") { t.Errorf("a log read for an environment with no name should fail, got %v", err) } } -func TestRemoteNodeRejectedCallIsAOutcomeNotAFailure(t *testing.T) { +func TestCloudNodeRejectedCallIsAOutcomeNotAFailure(t *testing.T) { stubAWSCreds(t) - srv := remoteControlServer(t, `{"error":"boom"}`, http.StatusInternalServerError) - node, _ := NewRemoteNode("env", remote.Config{StartURL: srv.URL, StopURL: srv.URL, Region: "us-east-1"}) + srv := cloudControlServer(t, `{"error":"boom"}`, http.StatusInternalServerError) + node, _ := NewCloudNode("env", cloud.Config{StartURL: srv.URL, StopURL: srv.URL, Region: "us-east-1"}) r := StatusCall(context.Background(), node) // A rejected call is a typed outcome carrying the reason, not an error that // would blank the rest of the node set. @@ -569,7 +569,7 @@ func TestRemoteNodeRejectedCallIsAOutcomeNotAFailure(t *testing.T) { } } -// A mixed set — a local daemon node and a remote environment — is observed +// A mixed set — a local daemon node and a cloud environment — is observed // through the one fan-out, in order, and one failing member does not stop the // rest. func TestFanOutNodesOverAMixedSet(t *testing.T) { @@ -583,8 +583,8 @@ func TestFanOutNodesOverAMixedSet(t *testing.T) { t.Fatal(err) } - remoteUp := remoteControlServer(t, `{"state":"running","healthy":true}`, http.StatusOK) - rnode, _ := NewRemoteNode("env", remote.Config{StartURL: remoteUp.URL, StopURL: remoteUp.URL, Region: "us-east-1"}) + cloudUp := cloudControlServer(t, `{"state":"running","healthy":true}`, http.StatusOK) + rnode, _ := NewCloudNode("env", cloud.Config{StartURL: cloudUp.URL, StopURL: cloudUp.URL, Region: "us-east-1"}) results := FanOutNodes(context.Background(), StatusCall, []Node{dnode, rnode}) if len(results) != 2 { @@ -598,14 +598,14 @@ func TestFanOutNodesOverAMixedSet(t *testing.T) { t.Errorf("daemon node result = %+v", results[0]) } if !results[1].OK() || results[1].Status.State != "running" { - t.Errorf("remote node result = %+v", results[1]) + t.Errorf("cloud node result = %+v", results[1]) } } func TestFanOutNodesKeepsAGoodNodeWhenAnotherFails(t *testing.T) { stubAWSCreds(t) - good := remoteControlServer(t, `{"state":"running"}`, http.StatusOK) - rnode, _ := NewRemoteNode("env", remote.Config{StartURL: good.URL, StopURL: good.URL, Region: "us-east-1"}) + good := cloudControlServer(t, `{"state":"running"}`, http.StatusOK) + rnode, _ := NewCloudNode("env", cloud.Config{StartURL: good.URL, StopURL: good.URL, Region: "us-east-1"}) results := FanOutNodes(context.Background(), StatusCall, []Node{&failingNode{name: "bad"}, rnode}) if len(results) != 2 { @@ -644,7 +644,7 @@ func (n *failingNode) Logs(context.Context, int64, int) (daemon.LogsResponse, er // StartWithProgress carries the boot up to the caller as phases: the attempt // going out, then the instance coming up once a reply reports it. -func TestRemoteNodeStartWithProgressReportsTheBoot(t *testing.T) { +func TestCloudNodeStartWithProgressReportsTheBoot(t *testing.T) { stubAWSCreds(t) var ( mu sync.Mutex @@ -667,14 +667,14 @@ func TestRemoteNodeStartWithProgressReportsTheBoot(t *testing.T) { }) srv := httptest.NewServer(mux) t.Cleanup(srv.Close) - cfg := remote.Config{StartURL: srv.URL + "/start", StopURL: srv.URL + "/start", Region: "us-east-1"} - node, err := NewRemoteNode("env", cfg) + cfg := cloud.Config{StartURL: srv.URL + "/start", StopURL: srv.URL + "/start", Region: "us-east-1"} + node, err := NewCloudNode("env", cfg) if err != nil { t.Fatal(err) } starter, ok := node.(ProgressStarter) if !ok { - t.Fatal("the remote node does not carry progress") + t.Fatal("the cloud node does not carry progress") } resp, err := starter.StartWithProgress(context.Background(), func(p StartPhase) { mu.Lock() @@ -716,7 +716,7 @@ func TestRemoteNodeStartWithProgressReportsTheBoot(t *testing.T) { // dashboard tile) would show a capacity wait for the remainder of the start, // alongside its own refreshes reporting the instance running. The last phase a // start reports must therefore not be the capacity wait. -func TestRemoteNodeStartWithProgressRetiresACapacityWait(t *testing.T) { +func TestCloudNodeStartWithProgressRetiresACapacityWait(t *testing.T) { stubAWSCreds(t) var ( mu sync.Mutex @@ -739,8 +739,8 @@ func TestRemoteNodeStartWithProgressRetiresACapacityWait(t *testing.T) { }) srv := httptest.NewServer(mux) t.Cleanup(srv.Close) - cfg := remote.Config{StartURL: srv.URL + "/start", StopURL: srv.URL + "/start", Region: "us-east-1"} - node, err := NewRemoteNode("env", cfg) + cfg := cloud.Config{StartURL: srv.URL + "/start", StopURL: srv.URL + "/start", Region: "us-east-1"} + node, err := NewCloudNode("env", cfg) if err != nil { t.Fatal(err) } @@ -770,10 +770,10 @@ func TestRemoteNodeStartWithProgressRetiresACapacityWait(t *testing.T) { } } -// A nil progress callback is valid. remote.Start invokes progress on every +// A nil progress callback is valid. cloud.Start invokes progress on every // retry path, so StartWithProgress must substitute a no-op rather than depend // on which paths a given start takes; this exercises a start that retries. -func TestRemoteNodeStartWithProgressAcceptsNoReporter(t *testing.T) { +func TestCloudNodeStartWithProgressAcceptsNoReporter(t *testing.T) { stubAWSCreds(t) var attempts int mux := http.NewServeMux() @@ -789,8 +789,8 @@ func TestRemoteNodeStartWithProgressAcceptsNoReporter(t *testing.T) { }) srv := httptest.NewServer(mux) t.Cleanup(srv.Close) - cfg := remote.Config{StartURL: srv.URL + "/start", StopURL: srv.URL + "/start", Region: "us-east-1"} - node, err := NewRemoteNode("env", cfg) + cfg := cloud.Config{StartURL: srv.URL + "/start", StopURL: srv.URL + "/start", Region: "us-east-1"} + node, err := NewCloudNode("env", cfg) if err != nil { t.Fatal(err) } @@ -801,18 +801,18 @@ func TestRemoteNodeStartWithProgressAcceptsNoReporter(t *testing.T) { // A start the control plane refuses, and a stop it refuses, both surface as // errors the node's caller can show on its row. -func TestRemoteNodeRejectedOperationsSurfaceAsErrors(t *testing.T) { +func TestCloudNodeRejectedOperationsSurfaceAsErrors(t *testing.T) { stubAWSCreds(t) - srv := remoteControlServer(t, `{"error":"boom"}`, http.StatusInternalServerError) - cfg := remote.Config{StartURL: srv.URL, StopURL: srv.URL, Region: "us-east-1"} - node, err := NewRemoteNode("env", cfg) + srv := cloudControlServer(t, `{"error":"boom"}`, http.StatusInternalServerError) + cfg := cloud.Config{StartURL: srv.URL, StopURL: srv.URL, Region: "us-east-1"} + node, err := NewCloudNode("env", cfg) if err != nil { t.Fatal(err) } ctx := context.Background() starter, ok := node.(ProgressStarter) if !ok { - t.Fatal("the remote node does not carry progress") + t.Fatal("the cloud node does not carry progress") } if _, err := starter.StartWithProgress(ctx, nil); err == nil { t.Error("a rejected start did not fail") @@ -825,7 +825,7 @@ func TestRemoteNodeRejectedOperationsSurfaceAsErrors(t *testing.T) { // A failing log read surfaces as an error, not as an empty tail: the caller // renders the detail rather than reporting "no logs" when the read itself // went wrong. -func TestRemoteNodeLogsErrorsSurface(t *testing.T) { +func TestCloudNodeLogsErrorsSurface(t *testing.T) { stubAWSCreds(t) srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) { w.Header().Set("Content-Type", "application/json") @@ -834,8 +834,8 @@ func TestRemoteNodeLogsErrorsSurface(t *testing.T) { })) t.Cleanup(srv.Close) t.Setenv("AWS_ENDPOINT_URL_CLOUDWATCH_LOGS", srv.URL) - cfg := remote.Config{StartURL: "http://x", StopURL: "http://x", Environment: "env-1", Region: "us-east-1"} - node, err := NewRemoteNode("env", cfg) + cfg := cloud.Config{StartURL: "http://x", StopURL: "http://x", Environment: "env-1", Region: "us-east-1"} + node, err := NewCloudNode("env", cfg) if err != nil { t.Fatal(err) } @@ -846,7 +846,7 @@ func TestRemoteNodeLogsErrorsSurface(t *testing.T) { // A log store that answers with an empty tail is a readable state — the // engine has not logged here — so the node reports missing, not an error. -func TestRemoteNodeLogsEmptyTailIsMissingNotFailure(t *testing.T) { +func TestCloudNodeLogsEmptyTailIsMissingNotFailure(t *testing.T) { stubAWSCreds(t) srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) { w.Header().Set("Content-Type", "application/x-amz-json-1.1") @@ -854,8 +854,8 @@ func TestRemoteNodeLogsEmptyTailIsMissingNotFailure(t *testing.T) { })) t.Cleanup(srv.Close) t.Setenv("AWS_ENDPOINT_URL_CLOUDWATCH_LOGS", srv.URL) - cfg := remote.Config{StartURL: "http://x", StopURL: "http://x", Environment: "env-1", Region: "us-east-1"} - node, err := NewRemoteNode("env", cfg) + cfg := cloud.Config{StartURL: "http://x", StopURL: "http://x", Environment: "env-1", Region: "us-east-1"} + node, err := NewCloudNode("env", cfg) if err != nil { t.Fatal(err) } diff --git a/internal/fleet/config.go b/internal/fleet/config.go index c9e04bc5..42c5c3be 100644 --- a/internal/fleet/config.go +++ b/internal/fleet/config.go @@ -17,9 +17,9 @@ import ( "gopkg.in/yaml.v3" + "github.com/spinloop-ai/spinloop/internal/cloud" "github.com/spinloop-ai/spinloop/internal/daemon" "github.com/spinloop-ai/spinloop/internal/opencode" - "github.com/spinloop-ai/spinloop/internal/remote" ) // SplitTag divides a tag named the way tags are named in a limit or an item — @@ -43,9 +43,9 @@ const ( // KindDaemon is a machine running `spinloop daemon`, reached over its // control API. Addressed by its `host`. KindDaemon = "daemon" - // KindRemote is an `spinloop remote` environment, driven through its cloud + // KindCloud is an `spinloop cloud` environment, driven through its cloud // control plane. Addressed by the registered environment it names. - KindRemote = "remote" + KindCloud = "cloud" ) // Prefer is how routing ranks several nodes that could all serve a request. @@ -216,8 +216,8 @@ type Config struct { // the file's behaviour changes. Concurrency *Concurrency `yaml:"concurrency"` // APIKeyEnv names the environment variable holding the key this fleet's - // remote nodes require, shared by every one of them: a remote's engine is - // always gated by its key, so a fleet of remotes can name the variable + // cloud nodes require, shared by every one of them: a remote's engine is + // always gated by its key, so a fleet of clouds can name the variable // once rather than on each node — a node's own EngineTokenEnv overrides // it. It is a remote-only default: a daemon gates on its own // EngineTokenEnv, and a fleet-wide key must not start gating an engine @@ -235,10 +235,10 @@ type Config struct { // client talks to is a Node (see node.go); this is just the entry. type NodeConfig struct { // Name identifies the node in output and to `fleet start|stop `. For - // a kind-remote node it is also the key of the registered environment it - // drives, /remotes//remote.json — the environment is - // already user-named at `spinloop remote deploy`, so a remote node has no - // separate address to give. The control URLs live in that env's remote.json + // a kind-cloud node it is also the key of the registered environment it + // drives, /clouds//cloud.json — the environment is + // already user-named at `spinloop cloud deploy`, so a cloud node has no + // separate address to give. The control URLs live in that env's cloud.json // anyway, so nothing identifying a deployment is written into the fleet file. Name string `yaml:"name"` // Host is where the daemon answers — a LAN name, a tailscale name, or an @@ -263,7 +263,7 @@ type NodeConfig struct { // reached through a tunnel. Engine *EngineOverride `yaml:"engine"` // File names the Spinloop file that describes what this node runs — - // what `spinloop fleet deploy` reads to create a kind: remote node's + // what `spinloop fleet deploy` reads to create a kind: cloud node's // environment, and what `spinloop fleet start` reads to tell a kind: // daemon node's engine what to run. Resolved relative to the fleet // file's directory. Optional: a node's own Name is tried as a @@ -271,9 +271,9 @@ type NodeConfig struct { // beside the fleet file, before either command gives up on it. Not // read by any other fleet command. File string `yaml:"file"` - // InstanceType names the EC2 instance type a kind: remote node's + // InstanceType names the EC2 instance type a kind: cloud node's // environment launches as, read by `spinloop fleet deploy` into the - // deploy config it derives. It is a property of the remote environment + // deploy config it derives. It is a property of the cloud environment // only — a kind: daemon node's hardware is the operator's to choose, so // naming one there is a configuration error. Empty means the node's // environment launches as the control plane's default type. @@ -287,7 +287,7 @@ type NodeConfig struct { // WakePolicy overrides the fleet-wide wake policy for this node alone, // in the same `on`/`off` shape. Empty means the fleet-wide setting // decides for this node, as it always has. It exists because waking is - // not free the same way on every node — a remote environment's wake + // not free the same way on every node — a cloud environment's wake // boots a cloud instance, unlike a local daemon's engine — so an // operator may want to decide one node's waking on its own terms rather // than through a single fleet-wide switch. @@ -411,7 +411,7 @@ func Resolve(flagPath string) (*Config, error) { } // ForEnvironment builds the fleet a `--env ` target names: one cloud -// node, named by the registered environment whose remote.json holds its +// node, named by the registered environment whose cloud.json holds its // control config. A registered environment and a one-node fleet file naming it // describe the same thing — `fleet deploy` registers an environment under its // node's name, which is why every other fleet command can find one by name — @@ -431,17 +431,17 @@ func Resolve(flagPath string) (*Config, error) { // them describe how several nodes are used, which a fleet of one has no // occasion for. func ForEnvironment(name string) (*Config, error) { - if !remote.IsEnvName(name) { + if !cloud.IsEnvName(name) { return nil, fmt.Errorf( "%q is not an environment name: an environment name is a plain identifier, with no path", name) } // Loading it is the check: every environment resolves the one way, by // name, and a name that resolves to nothing fails here rather than at the // first control call. - if _, err := remote.LoadEnvironment(name, os.Getenv); err != nil { + if _, err := cloud.LoadEnvironment(name, os.Getenv); err != nil { return nil, err } - cfg := &Config{Nodes: []NodeConfig{{Name: name, Kind: KindRemote}}} + cfg := &Config{Nodes: []NodeConfig{{Name: name, Kind: KindCloud}}} // The same validation a parsed file gets, so a fleet of one cannot reach a // command in a state a fleet file could not. if err := cfg.validate(); err != nil { @@ -549,16 +549,16 @@ func (c *Config) validate() error { "node %q is kind %q: instance-type names the cloud environment's machine, and a daemon's hardware is the operator's to choose, not the fleet file's", n.Name, KindDaemon) } - case KindRemote: + case KindCloud: // The node's name *is* the registered environment's key, so it must // be env-shaped; a path-like name would be read as a registry // subdirectory rather than named. - if !remote.IsEnvName(n.Name) { + if !cloud.IsEnvName(n.Name) { return fmt.Errorf( "node %q is kind %q: its name must be a registered environment name (no /, no .json)", - n.Name, KindRemote) + n.Name, KindCloud) } - if n.InstanceType != "" && !remote.IsInstanceType(n.InstanceType) { + if n.InstanceType != "" && !cloud.IsInstanceType(n.InstanceType) { return fmt.Errorf( "node %q has instance-type %q, which is not shaped like an EC2 instance type (a family and size separated by a dot, e.g. g6e.xlarge)", n.Name, n.InstanceType) @@ -566,7 +566,7 @@ func (c *Config) validate() error { default: return fmt.Errorf( "node %q has kind %q: supported kinds are %q and %q", - n.Name, n.Kind, KindDaemon, KindRemote) + n.Name, n.Kind, KindDaemon, KindCloud) } } return nil @@ -638,19 +638,19 @@ func (c *Config) EngineToken(n NodeConfig) (string, error) { return c.resolveTokenEnv(fmt.Sprintf("node %q", n.Name), n.EngineTokenEnv) } -// RemoteEngineToken resolves the key a remote node's engine requires: the +// CloudEngineToken resolves the key a cloud node's engine requires: the // variable the node names when it names one, else the fleet-wide APIKeyEnv. // A remote's engine is always gated by its key, so a node that names no // resolvable key fails here, before a launch depends on it — the way every // other missing secret in this file is named, the node and the fix. -func (c *Config) RemoteEngineToken(n NodeConfig) (string, error) { +func (c *Config) CloudEngineToken(n NodeConfig) (string, error) { name := n.EngineTokenEnv if name == "" { name = c.APIKeyEnv } if name == "" { return "", fmt.Errorf( - "node %q is a remote environment, so its engine key must be set: name the variable holding it, in this node's `engineTokenEnv` or the file's fleet-wide `apiKeyEnv` (%s)", + "node %q is a cloud environment, so its engine key must be set: name the variable holding it, in this node's `engineTokenEnv` or the file's fleet-wide `apiKeyEnv` (%s)", n.Name, c.Path) } return c.resolveTokenEnv(fmt.Sprintf("node %q", n.Name), name) diff --git a/internal/fleet/config_test.go b/internal/fleet/config_test.go index 69a7b50a..2508f187 100644 --- a/internal/fleet/config_test.go +++ b/internal/fleet/config_test.go @@ -97,10 +97,10 @@ func TestLoadRejectsIncompleteNodes(t *testing.T) { "no nodes": "nodes: []\n", "no name": "nodes:\n - host: a.local\n", "no host": "nodes:\n - name: studio\n", - "remote name is a path": "nodes:\n - name: a/b\n kind: remote\n", - "remote name has .json": "nodes:\n - name: prod.json\n kind: remote\n", + "cloud name is a path": "nodes:\n - name: a/b\n kind: cloud\n", + "cloud name has .json": "nodes:\n - name: prod.json\n kind: cloud\n", "daemon instance-type": "nodes:\n - name: studio\n host: a.local\n instance-type: g6e.xlarge\n", - "remote bad instance-type": "nodes:\n - name: prod\n kind: remote\n instance-type: g6exlarge\n", + "remote bad instance-type": "nodes:\n - name: prod\n kind: cloud\n instance-type: g6exlarge\n", } { t.Run(name, func(t *testing.T) { if _, err := Load(writeFleet(t, body, "")); err == nil { @@ -118,34 +118,34 @@ func TestLoadUnknownKindNamesIt(t *testing.T) { } } -// A kind-remote node's name is the registered environment it drives; it needs +// A kind-cloud node's name is the registered environment it drives; it needs // no host, and nothing else to name it with. -func TestLoadRemoteKindNamedByEnvironment(t *testing.T) { +func TestLoadCloudKindNamedByEnvironment(t *testing.T) { path := writeFleet(t, ` nodes: - name: prod - kind: remote + kind: cloud `, "") cfg, err := Load(path) if err != nil { t.Fatal(err) } n := cfg.Nodes[0] - if n.Kind != KindRemote || n.Name != "prod" { - t.Errorf("remote node = %+v, want kind remote named prod", n) + if n.Kind != KindCloud || n.Name != "prod" { + t.Errorf("cloud node = %+v, want kind remote named prod", n) } if n.Host != "" { - t.Errorf("remote node needs no host, got %q", n.Host) + t.Errorf("cloud node needs no host, got %q", n.Host) } } -// A kind-remote node may name the instance type its environment launches as; +// A kind-cloud node may name the instance type its environment launches as; // the field parses onto the node so `fleet deploy` can read it. -func TestLoadRemoteKindInstanceType(t *testing.T) { +func TestLoadCloudKindInstanceType(t *testing.T) { path := writeFleet(t, ` nodes: - name: prod - kind: remote + kind: cloud instance-type: g6e.2xlarge `, "") cfg, err := Load(path) @@ -157,17 +157,17 @@ nodes: } } -// A remote node naming a malformed instance type is a configuration error that +// A cloud node naming a malformed instance type is a configuration error that // names both the node and the value, not a silent deploy of junk. -func TestLoadRemoteKindBadInstanceTypeNamesIt(t *testing.T) { +func TestLoadCloudKindBadInstanceTypeNamesIt(t *testing.T) { _, err := Load(writeFleet(t, ` nodes: - name: prod - kind: remote + kind: cloud instance-type: g6exlarge `, "")) if err == nil { - t.Fatal("accepted a remote node with a malformed instance-type") + t.Fatal("accepted a cloud node with a malformed instance-type") } msg := err.Error() if !strings.Contains(msg, "prod") || !strings.Contains(msg, "g6exlarge") { @@ -404,14 +404,14 @@ nodes: // The fleet-wide key is a remote-only default: resolved like every other // secret in the file, a node's own reference overrides it, and a remote that // names no resolvable key is named for it. -func TestRemoteEngineTokenResolution(t *testing.T) { +func TestCloudEngineTokenResolution(t *testing.T) { path := writeFleet(t, ` apiKeyEnv: FLEET_KEY nodes: - name: shared - kind: remote + kind: cloud - name: own - kind: remote + kind: cloud engineTokenEnv: OWN_KEY - name: box host: box.local @@ -423,18 +423,18 @@ nodes: shared, _ := cfg.Node("shared") // The fleet's variable, from the .env beside the file. - if got, err := cfg.RemoteEngineToken(shared); err != nil || got != "from-dotenv" { + if got, err := cfg.CloudEngineToken(shared); err != nil || got != "from-dotenv" { t.Errorf("key = %q, %v; want the fleet .env value", got, err) } // An exported value wins, as everywhere else in spinloop. t.Setenv("FLEET_KEY", "exported") - if got, err := cfg.RemoteEngineToken(shared); err != nil || got != "exported" { + if got, err := cfg.CloudEngineToken(shared); err != nil || got != "exported" { t.Errorf("key = %q, %v; want the exported value", got, err) } // A node's own reference overrides the fleet-wide one. own, _ := cfg.Node("own") - if got, err := cfg.RemoteEngineToken(own); err != nil || got != "own-dotenv" { + if got, err := cfg.CloudEngineToken(own); err != nil || got != "own-dotenv" { t.Errorf("key = %q, %v; want the node's own value", got, err) } @@ -446,19 +446,19 @@ nodes: } } -func TestRemoteEngineTokenUnsetNamesTheVariable(t *testing.T) { +func TestCloudEngineTokenUnsetNamesTheVariable(t *testing.T) { path := writeFleet(t, ` apiKeyEnv: NOWHERE_FLEET_KEY nodes: - name: shared - kind: remote + kind: cloud `, "") cfg, err := Load(path) if err != nil { t.Fatal(err) } node, _ := cfg.Node("shared") - _, err = cfg.RemoteEngineToken(node) + _, err = cfg.CloudEngineToken(node) if err == nil { t.Fatal("an unset fleet key variable should be a config error") } @@ -469,18 +469,18 @@ nodes: } } -func TestRemoteEngineTokenMissingNamesBothPlaces(t *testing.T) { +func TestCloudEngineTokenMissingNamesBothPlaces(t *testing.T) { path := writeFleet(t, ` nodes: - name: shared - kind: remote + kind: cloud `, "") cfg, err := Load(path) if err != nil { t.Fatal(err) } node, _ := cfg.Node("shared") - _, err = cfg.RemoteEngineToken(node) + _, err = cfg.CloudEngineToken(node) if err == nil { t.Fatal("a remote naming no key anywhere should be a config error") } @@ -499,7 +499,7 @@ func TestLoadIgnoresUnknownFleetFields(t *testing.T) { apiKeyEnv: SHARED_KEY nodes: - name: shared - kind: remote + kind: cloud `, "") data, err := os.ReadFile(path) if err != nil { @@ -549,7 +549,7 @@ func TestFileField(t *testing.T) { path := writeFleet(t, ` nodes: - name: gpu-env - kind: remote + kind: cloud file: ./envs/gpu.Spinloop - name: dev-1 host: dev1.local @@ -561,9 +561,9 @@ nodes: if err != nil { t.Fatal(err) } - remote, _ := cfg.Node("gpu-env") - if remote.File != "./envs/gpu.Spinloop" { - t.Errorf("remote node File = %q", remote.File) + cloud, _ := cfg.Node("gpu-env") + if cloud.File != "./envs/gpu.Spinloop" { + t.Errorf("cloud node File = %q", cloud.File) } daemonNode, _ := cfg.Node("dev-1") if daemonNode.File != "../shared/dev.Spinloop" { diff --git a/internal/fleet/fanout.go b/internal/fleet/fanout.go index 16a20617..fbb8de44 100644 --- a/internal/fleet/fanout.go +++ b/internal/fleet/fanout.go @@ -146,7 +146,7 @@ func (c *Config) FanOut(ctx context.Context, call Call) []NodeResult { // FanOutNodes runs call over an explicit set of nodes concurrently and returns // one result per node, in the order the set is given. It is the seam that lets an // observable be driven regardless of where its nodes come from — a fleet file's -// daemon nodes, a remote environment, or a mix — through the one fan-out the rest +// daemon nodes, a cloud environment, or a mix — through the one fan-out the rest // of the client already shares. // // As with Config.FanOut it never returns an error: a node that cannot be reached diff --git a/internal/fleet/fleet_remote_test.go b/internal/fleet/fleet_cloud_test.go similarity index 65% rename from internal/fleet/fleet_remote_test.go rename to internal/fleet/fleet_cloud_test.go index 024157a3..3b1d8f0e 100644 --- a/internal/fleet/fleet_remote_test.go +++ b/internal/fleet/fleet_cloud_test.go @@ -9,34 +9,34 @@ import ( "testing" ) -// registerRemoteEnv points the environment registry (SPINLOOP_CONFIG_DIR) at a -// temp config directory and writes one environment's remote.json, whose control +// registerCloudEnv points the environment registry (SPINLOOP_CONFIG_DIR) at a +// temp config directory and writes one environment's cloud.json, whose control // plane is the start and stop servers it is handed. It returns the name a // kind-remote fleet node uses to refer to it. -func registerRemoteEnv(t *testing.T, name, startURL, stopURL string) { +func registerCloudEnv(t *testing.T, name, startURL, stopURL string) { t.Helper() home := t.TempDir() t.Setenv("SPINLOOP_CONFIG_DIR", home) - envDir := filepath.Join(home, "remotes", name) + envDir := filepath.Join(home, "clouds", name) if err := os.MkdirAll(envDir, 0o755); err != nil { t.Fatal(err) } body := fmt.Sprintf( `{"start_url":%q,"stop_url":%q,"region":"us-east-1","environment":%q}`, startURL, stopURL, name) - if err := os.WriteFile(filepath.Join(envDir, "remote.json"), []byte(body), 0o600); err != nil { + if err := os.WriteFile(filepath.Join(envDir, "cloud.json"), []byte(body), 0o600); err != nil { t.Fatal(err) } } -// A kind-remote node in the fleet file is built by the regular NewNode and +// A kind-cloud node in the fleet file is built by the regular NewNode and // observed through the regular fan-out, resolved from the environment registry. func TestFanOutLoadsARemoteNodeFromTheFile(t *testing.T) { stubAWSCreds(t) - up := remoteControlServer(t, `{"state":"running","healthy":true}`, http.StatusOK) - registerRemoteEnv(t, "prod", up.URL, up.URL) + up := cloudControlServer(t, `{"state":"running","healthy":true}`, http.StatusOK) + registerCloudEnv(t, "prod", up.URL, up.URL) - path := writeFleet(t, "nodes:\n - name: prod\n kind: remote\n", "") + path := writeFleet(t, "nodes:\n - name: prod\n kind: cloud\n", "") cfg, err := Load(path) if err != nil { t.Fatal(err) @@ -46,19 +46,19 @@ func TestFanOutLoadsARemoteNodeFromTheFile(t *testing.T) { t.Fatalf("got %d results, want 1", len(r)) } if !r[0].OK() || r[0].Status.State != "running" { - t.Errorf("remote node via the file = %+v (err %v)", r[0], r[0].Err) + t.Errorf("cloud node via the file = %+v (err %v)", r[0], r[0].Err) } } // A node for an environment that is not registered cannot be built, so the // fan-out reports it as a config error — naming the environment — rather than // blanking the view or surfacing a 401. -func TestFanOutRemotesANUnregisteredEnvAsAConfigError(t *testing.T) { +func TestFanOutCloudsANUnregisteredEnvAsAConfigError(t *testing.T) { stubAWSCreds(t) // A registry directory with no such environment in it. t.Setenv("SPINLOOP_CONFIG_DIR", t.TempDir()) - path := writeFleet(t, "nodes:\n - name: prod\n kind: remote\n", "") + path := writeFleet(t, "nodes:\n - name: prod\n kind: cloud\n", "") cfg, err := Load(path) if err != nil { t.Fatal(err) @@ -75,22 +75,22 @@ func TestFanOutRemotesANUnregisteredEnvAsAConfigError(t *testing.T) { } } -// A fleet file that mixes a local daemon node and a remote environment is +// A fleet file that mixes a local daemon node and a cloud environment is // observed through the one fan-out, in file order, as the same kind of row. func TestFanOutOverAMixedFile(t *testing.T) { stubAWSCreds(t) - remoteUp := remoteControlServer(t, `{"state":"running","healthy":true}`, http.StatusOK) - registerRemoteEnv(t, "prod", remoteUp.URL, remoteUp.URL) + cloudUp := cloudControlServer(t, `{"state":"running","healthy":true}`, http.StatusOK) + registerCloudEnv(t, "prod", cloudUp.URL, cloudUp.URL) // A daemon node, built the way the daemon tests build it. cfg := fleetFor(t, stubDaemon(t, "", "running"), "") - cfg.Nodes = append(cfg.Nodes, NodeConfig{Name: "prod", Kind: KindRemote}) + cfg.Nodes = append(cfg.Nodes, NodeConfig{Name: "prod", Kind: KindCloud}) r := cfg.FanOut(context.Background(), StatusCall) if len(r) != 2 { t.Fatalf("got %d results, want 2", len(r)) } - // File order: the daemon node comes first, the remote environment after. + // File order: the daemon node comes first, the cloud environment after. if r[0].Name != "box" || r[1].Name != "prod" { t.Fatalf("order = %q, %q; want box then prod", r[0].Name, r[1].Name) } @@ -98,6 +98,6 @@ func TestFanOutOverAMixedFile(t *testing.T) { t.Errorf("daemon node via the file = %+v (err %v)", r[0], r[0].Err) } if !r[1].OK() || r[1].Status.State != "running" { - t.Errorf("remote node via the file = %+v (err %v)", r[1], r[1].Err) + t.Errorf("cloud node via the file = %+v (err %v)", r[1], r[1].Err) } } diff --git a/internal/fleet/for_environment_test.go b/internal/fleet/for_environment_test.go index bd47f74b..9975ee97 100644 --- a/internal/fleet/for_environment_test.go +++ b/internal/fleet/for_environment_test.go @@ -13,8 +13,8 @@ import ( // cloud node the same name in a fleet file would build. func TestForEnvironmentBuildsAFleetOfOne(t *testing.T) { stubAWSCreds(t) - up := remoteControlServer(t, `{"state":"running","healthy":true}`, http.StatusOK) - registerRemoteEnv(t, "prod", up.URL, up.URL) + up := cloudControlServer(t, `{"state":"running","healthy":true}`, http.StatusOK) + registerCloudEnv(t, "prod", up.URL, up.URL) cfg, err := ForEnvironment("prod") if err != nil { @@ -26,8 +26,8 @@ func TestForEnvironmentBuildsAFleetOfOne(t *testing.T) { if cfg.Nodes[0].Name != "prod" { t.Errorf("name = %q, want prod", cfg.Nodes[0].Name) } - if cfg.Nodes[0].Kind != KindRemote { - t.Errorf("kind = %q, want %q", cfg.Nodes[0].Kind, KindRemote) + if cfg.Nodes[0].Kind != KindCloud { + t.Errorf("kind = %q, want %q", cfg.Nodes[0].Kind, KindCloud) } // The node builds and answers, which is the whole claim: the fan-out and // the renderers need nothing special for a fleet assembled this way. @@ -45,10 +45,10 @@ func TestForEnvironmentBuildsAFleetOfOne(t *testing.T) { // environment was named. func TestForEnvironmentMatchesTheSameNodeInAFile(t *testing.T) { stubAWSCreds(t) - up := remoteControlServer(t, `{"state":"running","healthy":true}`, http.StatusOK) - registerRemoteEnv(t, "prod", up.URL, up.URL) + up := cloudControlServer(t, `{"state":"running","healthy":true}`, http.StatusOK) + registerCloudEnv(t, "prod", up.URL, up.URL) - path := writeFleet(t, "nodes:\n - name: prod\n kind: remote\n", "") + path := writeFleet(t, "nodes:\n - name: prod\n kind: cloud\n", "") fromFile, err := Load(path) if err != nil { t.Fatal(err) @@ -79,11 +79,11 @@ func TestForEnvironmentRejects(t *testing.T) { env string want string }{ - {"a path", "./remote.json", "plain identifier"}, + {"a path", "./cloud.json", "plain identifier"}, {"a nested path", "envs/prod", "plain identifier"}, {"a json file", "prod.json", "plain identifier"}, {"empty", "", "plain identifier"}, - {"unregistered", "nope", "remotes/nope/remote.json"}, + {"unregistered", "nope", "clouds/nope/cloud.json"}, } for _, tt := range tests { t.Run(tt.name, func(t *testing.T) { @@ -106,7 +106,7 @@ func TestForEnvironmentUnregisteredNamesTheFix(t *testing.T) { if err == nil { t.Fatal("want an error") } - if !strings.Contains(err.Error(), "remotes/nope/remote.json") { + if !strings.Contains(err.Error(), "clouds/nope/cloud.json") { t.Errorf("error %q does not name the environment's registry path", err) } } @@ -116,8 +116,8 @@ func TestForEnvironmentUnregisteredNamesTheFix(t *testing.T) { // than inherited from anywhere. func TestForEnvironmentCarriesNoFileOrFleetWideSettings(t *testing.T) { stubAWSCreds(t) - up := remoteControlServer(t, `{"state":"running"}`, http.StatusOK) - registerRemoteEnv(t, "prod", up.URL, up.URL) + up := cloudControlServer(t, `{"state":"running"}`, http.StatusOK) + registerCloudEnv(t, "prod", up.URL, up.URL) cfg, err := ForEnvironment("prod") if err != nil { @@ -156,8 +156,8 @@ func TestForEnvironmentCarriesNoFileOrFleetWideSettings(t *testing.T) { // the target is the environment, and the file is not consulted. func TestForEnvironmentIgnoresADirectoryFleetFile(t *testing.T) { stubAWSCreds(t) - up := remoteControlServer(t, `{"state":"running"}`, http.StatusOK) - registerRemoteEnv(t, "prod", up.URL, up.URL) + up := cloudControlServer(t, `{"state":"running"}`, http.StatusOK) + registerCloudEnv(t, "prod", up.URL, up.URL) dir := t.TempDir() body := "prefer: active\nnodes:\n - name: other\n host: elsewhere\n" @@ -182,8 +182,8 @@ func TestForEnvironmentIgnoresADirectoryFleetFile(t *testing.T) { // never looks for a .env — which is what lets it carry no directory at all. func TestForEnvironmentReadsNoAdjacentEnvFile(t *testing.T) { stubAWSCreds(t) - up := remoteControlServer(t, `{"state":"running"}`, http.StatusOK) - registerRemoteEnv(t, "prod", up.URL, up.URL) + up := cloudControlServer(t, `{"state":"running"}`, http.StatusOK) + registerCloudEnv(t, "prod", up.URL, up.URL) dir := t.TempDir() // A .env that would be read if the lookup ever ran, holding a value no diff --git a/internal/fleet/node.go b/internal/fleet/node.go index cae44afe..a4420177 100644 --- a/internal/fleet/node.go +++ b/internal/fleet/node.go @@ -7,10 +7,10 @@ import ( "os" "time" + "github.com/spinloop-ai/spinloop/internal/cloud" "github.com/spinloop-ai/spinloop/internal/daemon" "github.com/spinloop-ai/spinloop/internal/inference" "github.com/spinloop-ai/spinloop/internal/metrics" - "github.com/spinloop-ai/spinloop/internal/remote" ) // Outcome classifies how a node call ended. A fleet view renders these as rows @@ -154,7 +154,7 @@ type Instance struct { } // Node is one member of the fleet. Only daemonNode implements it today; the -// interface exists so a remote-environment kind (an `spinloop remote` +// interface exists so a remote-environment kind (an `spinloop cloud` // environment read through its stats Lambda, which already yields // metrics.Stats) can be added without reworking the fan-out or the renderers. type Node interface { @@ -207,18 +207,18 @@ func (n *daemonNode) Logs(ctx context.Context, offset int64, limit int) (daemon. // NewNode builds the live Node for one fleet-file entry. A daemon node resolves // its bearer token here — a reference that resolves to nothing fails before any // call is attempted, naming the variable rather than surfacing later as a 401. -// A remote node loads its registered environment's control config; a missing +// A cloud node loads its registered environment's control config; a missing // environment fails the same way, as a per-node error the fan-out renders as a // row rather than a blanked view. func (c *Config) NewNode(entry NodeConfig) (Node, error) { - if entry.Kind == KindRemote { + if entry.Kind == KindCloud { // The node's name is the registered environment's key, and every // environment resolves the one way: by name, from the registry. - cfg, err := remote.LoadEnvironment(entry.Name, os.Getenv) + cfg, err := cloud.LoadEnvironment(entry.Name, os.Getenv) if err != nil { return nil, err } - return NewRemoteNode(entry.Name, cfg) + return NewCloudNode(entry.Name, cfg) } token, err := c.Token(entry) if err != nil { diff --git a/internal/fleet/node_test.go b/internal/fleet/node_test.go index ddde7cbf..a6903ac5 100644 --- a/internal/fleet/node_test.go +++ b/internal/fleet/node_test.go @@ -275,12 +275,12 @@ func TestResultCarriesTheStatusAndTheVerdict(t *testing.T) { // NewNode for a remote entry loads the registered environment's config, and // whatever is missing is a per-node error the way every other missing thing // is: no config directory to find it in, or the environment never registered. -func TestNewNodeForAUnregisteredRemoteEnvironment(t *testing.T) { +func TestNewNodeForAUnregisteredCloudEnvironment(t *testing.T) { cfg := &Config{} t.Setenv("SPINLOOP_CONFIG_DIR", "") t.Setenv("XDG_CONFIG_HOME", "") t.Setenv("HOME", "") - if _, err := cfg.NewNode(NodeConfig{Name: "env", Kind: KindRemote}); err == nil { + if _, err := cfg.NewNode(NodeConfig{Name: "env", Kind: KindCloud}); err == nil { t.Error("a remote entry with no config directory should fail") } // With a registry to look in, an unregistered environment fails the node, @@ -288,7 +288,7 @@ func TestNewNodeForAUnregisteredRemoteEnvironment(t *testing.T) { // one simply finds nothing. t.Setenv("SPINLOOP_CONFIG_DIR", t.TempDir()) for _, name := range []string{"env", "a/b"} { - if _, err := cfg.NewNode(NodeConfig{Name: name, Kind: KindRemote}); err == nil { + if _, err := cfg.NewNode(NodeConfig{Name: name, Kind: KindCloud}); err == nil { t.Errorf("unregistered environment %q built a node", name) } } diff --git a/internal/fleet/select.go b/internal/fleet/select.go index e8fe66f6..5169bc88 100644 --- a/internal/fleet/select.go +++ b/internal/fleet/select.go @@ -406,7 +406,7 @@ func (c *Config) EngineBaseURL(n NodeConfig, status daemon.StatusResponse) (stri host, port, path := n.Host, 0, "" if ep := status.Engine; ep != nil { port, path = ep.Port, ep.Path - // A node that reports its engine's host — a remote environment, whose + // A node that reports its engine's host — a cloud environment, whose // control plane knows the instance's published address — is reached // there, in place of the host the fleet file supplies. if ep.Host != "" { @@ -484,8 +484,8 @@ func hostIsLoopback(host string) bool { // the control plane reports the instance, never the gate, so its key is looked // up whatever the status says — from the node's own reference or the fleet's. func (c *Config) engineKeyFor(n NodeConfig, status daemon.StatusResponse) (string, error) { - if n.Kind == KindRemote { - return c.RemoteEngineToken(n) + if n.Kind == KindCloud { + return c.CloudEngineToken(n) } if status.Engine == nil || !status.Engine.RequiresKey { return "", nil diff --git a/internal/fleet/select_test.go b/internal/fleet/select_test.go index a365b21b..8e4fa326 100644 --- a/internal/fleet/select_test.go +++ b/internal/fleet/select_test.go @@ -301,7 +301,7 @@ func TestEngineBaseURL(t *testing.T) { }, { name: "a reported engine host is used in place of the fleet file's", - node: NodeConfig{Name: "env", Kind: "remote"}, + node: NodeConfig{Name: "env", Kind: "cloud"}, status: daemon.StatusResponse{ Engine: &daemon.EngineEndpoint{Host: "1.2.3.4", Port: 8000, Path: "/v1"}, }, @@ -309,7 +309,7 @@ func TestEngineBaseURL(t *testing.T) { }, { name: "an override still beats a reported engine host", - node: NodeConfig{Name: "env", Kind: "remote", Engine: &EngineOverride{Host: "proxy"}}, + node: NodeConfig{Name: "env", Kind: "cloud", Engine: &EngineOverride{Host: "proxy"}}, status: daemon.StatusResponse{ Engine: &daemon.EngineEndpoint{Host: "1.2.3.4", Port: 8000, Path: "/v1"}, }, @@ -329,7 +329,7 @@ func TestEngineBaseURL(t *testing.T) { } } -// A loopback-bound engine on a remote node is refused with both remedies, +// A loopback-bound engine on a cloud node is refused with both remedies, // rather than handed over as an address that cannot connect. func TestLoopbackEngineIsRefused(t *testing.T) { cfg := &Config{Path: "fleet.yaml"} @@ -337,7 +337,7 @@ func TestLoopbackEngineIsRefused(t *testing.T) { _, err := cfg.EngineBaseURL(NodeConfig{Name: "gpu", Host: "gpu-box"}, status) if err == nil { - t.Fatal("a loopback engine on a remote node should be refused") + t.Fatal("a loopback engine on a cloud node should be refused") } for _, want := range []string{"gpu", "loopback", "--host", "fleet.yaml"} { if !strings.Contains(err.Error(), want) { @@ -410,7 +410,7 @@ func TestEngineKeyResolution(t *testing.T) { // instance, never the gate — so its key is looked up whatever the status says, // from the node's own reference or the fleet's, and a remote with no resolvable // key fails before a launch depends on it. -func TestRemoteEngineKeyResolution(t *testing.T) { +func TestCloudEngineKeyResolution(t *testing.T) { cfg := &Config{Path: "fleet.yaml", Dir: t.TempDir(), APIKeyEnv: "FLEET_KEY"} t.Setenv("FLEET_KEY", "sk-fleet") t.Setenv("NODE_ENGINE_KEY", "sk-node") @@ -418,20 +418,20 @@ func TestRemoteEngineKeyResolution(t *testing.T) { empty := daemon.StatusResponse{} // The fleet's key, by default. - remote := NodeConfig{Name: "cloud", Kind: KindRemote} + remote := NodeConfig{Name: "cloud", Kind: KindCloud} if key, err := cfg.engineKeyFor(remote, empty); err != nil || key != "sk-fleet" { t.Errorf("key = %q, %v; want sk-fleet", key, err) } // The node's own reference overrides it. - own := NodeConfig{Name: "cloud", Kind: KindRemote, EngineTokenEnv: "NODE_ENGINE_KEY"} + own := NodeConfig{Name: "cloud", Kind: KindCloud, EngineTokenEnv: "NODE_ENGINE_KEY"} if key, err := cfg.engineKeyFor(own, empty); err != nil || key != "sk-node" { t.Errorf("key = %q, %v; want sk-node", key, err) } // No key named anywhere fails, naming the node and both places to fix it. cfg.APIKeyEnv = "" - _, err := cfg.engineKeyFor(NodeConfig{Name: "cloud", Kind: KindRemote}, empty) + _, err := cfg.engineKeyFor(NodeConfig{Name: "cloud", Kind: KindCloud}, empty) if err == nil { t.Fatal("a remote with no key named should fail") } diff --git a/internal/fleet/start_phase.go b/internal/fleet/start_phase.go index 6598f2ba..21071ff2 100644 --- a/internal/fleet/start_phase.go +++ b/internal/fleet/start_phase.go @@ -16,7 +16,7 @@ import ( "strings" "time" - "github.com/spinloop-ai/spinloop/internal/remote" + "github.com/spinloop-ai/spinloop/internal/cloud" ) // StartPhaseKind identifies what a start is currently doing. Exactly one value @@ -66,7 +66,7 @@ const ( // RenderPhase is the phase's line at time now. A wait counts down towards // RetryAt and a boot counts up from Since, so nothing here is fixed when the -// phase is built. The dashboard tile and `spinloop remote start` both draw +// phase is built. The dashboard tile and `spinloop cloud start` both draw // their line from this, so the two cannot word one phase differently. func RenderPhase(p StartPhase, now time.Time) string { switch p.Kind { @@ -126,15 +126,15 @@ func formatPhaseDuration(d time.Duration) string { } } -// StartPhases adapts remote.Start's progress and onState callbacks onto a -// stream of phases: it returns the pair to hand remote.Start, and calls report +// StartPhases adapts cloud.Start's progress and onState callbacks onto a +// stream of phases: it returns the pair to hand cloud.Start, and calls report // once per transition. The two callbacks are separate and neither carries a // phase on its own — onState carries the state of a reply, and the progress // line that follows a 503 carries that reply's retry-after — so the mapping // holds the state between them. // -// It is here rather than in either caller because both `spinloop remote start` -// and the dashboard drive remote.Start and render the result. +// It is here rather than in either caller because both `spinloop cloud start` +// and the dashboard drive cloud.Start and render the result. func StartPhases(report func(StartPhase)) (progress func(string), onState func(string)) { t := &startPhases{report: report, now: time.Now} return t.progress, t.state @@ -168,7 +168,7 @@ func (t *startPhases) enter(kind StartPhaseKind, detail string) { // state maps one reply's state onto a phase. func (t *startPhases) state(s string) { switch { - case s == remote.StateInFlight: + case s == cloud.StateInFlight: // A fresh attempt supersedes a capacity wait and a dropped // connection: each described the attempt before it. It does not // supersede a boot — once a reply has reported the instance coming @@ -185,7 +185,7 @@ func (t *startPhases) state(s string) { // that is over before it can be read. case s == stateNoCapacity: // The reply's retry-after reaches this caller only on the progress - // line remote.Start writes next, so the wait carries no due time + // line cloud.Start writes next, so the wait carries no due time // until that line arrives. t.enter(PhaseWaitingCapacity, s) case s == stateSeeding: @@ -198,11 +198,11 @@ func (t *startPhases) state(s string) { } } -// droppedPrefix is how remote.Start opens the line it writes when an attempt's +// droppedPrefix is how cloud.Start opens the line it writes when an attempt's // connection drops mid-request. const droppedPrefix = "connection dropped" -// progress maps one of remote.Start's status lines onto a phase. The lines are +// progress maps one of cloud.Start's status lines onto a phase. The lines are // the only place the 503's retry-after and a transport error reach this // caller; which state a retry line refers to came through onState immediately // before it, so only the delay is read off the line itself. @@ -222,12 +222,12 @@ func (t *startPhases) progress(line string) { } } -// retryInMarker precedes the delay in every line remote.Start writes before a +// retryInMarker precedes the delay in every line cloud.Start writes before a // wait. const retryInMarker = "retrying in " // parseRetryIn reads the delay a status line names, in either of the forms -// remote.Start writes it — a whole number of seconds from the reply's +// cloud.Start writes it — a whole number of seconds from the reply's // retry-after, or a Go duration for the fixed wait after a dropped connection. // A line naming no delay reports false, and the phase then carries no due time // rather than a wrong one. diff --git a/internal/fleet/start_phase_test.go b/internal/fleet/start_phase_test.go index f0167034..35917453 100644 --- a/internal/fleet/start_phase_test.go +++ b/internal/fleet/start_phase_test.go @@ -5,7 +5,7 @@ import ( "testing" "time" - "github.com/spinloop-ai/spinloop/internal/remote" + "github.com/spinloop-ai/spinloop/internal/cloud" ) // Every line a phase renders is computed from the phase and the time it is @@ -63,12 +63,12 @@ func TestStartPhasesSeeding(t *testing.T) { var got []StartPhase progress, onState := StartPhases(func(p StartPhase) { got = append(got, p) }) - onState(remote.StateInFlight) + onState(cloud.StateInFlight) onState(stateSeeding) progress("seeding the weights (seed llamacpp--org-model--Q4_K_M); retrying in 60s") // The next attempt supersedes nothing: it polls the fetch the seeding // reply reported, and the reply that follows reports the same fetch. - onState(remote.StateInFlight) + onState(cloud.StateInFlight) onState(stateSeeding) want := []StartPhaseKind{PhaseAttempting, PhaseSeeding} @@ -110,11 +110,11 @@ func TestStartPhasesHandsTheSeedingOffToTheBoot(t *testing.T) { var got []StartPhase progress, onState := StartPhases(func(p StartPhase) { got = append(got, p) }) - onState(remote.StateInFlight) + onState(cloud.StateInFlight) onState(stateSeeding) progress("seeding the weights (seed llamacpp--org-model--Q4_K_M); retrying in 60s") // The weights are in: the next reply reports the instance coming up. - onState(remote.StateInFlight) + onState(cloud.StateInFlight) onState("starting") progress("instance starting; retrying in 5s") @@ -141,7 +141,7 @@ func TestStartPhasesHandsTheSeedingOffToTheBoot(t *testing.T) { } } -// The mapping from remote.Start's two callbacks onto phases: an attempt goes +// The mapping from cloud.Start's two callbacks onto phases: an attempt goes // out, is refused for capacity with a due time for the next one, and the // attempt that follows retires the refusal rather than leaving it standing // while the instance boots — the defect this phase stream exists for. @@ -149,10 +149,10 @@ func TestStartPhasesRetiresACapacityWait(t *testing.T) { var got []StartPhase progress, onState := StartPhases(func(p StartPhase) { got = append(got, p) }) - onState(remote.StateInFlight) + onState(cloud.StateInFlight) onState("no-capacity") progress("instance no-capacity; retrying in 120s") - onState(remote.StateInFlight) + onState(cloud.StateInFlight) kinds := make([]StartPhaseKind, len(got)) for i, p := range got { @@ -185,10 +185,10 @@ func TestStartPhasesHoldsTheBootAcrossPolls(t *testing.T) { var got []StartPhase progress, onState := StartPhases(func(p StartPhase) { got = append(got, p) }) - onState(remote.StateInFlight) + onState(cloud.StateInFlight) onState("starting") progress("instance starting; retrying in 5s") - onState(remote.StateInFlight) + onState(cloud.StateInFlight) onState("starting") if len(got) != 2 { @@ -205,7 +205,7 @@ func TestStartPhasesReportsADroppedConnection(t *testing.T) { var got []StartPhase progress, onState := StartPhases(func(p StartPhase) { got = append(got, p) }) - onState(remote.StateInFlight) + onState(cloud.StateInFlight) progress("connection dropped (unexpected EOF); retrying in 5s") last := got[len(got)-1] diff --git a/internal/fleet/wake.go b/internal/fleet/wake.go index 565a0dfd..46f41203 100644 --- a/internal/fleet/wake.go +++ b/internal/fleet/wake.go @@ -33,7 +33,7 @@ var wakePoll = 2 * time.Second // fleet file into one actual start. Two requests racing to wake the same // node is the ordinary shape of two agents starting near enough together, // and a daemon node's own 409 already turns the loser into a joiner — but a -// remote environment's control plane has no equivalent guard: its instance +// cloud environment's control plane has no equivalent guard: its instance // lookup is eventually consistent right after a launch, so two wakes that // race within that window can each miss the other's not-yet-visible // instance and each launch one, doubling the bill for what should have been @@ -269,7 +269,7 @@ func (c *Config) WaitLoading(ctx context.Context, w Want, results []NodeResult, // // A daemon that reports its own readiness reading — the engine has answered // its health check — is taken on that word; it checked from the same machine -// the engine runs on, and a remote node's reading is the control plane's own +// the engine runs on, and a cloud node's reading is the control plane's own // equivalent check. A ReadyNo reading is taken on its word too: the engine's // port can accept a connection well before the engine can answer a request // — llama.cpp and vLLM both open it early and answer their own health check diff --git a/internal/fleet/wake_test.go b/internal/fleet/wake_test.go index c2fd3cb3..ea980f71 100644 --- a/internal/fleet/wake_test.go +++ b/internal/fleet/wake_test.go @@ -342,9 +342,9 @@ func TestWakeTimesOutWithoutStopping(t *testing.T) { // Two Wake calls racing to wake the same node coalesce into one actual // start. This fixture's own /v1/start does not itself reject a concurrent // call the way a real daemon's supervisor mutex does — unlike a daemon -// node, a remote environment's control plane has no such guard at all — so +// node, a cloud environment's control plane has no such guard at all — so // without wakeSingleflight both calls would reach StartWith and each start -// their own engine (or, for a remote node, each launch their own instance). +// their own engine (or, for a cloud node, each launch their own instance). func TestWakeCoalescesConcurrentCallsForTheSameNode(t *testing.T) { shortWake(t) node := newFakeNode(t, string(daemon.StateIdle), "") @@ -501,11 +501,11 @@ func TestWakeWithoutAKeyIsUngated(t *testing.T) { } // The wake path stays daemon-only: a remote is never woken — what it serves is -// set by `spinloop remote deploy`. Wake boots its instance and waits for its +// set by `spinloop cloud deploy`. Wake boots its instance and waits for its // engine to answer the same way it does for a daemon node, without pushing // the candidate resolver's config onto it — the environment already knows // what it serves. -func TestWakeStartsADeployedRemoteNode(t *testing.T) { +func TestWakeStartsADeployedCloudNode(t *testing.T) { shortWake(t) stubAWSCreds(t) @@ -543,7 +543,7 @@ func TestWakeStartsADeployedRemoteNode(t *testing.T) { engine = ln mu.Unlock() w.Header().Set("Content-Type", "application/json") - // remote.Start only accepts HTTP 200 with state "ready" as done; the + // cloud.Start only accepts HTTP 200 with state "ready" as done; the // engine's own running state comes from the status polls waitReady // makes afterwards, not from this reply. fmt.Fprintf(w, `{"state":"ready","healthy":true,"runner":"llamacpp","modelId":"org/m","servedName":"m","base_url":"http://%s/v1"}`, ln.Addr()) @@ -557,9 +557,9 @@ func TestWakeStartsADeployedRemoteNode(t *testing.T) { engine.Close() } }) - registerRemoteEnv(t, "cloud", srv.URL, srv.URL) + registerCloudEnv(t, "cloud", srv.URL, srv.URL) - path := writeFleet(t, "nodes:\n - name: cloud\n kind: remote\n", "") + path := writeFleet(t, "nodes:\n - name: cloud\n kind: cloud\n", "") cfg, err := Load(path) if err != nil { t.Fatal(err) @@ -584,7 +584,7 @@ func TestWakeStartsADeployedRemoteNode(t *testing.T) { } } -// A remote node's wake waits for the control plane's own health check to +// A cloud node's wake waits for the control plane's own health check to // say ready, not just for its port to accept a connection: this fake // reports running-but-unhealthy for a stretch after boot, the way an engine // that has opened its port but is still loading weights does, before @@ -638,9 +638,9 @@ func TestWakeWaitsForARemoteEngineToBecomeHealthy(t *testing.T) { engine.Close() } }) - registerRemoteEnv(t, "cloud", srv.URL, srv.URL) + registerCloudEnv(t, "cloud", srv.URL, srv.URL) - path := writeFleet(t, "nodes:\n - name: cloud\n kind: remote\n", "") + path := writeFleet(t, "nodes:\n - name: cloud\n kind: cloud\n", "") cfg, err := Load(path) if err != nil { t.Fatal(err) diff --git a/internal/gateway/gateway.go b/internal/gateway/gateway.go index fcb0bd3f..6403af5d 100644 --- a/internal/gateway/gateway.go +++ b/internal/gateway/gateway.go @@ -300,9 +300,9 @@ func (h *Handler) handleModels(w http.ResponseWriter, r *http.Request) { // waking is allowed (its own `wake` setting, or the fleet's when it names // none), the model it would be started with, under the served-name-first // naming a running node reports. A daemon node's model comes from its own -// Spinloop source; a remote node's comes from its own stats reply. It is +// Spinloop source; a cloud node's comes from its own stats reply. It is // resolved at most once per sourcesTTL, shared by every models request — -// which bounds how often a remote node's resolution pays a live control- +// which bounds how often a cloud node's resolution pays a live control- // plane call, the same way it bounds how often a daemon node's pays a // Spinloop file read. func (h *Handler) wakeableModels(ctx context.Context) map[string]string { @@ -334,7 +334,7 @@ func (h *Handler) wakeableModels(ctx context.Context) map[string]string { return m } -// remoteConfigFor resolves what a kind: remote node would be started with: +// cloudConfigFor resolves what a kind: cloud node would be started with: // the environment's stored deploy config, read from its stats reply. The // status reply a fan-out already holds is no good for this — the control // plane only relays deploy facts on the status reply while the environment @@ -346,9 +346,9 @@ func (h *Handler) wakeableModels(ctx context.Context) map[string]string { // environment's stats read fails outright (no config to read), which is // what "nothing deployed" looks like here. It carries the served name // alongside the model id too, the same field the deploy config's ALIAS -// sets, so a stopped remote node's wakeable name matches what it reported +// sets, so a stopped cloud node's wakeable name matches what it reported // while running rather than falling back to the bare model id. -func (h *Handler) remoteConfigFor(ctx context.Context) fleet.ConfigFor { +func (h *Handler) cloudConfigFor(ctx context.Context) fleet.ConfigFor { return func(entry fleet.NodeConfig) (inference.DeployConfig, error) { node, err := h.cfg.NewNode(entry) if err != nil { @@ -356,7 +356,7 @@ func (h *Handler) remoteConfigFor(ctx context.Context) fleet.ConfigFor { } stats, err := node.Metrics(ctx) if err != nil { - return inference.DeployConfig{}, fmt.Errorf("%s: %w (run `spinloop remote deploy` if nothing is deployed)", entry.Name, err) + return inference.DeployConfig{}, fmt.Errorf("%s: %w (run `spinloop cloud deploy` if nothing is deployed)", entry.Name, err) } return inference.DeployConfig{ModelID: stats.ModelID, ServedModelName: stats.ServedName}, nil } @@ -364,15 +364,15 @@ func (h *Handler) remoteConfigFor(ctx context.Context) fleet.ConfigFor { // combinedConfigFor resolves what any node — daemon or remote — would be // started with: a daemon node through the gateway's own cfgFor (its -// Spinloop source), a remote node through its own stats reply -// (remoteConfigFor). A daemon node fails the way it always has when the -// gateway holds no cfgFor at all; a remote node's resolution does not +// Spinloop source), a cloud node through its own stats reply +// (cloudConfigFor). A daemon node fails the way it always has when the +// gateway holds no cfgFor at all; a cloud node's resolution does not // depend on cfgFor, so it still works when the gateway was built with none. func (h *Handler) combinedConfigFor(ctx context.Context) fleet.ConfigFor { - remoteFor := h.remoteConfigFor(ctx) + cloudFor := h.cloudConfigFor(ctx) return func(entry fleet.NodeConfig) (inference.DeployConfig, error) { - if entry.Kind == fleet.KindRemote { - return remoteFor(entry) + if entry.Kind == fleet.KindCloud { + return cloudFor(entry) } if h.cfgFor == nil { return inference.DeployConfig{}, fmt.Errorf( diff --git a/internal/gateway/gateway_test.go b/internal/gateway/gateway_test.go index b3208908..b152ce83 100644 --- a/internal/gateway/gateway_test.go +++ b/internal/gateway/gateway_test.go @@ -24,13 +24,13 @@ import ( "github.com/spinloop-ai/spinloop/internal/inference" ) -// registerRemoteEnv points the environment registry (SPINLOOP_CONFIG_DIR) at -// a temp config directory and writes one environment's remote.json, whose +// registerCloudEnv points the environment registry (SPINLOOP_CONFIG_DIR) at +// a temp config directory and writes one environment's cloud.json, whose // control plane — start, stop and stats alike — is the server url given, and // stubs the AWS credential chain so a signed control call reaches it. // Reproduced from internal/fleet's own helper of the same name because it // lives in a different package. -func registerRemoteEnv(t *testing.T, name, url string) { +func registerCloudEnv(t *testing.T, name, url string) { t.Helper() t.Setenv("AWS_ACCESS_KEY_ID", "AKIATESTTESTTESTTEST") t.Setenv("AWS_SECRET_ACCESS_KEY", "test-secret") @@ -42,13 +42,13 @@ func registerRemoteEnv(t *testing.T, name, url string) { home := t.TempDir() t.Setenv("SPINLOOP_CONFIG_DIR", home) - envDir := filepath.Join(home, "remotes", name) + envDir := filepath.Join(home, "clouds", name) if err := os.MkdirAll(envDir, 0o755); err != nil { t.Fatal(err) } body := fmt.Sprintf(`{"start_url":%q,"stop_url":%q,"stats_url":%q,"region":"us-east-1","environment":%q}`, url, url, url+"/stats", name) - if err := os.WriteFile(filepath.Join(envDir, "remote.json"), []byte(body), 0o600); err != nil { + if err := os.WriteFile(filepath.Join(envDir, "cloud.json"), []byte(body), 0o600); err != nil { t.Fatal(err) } } @@ -472,7 +472,7 @@ func TestModelsListsOnlyWhatRunsWhenWakeIsOff(t *testing.T) { } } -// A deployed-but-stopped remote environment's model comes from its own +// A deployed-but-stopped cloud environment's model comes from its own // stats reply — the environment's stored deploy config, read directly by // the stats Lambda — not from its status reply, which carries no deploy // facts while the environment is stopped, and not from a Spinloop source. @@ -480,11 +480,11 @@ func TestModelsListsOnlyWhatRunsWhenWakeIsOff(t *testing.T) { // served-name-first naming: the stats reply carries the served name beside // the model id, so a stopped environment lists the same name it would report // while running rather than the bare model id. -func TestModelsListsADeployedRemoteEnvironment(t *testing.T) { - url, _ := remoteControlServer(t) - registerRemoteEnv(t, "env", url) +func TestModelsListsADeployedCloudEnvironment(t *testing.T) { + url, _ := cloudControlServer(t) + registerCloudEnv(t, "env", url) cfg := &fleet.Config{Path: "fleet.yaml", Dir: t.TempDir(), Nodes: []fleet.NodeConfig{ - {Name: "env", Kind: fleet.KindRemote}, + {Name: "env", Kind: fleet.KindCloud}, }} h := New(cfg, "", Options{}) m := h.wakeableModels(context.Background()) @@ -493,10 +493,10 @@ func TestModelsListsADeployedRemoteEnvironment(t *testing.T) { } } -// An undeployed remote environment's stats read fails outright — the stats +// An undeployed cloud environment's stats read fails outright — the stats // Lambda has no deploy config to read — so it has nothing to be woken with // and is not listed. -func TestModelsLeavesOutAnUndeployedRemoteEnvironment(t *testing.T) { +func TestModelsLeavesOutAnUndeployedCloudEnvironment(t *testing.T) { mux := http.NewServeMux() mux.HandleFunc("GET /", func(w http.ResponseWriter, r *http.Request) { w.Header().Set("Content-Type", "application/json") @@ -504,28 +504,28 @@ func TestModelsLeavesOutAnUndeployedRemoteEnvironment(t *testing.T) { }) mux.HandleFunc("GET /stats", func(w http.ResponseWriter, r *http.Request) { w.WriteHeader(http.StatusBadRequest) - w.Write([]byte(`{"error":"cannot read deploy config: run spinloop remote deploy first"}`)) + w.Write([]byte(`{"error":"cannot read deploy config: run spinloop cloud deploy first"}`)) }) srv := httptest.NewServer(mux) t.Cleanup(srv.Close) - registerRemoteEnv(t, "env", srv.URL) + registerCloudEnv(t, "env", srv.URL) cfg := &fleet.Config{Path: "fleet.yaml", Dir: t.TempDir(), Nodes: []fleet.NodeConfig{ - {Name: "env", Kind: fleet.KindRemote}, + {Name: "env", Kind: fleet.KindCloud}, }} h := New(cfg, "", Options{}) if m := h.wakeableModels(context.Background()); len(m) != 0 { - t.Errorf("an undeployed remote environment has nothing to start it with, got %v", m) + t.Errorf("an undeployed cloud environment has nothing to start it with, got %v", m) } } -// A remote node with its own wake disabled is not listed, even though it is +// A cloud node with its own wake disabled is not listed, even though it is // deployed, and even under a fleet that wakes. func TestModelsLeavesOutARemoteEnvironmentWithWakeDisabled(t *testing.T) { - url, _ := remoteControlServer(t) - registerRemoteEnv(t, "env", url) + url, _ := cloudControlServer(t) + registerCloudEnv(t, "env", url) cfg := &fleet.Config{Path: "fleet.yaml", Dir: t.TempDir(), Nodes: []fleet.NodeConfig{ - {Name: "env", Kind: fleet.KindRemote, WakePolicy: fleet.WakeOff}, + {Name: "env", Kind: fleet.KindCloud, WakePolicy: fleet.WakeOff}, }} h := New(cfg, "", Options{}) if m := h.wakeableModels(context.Background()); len(m) != 0 { @@ -533,17 +533,17 @@ func TestModelsLeavesOutARemoteEnvironmentWithWakeDisabled(t *testing.T) { } } -// A remote node whose environment is not registered fails before any network +// A cloud node whose environment is not registered fails before any network // call, naming the node the way a daemon node with no resolvable Spinloop // source would. -func TestRemoteConfigForUnregisteredEnvironment(t *testing.T) { +func TestCloudConfigForUnregisteredEnvironment(t *testing.T) { t.Setenv("SPINLOOP_CONFIG_DIR", t.TempDir()) cfg := &fleet.Config{Path: "fleet.yaml", Dir: t.TempDir(), Nodes: []fleet.NodeConfig{ - {Name: "env", Kind: fleet.KindRemote}, + {Name: "env", Kind: fleet.KindCloud}, }} h := New(cfg, "", Options{}) - cfgFor := h.remoteConfigFor(context.Background()) - if _, err := cfgFor(fleet.NodeConfig{Name: "env", Kind: fleet.KindRemote}); err == nil || + cfgFor := h.cloudConfigFor(context.Background()) + if _, err := cfgFor(fleet.NodeConfig{Name: "env", Kind: fleet.KindCloud}); err == nil || !strings.Contains(err.Error(), "env") { t.Errorf("an unregistered environment should fail naming it, got %v", err) } @@ -928,14 +928,14 @@ func TestLoopbackBoundEngineFailsNamingTheFix(t *testing.T) { } } -// A remote environment's engine address arrives as its status's reported host — +// A cloud environment's engine address arrives as its status's reported host — // the control plane's published endpoint — so it is a candidate, not marked // unreachable the way a loopback-bound engine is. Without this the gateway -// could list a remote node's model and then refuse to route a request to it. +// could list a cloud node's model and then refuse to route a request to it. func TestReachableRoutesToARemoteNode(t *testing.T) { t.Setenv("TEST_ENGINE_KEY", "secret") cfg := &fleet.Config{Path: "fleet.yaml", Dir: t.TempDir(), APIKeyEnv: "TEST_ENGINE_KEY", Nodes: []fleet.NodeConfig{ - {Name: "env", Kind: fleet.KindRemote}, + {Name: "env", Kind: fleet.KindCloud}, }} h := New(cfg, "", Options{}) results := []fleet.NodeResult{ @@ -948,7 +948,7 @@ func TestReachableRoutesToARemoteNode(t *testing.T) { } choice, err := cfg.Choose(h.reachable(results), fleet.Want{Model: "wanted"}) if err != nil { - t.Fatalf("a request for the remote node's model should route to it: %v", err) + t.Fatalf("a request for the cloud node's model should route to it: %v", err) } if choice.Node.Name != "env" { t.Errorf("choice = %q, want env", choice.Node.Name) @@ -1163,9 +1163,9 @@ func TestNothingCanServeNamesEveryRefusal(t *testing.T) { } } -// --- waking a remote node ----------------------------------------------------- +// --- waking a cloud node ----------------------------------------------------- -// remoteControlServer serves a deployed-but-stopped environment's status +// cloudControlServer serves a deployed-but-stopped environment's status // until its instance is booted (POST), after which it reports running with // an engine that actually answers an OpenAI-compatible completion — so a // request proxied to it end to end gets a real reply, the same way it does @@ -1180,7 +1180,7 @@ func TestNothingCanServeNamesEveryRefusal(t *testing.T) { // Lambda reads the deploy config directly rather than relaying it alongside // instance state; that is what the gateway now resolves a stopped remote // node's wakeable model from. -func remoteControlServer(t *testing.T) (url string, started func() bool) { +func cloudControlServer(t *testing.T) (url string, started func() bool) { t.Helper() var ( mu sync.Mutex @@ -1236,22 +1236,22 @@ func remoteControlServer(t *testing.T) (url string, started func() bool) { return srv.URL, func() bool { mu.Lock(); defer mu.Unlock(); return wasSent } } -// A request for a model only a deployed-but-stopped remote node serves is +// A request for a model only a deployed-but-stopped cloud node serves is // held and answered once that node's instance boots — the gateway sources // its wakeable model from its own last status, not a Spinloop source, and // StartWith boots it without pushing any config. -func TestColdRequestWakesADeployedRemoteNode(t *testing.T) { +func TestColdRequestWakesADeployedCloudNode(t *testing.T) { shortWake := func(t *testing.T) { old := fleet.WakeTimeout fleet.WakeTimeout = 3 * time.Second t.Cleanup(func() { fleet.WakeTimeout = old }) } shortWake(t) - url, started := remoteControlServer(t) - registerRemoteEnv(t, "cloud", url) + url, started := cloudControlServer(t) + registerCloudEnv(t, "cloud", url) cfg := &fleet.Config{Path: "fleet.yaml", Dir: t.TempDir(), Nodes: []fleet.NodeConfig{ - {Name: "cloud", Kind: fleet.KindRemote}, + {Name: "cloud", Kind: fleet.KindCloud}, }} h := New(cfg, "", Options{}) @@ -1264,23 +1264,23 @@ func TestColdRequestWakesADeployedRemoteNode(t *testing.T) { } } -// A request for a deployed-but-stopped remote node's served name — the +// A request for a deployed-but-stopped cloud node's served name — the // alias a caller knew it by while it was running, rather than its bare model // id — still matches: the stats reply the wake path reads carries the served // name from the deploy config, the same field the status reply carries while // running, so a caller need not learn a second name once the node stops. -func TestColdRequestWakesADeployedRemoteNodeByItsServedName(t *testing.T) { +func TestColdRequestWakesADeployedCloudNodeByItsServedName(t *testing.T) { shortWake := func(t *testing.T) { old := fleet.WakeTimeout fleet.WakeTimeout = 3 * time.Second t.Cleanup(func() { fleet.WakeTimeout = old }) } shortWake(t) - url, started := remoteControlServer(t) - registerRemoteEnv(t, "cloud", url) + url, started := cloudControlServer(t) + registerCloudEnv(t, "cloud", url) cfg := &fleet.Config{Path: "fleet.yaml", Dir: t.TempDir(), Nodes: []fleet.NodeConfig{ - {Name: "cloud", Kind: fleet.KindRemote}, + {Name: "cloud", Kind: fleet.KindCloud}, }} h := New(cfg, "", Options{}) @@ -1293,10 +1293,10 @@ func TestColdRequestWakesADeployedRemoteNodeByItsServedName(t *testing.T) { } } -// An undeployed remote node's stats read fails outright — nothing deployed +// An undeployed cloud node's stats read fails outright — nothing deployed // to read — so it does not match any request and the failure says so, // naming the deploy path, without holding the request for the wake timeout. -func TestWakeRefusesAnUndeployedRemoteNode(t *testing.T) { +func TestWakeRefusesAnUndeployedCloudNode(t *testing.T) { mux := http.NewServeMux() mux.HandleFunc("GET /", func(w http.ResponseWriter, r *http.Request) { w.Header().Set("Content-Type", "application/json") @@ -1304,14 +1304,14 @@ func TestWakeRefusesAnUndeployedRemoteNode(t *testing.T) { }) mux.HandleFunc("GET /stats", func(w http.ResponseWriter, r *http.Request) { w.WriteHeader(http.StatusBadRequest) - w.Write([]byte(`{"error":"cannot read deploy config: run spinloop remote deploy first"}`)) + w.Write([]byte(`{"error":"cannot read deploy config: run spinloop cloud deploy first"}`)) }) srv := httptest.NewServer(mux) t.Cleanup(srv.Close) - registerRemoteEnv(t, "cloud", srv.URL) + registerCloudEnv(t, "cloud", srv.URL) cfg := &fleet.Config{Path: "fleet.yaml", Dir: t.TempDir(), Nodes: []fleet.NodeConfig{ - {Name: "cloud", Kind: fleet.KindRemote}, + {Name: "cloud", Kind: fleet.KindCloud}, }} h := New(cfg, "", Options{}) @@ -1319,20 +1319,20 @@ func TestWakeRefusesAnUndeployedRemoteNode(t *testing.T) { if resp.StatusCode != http.StatusServiceUnavailable { t.Fatalf("HTTP %d, want 503: %s", resp.StatusCode, body) } - if !strings.Contains(body, "spinloop remote deploy") { + if !strings.Contains(body, "spinloop cloud deploy") { t.Errorf("the failure should name the deploy path: %s", body) } } -// A remote node whose own wake is disabled is not started even though its +// A cloud node whose own wake is disabled is not started even though its // last status matches the request, and the failure says waking is off for // it rather than that nothing can serve the model at all. -func TestWakeDisabledForAMatchingRemoteNodeNamesIt(t *testing.T) { - url, started := remoteControlServer(t) - registerRemoteEnv(t, "cloud", url) +func TestWakeDisabledForAMatchingCloudNodeNamesIt(t *testing.T) { + url, started := cloudControlServer(t) + registerCloudEnv(t, "cloud", url) cfg := &fleet.Config{Path: "fleet.yaml", Dir: t.TempDir(), Nodes: []fleet.NodeConfig{ - {Name: "cloud", Kind: fleet.KindRemote, WakePolicy: fleet.WakeOff}, + {Name: "cloud", Kind: fleet.KindCloud, WakePolicy: fleet.WakeOff}, }} h := New(cfg, "", Options{}) diff --git a/internal/harness/adapters.go b/internal/harness/adapters.go index cc8c1176..a2517577 100644 --- a/internal/harness/adapters.go +++ b/internal/harness/adapters.go @@ -44,9 +44,9 @@ func opencodeBlock(p *catalog.Provider, sel spinloop.Selection, contextWindow, o if !setDefaultModel { defaultModel = "" } - // A remote selection renames the provider after its environment and carries a + // A cloud selection renames the provider after its environment and carries a // display name to match; opencode's model picker lists providers by that - // name, so use it in place of the catalogue engine's name to tell the remote + // name, so use it in place of the catalogue engine's name to tell the cloud // provider apart from a local engine of the same kind. if sel.DisplayName != "" { block["name"] = sel.DisplayName @@ -281,7 +281,7 @@ func (lucinateHarness) Apply(p *catalog.Provider, sel spinloop.Selection, contex } // The connection's display name is the provider's, or the selection's display - // name for a remote endpoint, so it reads distinctly from a local engine of + // name for a cloud endpoint, so it reads distinctly from a local engine of // the same kind — mirroring the opencode adapter. name := p.Name if name == "" { diff --git a/internal/harness/harness_test.go b/internal/harness/harness_test.go index 4a711b65..72862e8a 100644 --- a/internal/harness/harness_test.go +++ b/internal/harness/harness_test.go @@ -158,7 +158,7 @@ func TestNamesAndLookup(t *testing.T) { } } -// A config written with no key for a remote endpoint succeeds and then fails on +// A config written with no key for a cloud endpoint succeeds and then fails on // the first request, so it has to be called out at the time. func TestMissingKeyWarning(t *testing.T) { keyed := &catalog.Provider{APIKeyEnv: "OPENAI_API_KEY"} @@ -166,7 +166,7 @@ func TestMissingKeyWarning(t *testing.T) { set := func(string) string { return "sk-test" } if w := missingKeyWarning(keyed, "http://198.51.100.1:8000/v1", unset); w == "" { - t.Error("a remote endpoint with no key should warn") + t.Error("a cloud endpoint with no key should warn") } else if !strings.Contains(w, "OPENAI_API_KEY") { t.Errorf("the warning should name the variable to set, got %q", w) } diff --git a/internal/inference/inference.go b/internal/inference/inference.go index 3acae97f..a03c0f0b 100644 --- a/internal/inference/inference.go +++ b/internal/inference/inference.go @@ -5,7 +5,7 @@ // It is deliberately a leaf — standard library only — so a node kind can be // described without depending on how any other kind is reached. What is // specific to reaching one kind stays with that kind: the AWS control plane in -// internal/remote, the control API in internal/daemon. +// internal/cloud, the control API in internal/daemon. package inference // DeployConfig is what the deploy Lambda accepts: the runner-neutral diff --git a/internal/metrics/metrics.go b/internal/metrics/metrics.go index 7ae44633..81c1adab 100644 --- a/internal/metrics/metrics.go +++ b/internal/metrics/metrics.go @@ -2,7 +2,7 @@ // host's GPU/CPU/RAM figures, in process. It is the Go home of the collection // that first shipped in the remote stats Lambda (remote/lambda/shared/stats.ts): // the parsers here are ports of that Lambda's, kept value-for-value compatible -// so the remote path can later delegate to a daemon running this code and the +// so the cloud path can later delegate to a daemon running this code and the // TypeScript collectors can be deleted. Every stat is optional — a host // without a source for one (no nvidia-smi, say) simply omits it. A macOS node // reports GPU utilisation and name from the I/O Kit accelerator service, and @@ -12,7 +12,7 @@ package metrics // Stats is the collected state of one serving host: what is running, its // engine counters, and its system figures. It mirrors the stats Lambda's // response field-for-field (minus the Lambda's transport fields), so the -// existing `spinloop remote metrics` formats render it unchanged. +// existing `spinloop cloud metrics` formats render it unchanged. type Stats struct { State string `json:"state"` Runner string `json:"runner,omitempty"` @@ -61,7 +61,7 @@ type Stats struct { // RetainUntil is the environment's retention deadline, RFC 3339: the idle // sweep will not terminate the instance before it. It is a property of the // cloud instance, not the engine, so it is empty for local daemon nodes and - // for remote environments without an update URL. Empty once the deadline + // for cloud environments without an update URL. Empty once the deadline // has passed, because the stats reply drops it there — a past tag keeps // nothing. Formatters omit the line when it is empty. RetainUntil string `json:"retainUntil,omitempty"` diff --git a/internal/opencode/opencode.go b/internal/opencode/opencode.go index 7d32f19f..7ce75d27 100644 --- a/internal/opencode/opencode.go +++ b/internal/opencode/opencode.go @@ -21,7 +21,7 @@ import ( // project's own `.env`. // // The process environment wins so an exported variable always beats the `.env`, -// which only fills a gap — the same precedence the remote commands follow, so +// which only fills a gap — the same precedence the cloud commands follow, so // the whole tool resolves local variables the same way. // // The file sits beside the Spinloop rather than beside the binary because that is diff --git a/internal/preset/preset.go b/internal/preset/preset.go index 49aeea21..ea8f0c19 100644 --- a/internal/preset/preset.go +++ b/internal/preset/preset.go @@ -266,7 +266,7 @@ var canonical = map[string]string{ "mu": "model-url", "tb": "threads-batch", "to": "timeout", "kvu": "kv-unified", // Speculative decoding. The draft-model spellings matter beyond rendering: - // `spinloop remote deploy` drops the flags the cloud sets itself by canonical + // `spinloop cloud deploy` drops the flags the cloud sets itself by canonical // name, and a drafter path written as `md` would otherwise slip past that // check and reach the instance, where the local path does not exist. "md": "spec-draft-model", "model-draft": "spec-draft-model", diff --git a/internal/preset/preset_test.go b/internal/preset/preset_test.go index 8e3b6bd0..c51b4446 100644 --- a/internal/preset/preset_test.go +++ b/internal/preset/preset_test.go @@ -251,7 +251,7 @@ func TestDialectOMLXHasNoBareBooleans(t *testing.T) { } // TestPackageFlagsStillLlamaCpp guards the callers that predate dialects: the -// package-level helpers, and remote.go's CanonicalKey, must keep rendering +// package-level helpers, and cloud.go's CanonicalKey, must keep rendering // llama.cpp. func TestPackageFlagsStillLlamaCpp(t *testing.T) { got := strings.Join(Flags([]Param{{Key: "hf", Value: "org/m"}, {Key: "mmap", Value: "1"}}), " ") diff --git a/internal/spinloop/spinloop.go b/internal/spinloop/spinloop.go index 4c641a34..a0fcc2d5 100644 --- a/internal/spinloop/spinloop.go +++ b/internal/spinloop/spinloop.go @@ -23,7 +23,7 @@ // // ENV sets an environment variable for the local `spinloop` process — the one // keyword that may appear more than once. It carries a single KEY=VALUE token -// and is used by the remote commands (which read it before signing AWS calls); +// and is used by the cloud commands (which read it before signing AWS calls); // it is local-only and never reaches a deployed instance. // // Keywords are matched case-insensitively, but UPPERCASE is canonical (it is @@ -129,7 +129,7 @@ func Parse(data []byte) (Selection, error) { if canon == "" { if kw == "remote" { return Selection{}, fmt.Errorf( - "line %d: the REMOTE instruction was removed: name the environment with `spinloop remote deploy --env ` at deploy time, and pass --env to the commands that act on it (remote subcommands, apply, unapply, harness)", + "line %d: the REMOTE instruction was removed: name the environment with `spinloop cloud deploy --env ` at deploy time, and pass --env to the commands that act on it (cloud subcommands, apply, unapply, harness)", line) } return Selection{}, fmt.Errorf("line %d: unknown keyword %q (expected PROVIDER, MODEL, ALIAS, CONTEXT, OUTPUT, PARALLEL, BASEURL, PRESET, or ENV)", line, fields[0]) diff --git a/internal/spinloopsrc/spinloopsrc.go b/internal/spinloopsrc/spinloopsrc.go index 25f5041e..65faba49 100644 --- a/internal/spinloopsrc/spinloopsrc.go +++ b/internal/spinloopsrc/spinloopsrc.go @@ -66,7 +66,7 @@ func Resolve(base, ref string) (string, error) { const fetchTimeout = 15 * time.Second // maxFetchSize caps how much of a response body Fetch reads. Every reference -// this package fetches — a Spinloop, a preset .ini, a remote.json — is a small, +// this package fetches — a Spinloop, a preset .ini, a cloud.json — is a small, // hand-editable text file, so this is generous headroom, not a tight limit. const maxFetchSize = 1 << 20 // 1 MiB diff --git a/mkdocs.yml b/mkdocs.yml index 5b393968..5ff0dd85 100644 --- a/mkdocs.yml +++ b/mkdocs.yml @@ -62,7 +62,7 @@ nav: - Open your coding agent: guides/harness.md - From a Hugging Face model: guides/hugging-face.md - Run a daemon node: guides/daemon.md - - Deploy to a cloud GPU: guides/remote.md + - Deploy to a cloud GPU: guides/cloud.md - Run a fleet: guides/fleet.md - Serve the fleet as a gateway: guides/gateway.md - Work a backlog: guides/work-items.md @@ -81,7 +81,7 @@ nav: - spinloop hf: commands/hf.md - spinloop alias: commands/alias.md - spinloop unalias: commands/unalias.md - - spinloop remote: commands/remote.md + - spinloop cloud: commands/cloud.md - spinloop fleet: commands/fleet.md - spinloop gateway: commands/gateway.md - spinloop orchestrator: commands/orchestrator.md diff --git a/openspec/changes/rename-remote-to-cloud/.openspec.yaml b/openspec/changes/rename-remote-to-cloud/.openspec.yaml new file mode 100644 index 00000000..e3966d7a --- /dev/null +++ b/openspec/changes/rename-remote-to-cloud/.openspec.yaml @@ -0,0 +1,2 @@ +schema: spec-driven +created: 2026-10-05 diff --git a/openspec/changes/rename-remote-to-cloud/design.md b/openspec/changes/rename-remote-to-cloud/design.md new file mode 100644 index 00000000..07b381cb --- /dev/null +++ b/openspec/changes/rename-remote-to-cloud/design.md @@ -0,0 +1,59 @@ +## Context + +"remote" currently means four things: this cloud GPU feature, a Spinloop +fetched over HTTP, a daemon on another host, and the `remote/` CDK project. +Only the first is being renamed. About 120 Go files and 35 specs mention +"remote"; most of the work is mechanical. + +## Goals / Non-Goals + +**Goals** +- `cloud` is the only name for this feature in the CLI, help, completion, + docs, errors, environment variables, fleet files and on-disk files. + +**Non-Goals** +- Changing any behaviour of the commands. +- Keeping the old names working (no alias, no fallback, no migration). +- Renaming the `remote/` CDK directory, `remote-spinloop-sources`, or the + capability directories under `openspec/specs/`. +- Removing the leftover path-form `REMOTE` references (issue 259). + +## Decisions + +1. **Clean break.** The old command, fleet kind, variables and paths are + removed outright; there is no deprecation period. The rewrite table in + `commands.go` (`"remote status": "status --env "`) is re-keyed to + `cloud`, and the `remote` command is not registered. +2. **Package rename with `git mv`.** `internal/remote` → `internal/cloud` + and `internal/fleet/remote_node.go` → `cloud_node.go`; identifiers such as + `KindRemote` → `KindCloud` change in the same commit so the build never has + two names for one thing. +3. **Environment variables.** `cli_viper.go` binds `SPINLOOP_CLOUD_*`; callers + ask for `cloud_`. +4. **Storage path.** `remotes//remote.json` → `clouds//cloud.json`; + keystore `remote-.json` → `cloud-.json`. Existing + environments must be re-registered with `spinloop cloud deploy`. +5. **Keep `remote/` (CDK).** It is referenced by the release asset the + bootstrap command downloads (`internal/cloud/source.go` strips the + `remote/` prefix), by CI, and by `scripts/check-no-cloud-identifiers.sh`. + Renaming it changes the release layout for no user-visible gain. Left as a + possible follow-up. +6. **Spec directories keep their names.** Renaming 17 capability + directories would break archive history and links. Main-spec wording is + updated; the new `cloud-command` capability holds the naming contract. + +## Risks / Trade-offs + +- A blind find-and-replace would hit the unrelated meanings above. + Mitigation: rename per package, review each diff, keep the "unchanged" + requirement as a test (`spinloop apply ` still works). +- Tab completion and the root dispatch table are string-keyed; a missed + entry fails at runtime. Mitigation: `complete_test.go` and + `root_dispatch_test.go` cases for the new name and for `remote` being unknown. +- Existing users lose access to registered environments and keystore files + until they re-register. Accepted; called out as **BREAKING** in the proposal. + +## Migration Plan + +There is no automatic migration. Users rename their scripts and fleet files, set +`SPINLOOP_CLOUD_*`, and re-run `spinloop cloud deploy` for each environment. diff --git a/openspec/changes/rename-remote-to-cloud/proposal.md b/openspec/changes/rename-remote-to-cloud/proposal.md new file mode 100644 index 00000000..488bd4a0 --- /dev/null +++ b/openspec/changes/rename-remote-to-cloud/proposal.md @@ -0,0 +1,55 @@ +## Why + +The `spinloop remote` command group deploys and drives scale-to-zero GPU +instances in the cloud, but "remote" is also used for other things: a Spinloop +fetched over HTTP, a remote node in a fleet, a remote machine running a +daemon. Naming the group `cloud` says what it does and removes the clash +(GitHub issue 117). + +## What Changes + +- Rename the `spinloop remote` command group to `spinloop cloud`. The + subcommands (`bootstrap`, `auth`, `bake`, `deploy`, `start`, `pause`, + `restart`, `keep`, `stop`, `seed`, `list`, …) keep their names. +- Rename the fleet node kind `remote` to `cloud` (`kind: cloud` in a fleet file). +- Rename the environment variables `SPINLOOP_REMOTE_*` to `SPINLOOP_CLOUD_*`. +- Rename the on-disk names: the `remotes/` environments directory, each + environment's `remote.json`, and the `remote-.json` keystore files. +- Rename the Go package `internal/remote` to `internal/cloud`, and the fleet + node type, file names, test names, error messages and help text to match. +- Reword the docs, examples, specs and `AGENTS.md` to say "cloud" wherever + they mean this feature. +- **BREAKING**: the old names are removed with no fallback. `spinloop remote`, + `kind: remote`, `SPINLOOP_REMOTE_*` and the old on-disk names stop working; + environments registered under `remotes/` must be re-registered. +- Not renamed: the `remote/` CDK project directory (see design.md), and + a Spinloop *source* fetched over HTTP (`remote-spinloop-sources`, + `internal/spinloopsrc`), which is unrelated. Removing the leftover path-form + `REMOTE` references is tracked in issue 259. + +## Capabilities + +### New Capabilities + +- `cloud-command`: the `cloud` command group name, the `cloud` fleet node + kind, the `SPINLOOP_CLOUD_*` variables, the on-disk names, and the removal + of the old names. + +### Modified Capabilities + +None as delta specs. The existing `remote-*` capability specs describe the +same behaviour under the old name; their wording is updated in the main specs +as part of the change (see tasks.md), and their directory names are left alone +so that history and links stay valid. + +## Impact + +- Code: `cmd/spinloop` (command tree, completion, root dispatch, help text, + `remote*.go` files), `internal/remote` → `internal/cloud`, `internal/fleet` + (node kind, `remote_node.go`), `internal/harness`, `internal/config`, + `internal/gateway`, and the Viper env binding. +- Docs: `docs/commands/remote.md`, `docs/guides/remote.md`, fleet, env-vars, + FAQ, troubleshooting, `mkdocs.yml`, `README.md`, examples (`fleet-remote`). +- Existing users: scripts, fleet files and environment variables using the old + names must be updated, and environments re-registered. +- `CHANGELOG.md` is not touched (the release process does it). diff --git a/openspec/changes/rename-remote-to-cloud/specs/cloud-command/spec.md b/openspec/changes/rename-remote-to-cloud/specs/cloud-command/spec.md new file mode 100644 index 00000000..fddb9caa --- /dev/null +++ b/openspec/changes/rename-remote-to-cloud/specs/cloud-command/spec.md @@ -0,0 +1,92 @@ +## Purpose + +Fix the names the cloud GPU feature is addressed by: the `cloud` command group, +the `cloud` fleet node kind, the `SPINLOOP_CLOUD_*` environment variables and +the on-disk files. "remote" also means a Spinloop fetched over HTTP and a +daemon on another machine, so the old `remote` name was ambiguous. The old +names are removed rather than kept as aliases. + +## ADDED Requirements + +### Requirement: The cloud command group + +The CLI SHALL provide a `cloud` command group containing every subcommand the +`remote` group provided, with the same arguments, flags, output and exit +codes. Help text, error messages and progress output SHALL say "cloud" where +they previously said "remote" for this feature. + +#### Scenario: A subcommand runs under the new name +- **WHEN** the user runs `spinloop cloud deploy staging ./Spinloop` +- **THEN** it behaves exactly as `spinloop remote deploy staging ./Spinloop` did before this change + +#### Scenario: Help lists the new name +- **WHEN** the user runs `spinloop --help` +- **THEN** the command list shows `cloud` and does not show `remote` + +### Requirement: The remote command group is removed + +The CLI SHALL NOT provide a `remote` command group. Running `spinloop remote …` +SHALL fail as any other unknown command does, and shell completion SHALL NOT +offer `remote`. + +#### Scenario: Old command is unknown +- **WHEN** the user runs `spinloop remote status` +- **THEN** the CLI exits non-zero with an unknown-command error + +#### Scenario: Completion offers only the new name +- **WHEN** shell completion is requested for the root command's subcommands +- **THEN** `cloud` is offered and `remote` is not + +### Requirement: The cloud fleet node kind + +A fleet file node SHALL accept `kind: cloud` for a registered cloud +environment. `kind: remote` SHALL be rejected with the same error as any other +unknown kind. Fleet output, status and dashboard SHALL label such nodes `cloud`. + +#### Scenario: New kind parses +- **WHEN** a fleet file declares a node with `kind: cloud` +- **THEN** it is loaded as a cloud environment node + +#### Scenario: Old kind is rejected +- **WHEN** a fleet file declares a node with `kind: remote` +- **THEN** loading fails with an unknown-kind error naming the node + +### Requirement: Cloud environment variables + +The CLI SHALL read its cloud settings from `SPINLOOP_CLOUD_*` variables +(`SPINLOOP_CLOUD_REGION`, `_START_URL`, `_STOP_URL`, `_STATS_URL`, +`_DEPLOY_URL`, `_ENV_URL`, `_SEED_URL`, `_UPDATE_URL`, `_KEYSTORE`, +`_PACKAGE_MANAGER`). `SPINLOOP_REMOTE_*` variables SHALL be ignored. + +#### Scenario: New variable is read +- **WHEN** `SPINLOOP_CLOUD_REGION=us-east-1` is set +- **THEN** the region used is `us-east-1` + +#### Scenario: Old variable is ignored +- **WHEN** only `SPINLOOP_REMOTE_REGION=eu-west-2` is set +- **THEN** the region is treated as unset + +### Requirement: Cloud environment storage names + +Environments SHALL be stored under `~/.config/spinloop/clouds//cloud.json` +and credentials in `cloud-.json` keystore files, with the same file +modes as before. The CLI SHALL NOT read or write `remotes//remote.json` +or `remote-.json`. + +#### Scenario: Writes use the new path +- **WHEN** `spinloop cloud deploy staging` finishes +- **THEN** `clouds/staging/cloud.json` exists with mode 0600 + +#### Scenario: Old environment is not found +- **WHEN** `remotes/staging/remote.json` exists and `clouds/staging/` does not +- **THEN** `spinloop status --env staging` reports that environment `staging` is not registered + +### Requirement: Unrelated uses of remote are unchanged + +Names that mean "fetched over HTTP" or "on another machine" SHALL keep +`remote` in their name: the `remote-spinloop-sources` capability, and daemon +base URLs on other hosts. + +#### Scenario: A Spinloop is fetched from a URL +- **WHEN** the user runs `spinloop apply https://example.com/Spinloop` +- **THEN** it behaves as before this change diff --git a/openspec/changes/rename-remote-to-cloud/tasks.md b/openspec/changes/rename-remote-to-cloud/tasks.md new file mode 100644 index 00000000..4ceb14ef --- /dev/null +++ b/openspec/changes/rename-remote-to-cloud/tasks.md @@ -0,0 +1,22 @@ +## 1. Code rename + +- [x] 1.1 `git mv internal/remote internal/cloud`; update package name, imports and identifiers +- [x] 1.2 Rename `cmd/spinloop/remote*.go` and their tests to `cloud*.go`; rename the command, subcommand wiring, help text and error messages; do not register a `remote` command +- [x] 1.3 Rename `internal/fleet/remote_node.go` (and test files), `KindRemote` → `KindCloud`, node type names, the `kind` value, and labels in fleet status and dashboard +- [x] 1.4 Update `internal/harness`, `internal/config`, `internal/gateway`, `internal/catalog`, `internal/daemon` and `cmd/spinloop/{route,code,serve,status*,metrics*,logs,follow,dashboard*,palette,complete}.go` references that mean this feature (leave `spinloopsrc` and URL-source "remote" alone) +- [x] 1.5 Re-key the rewrite table and root dispatch in `commands.go` / `main.go` to `cloud` +- [x] 1.6 Rename `SPINLOOP_REMOTE_*` to `SPINLOOP_CLOUD_*` in `cli_viper.go`, code, `.github` workflows and example scripts +- [x] 1.7 Rename the on-disk names to `clouds//cloud.json` and `cloud-.json` + +## 2. Specs and docs + +- [x] 2.1 Update wording in the existing `remote-*` main specs and other specs that mention this feature (leave `remote-spinloop-sources`); update `openspec/config.yaml` context +- [x] 2.2 Rename `docs/commands/remote.md` → `cloud.md`, `docs/guides/remote.md` → `cloud.md`, `examples/fleet-remote` → `fleet-cloud`; update `mkdocs.yml`, README, FAQ, troubleshooting, env-vars, fleet file, openapi and links +- [x] 2.3 Update `AGENTS.md` layout section and `docs/maintainer/internals.md`; note the kept `remote/` directory +- [x] 2.4 Check `scripts/check-no-cloud-identifiers.sh`, CI workflows and `.dockerignore` still match + +## 3. Tests + +- [x] 3.1 Tests for each scenario in `cloud-command`: `cloud` runs, `remote` is an unknown command and not completed, `kind: cloud` loads and `kind: remote` is rejected, `SPINLOOP_CLOUD_*` is read and `SPINLOOP_REMOTE_*` ignored, new paths written and old paths not read +- [x] 3.2 Test that `spinloop apply ` is unchanged +- [x] 3.3 `go build ./... && go vet ./... && go test ./... -cover` (total ≥ 80%), `gofmt -l .` clean, `scripts/check-no-cloud-identifiers.sh`, and a `grep -ri remote` review of what is left diff --git a/openspec/config.yaml b/openspec/config.yaml index eddc2ba0..88ad66e1 100644 --- a/openspec/config.yaml +++ b/openspec/config.yaml @@ -6,16 +6,16 @@ schema: spec-driven context: | spinloop is a Go CLI that configures a coding agent — a "harness" (opencode, Pi or lucinate) — to use a model provider, by deep-merging provider settings - into that harness's config; it also serves, fleets and remotely deploys local + into that harness's config; it also serves, fleets and deploys local inference engines. Tech stack: Go; domain logic in internal/ packages (catalog, config, contextsize, daemon, fleet, harness, lucinate, opencode, spinloop, spinloopsrc, - pi, preset, remote), CLI in cmd/spinloop on Cobra (command tree + pflag flags) - with Viper as the CLI-layer binding surface for SPINLOOP_ALIAS, SPINLOOP_REMOTE_* - and SPINLOOP_REMOTE_PACKAGE_MANAGER; shell tab completion is Cobra's own + pi, preset, cloud), CLI in cmd/spinloop on Cobra (command tree + pflag flags) + with Viper as the CLI-layer binding surface for SPINLOOP_ALIAS, SPINLOOP_CLOUD_* + and SPINLOOP_CLOUD_PACKAGE_MANAGER; shell tab completion is Cobra's own __complete engine plus the scripts it generates. The only AWS/network - dependency is aws-sdk-go-v2 in internal/remote. Providers/models are data in + dependency is aws-sdk-go-v2 in internal/cloud. Providers/models are data in internal/catalog/providers.yaml, not Go code. Conventions: keep test coverage >= 80% (go test ./... -cover); gofmt before diff --git a/openspec/specs/alias-registry/spec.md b/openspec/specs/alias-registry/spec.md index 990a3bdf..9323d18c 100644 --- a/openspec/specs/alias-registry/spec.md +++ b/openspec/specs/alias-registry/spec.md @@ -4,7 +4,7 @@ Define the alias registry: naming a Spinloop once with `spinloop alias` so the name stands in for its path in every command that takes one (`apply`, -`unapply`, `serve`, `harness`, and the `remote` control commands `deploy`, +`unapply`, `serve`, `harness`, and the `cloud` control commands `deploy`, `start`, `stop`, `status`, `stats`), and the rules that keep aliases from ever changing what an already-working command does. @@ -75,7 +75,7 @@ fetched. When an alias decides the path, the command SHALL say so. That report SHALL go to stderr. It is prose about how the command was resolved rather than the command's result, and the same resolution serves -`spinloop remote env`, whose stdout is meant to be evaluated by a shell. +`spinloop cloud env`, whose stdout is meant to be evaluated by a shell. #### Scenario: Alias used from anywhere @@ -86,7 +86,7 @@ rather than the command's result, and the same resolution serves #### Scenario: The alias note stays out of stdout - **WHEN** an alias resolves the Spinloop for a command whose stdout is consumed - by a shell, such as `spinloop remote env` + by a shell, such as `spinloop cloud env` - **THEN** the note naming the alias is written to stderr and stdout carries only the command's own output @@ -131,14 +131,14 @@ When `SPINLOOP_ALIAS` decides the Spinloop, the command SHALL say so on stderr, naming the variable, the alias and the resolved path. A command that consults a Spinloop only when there is one to consult — the -`remote` subcommands, which otherwise act on the `default` environment, and +`cloud` subcommands, which otherwise act on the `default` environment, and `daemon`, which otherwise starts idle — SHALL count `SPINLOOP_ALIAS` as naming one. A set variable SHALL NOT be passed over in favour of that fallback. #### Scenario: The variable counts as having a Spinloop - **WHEN** `SPINLOOP_ALIAS` names a Spinloop whose `ENV` instructions set - `AWS_PROFILE` and the user runs `spinloop remote status` in a directory with + `AWS_PROFILE` and the user runs `spinloop cloud status` in a directory with no `Spinloop` - **THEN** that Spinloop's `ENV` instructions are applied to the process environment before the control call, rather than being skipped because no diff --git a/openspec/specs/config-location/spec.md b/openspec/specs/config-location/spec.md index 9ef4796a..671e3f3e 100644 --- a/openspec/specs/config-location/spec.md +++ b/openspec/specs/config-location/spec.md @@ -5,7 +5,7 @@ How spinloop resolves its own config directory, and why that resolution is explicit rather than inferred. -Everything spinloop owns lives under one directory — `config.json`, `remote.json`, +Everything spinloop owns lives under one directory — `config.json`, `cloud.json`, the environment registry, the daemon state dir, the CDK source cache — and the obvious way to find it leans on `$HOME`. A systemd service does not get one. On the cloud instance that meant the boot script wrote to `/root/.config/spinloop` @@ -20,14 +20,14 @@ silently resolving to a bogus relative path. spinloop SHALL resolve one config directory and place every file it owns under it: its own `config.json` (default-harness preference and alias registry), the -`remotes//` environment registry, the daemon state directory, and the CDK +`clouds//` environment registry, the daemon state directory, and the CDK source directory. There SHALL be one resolver; the location SHALL NOT be computed independently in more than one place. #### Scenario: All spinloop-owned state shares one root - **WHEN** the config directory resolves to a given path -- **THEN** `config.json`, the `remotes//` registry, and the daemon state +- **THEN** `config.json`, the `clouds//` registry, and the daemon state directory all resolve beneath that same path ### Requirement: SPINLOOP_CONFIG_DIR override diff --git a/openspec/specs/daemon-api/spec.md b/openspec/specs/daemon-api/spec.md index f1c12450..78f14819 100644 --- a/openspec/specs/daemon-api/spec.md +++ b/openspec/specs/daemon-api/spec.md @@ -173,7 +173,7 @@ Status SHALL also report the daemon's spinloop version as a string, set from the ### Requirement: Deploy config push -The API SHALL accept a deploy config in the same shape `spinloop remote deploy` +The API SHALL accept a deploy config in the same shape `spinloop cloud deploy` derives from a Spinloop and its preset (runner, model, context, alias, serve args — the preset already resolved by the pusher). The daemon SHALL validate that the runner names an engine it can serve, persist the config, and use it diff --git a/openspec/specs/endpoint-lifecycle/spec.md b/openspec/specs/endpoint-lifecycle/spec.md index b6cd999b..07c162f5 100644 --- a/openspec/specs/endpoint-lifecycle/spec.md +++ b/openspec/specs/endpoint-lifecycle/spec.md @@ -2,7 +2,7 @@ ## Purpose -Define when the remote endpoint's instance exists — how it is started on +Define when the cloud endpoint's instance exists — how it is started on demand, how it is judged to be still wanted, and the bounds that decide when it is torn down. ## Requirements @@ -216,8 +216,8 @@ stronger guarantee always wins: 1. A **retention override** — an instance marked to be retained until a stated time SHALL NOT be terminated automatically before it, for any reason. The - tag is set from the CLI via `spinloop remote keep DURATION` or - `spinloop remote start --keep DURATION`, which compute an absolute deadline + tag is set from the CLI via `spinloop cloud keep DURATION` or + `spinloop cloud start --keep DURATION`, which compute an absolute deadline from the provided duration and apply it as the `Retain-Until` EC2 tag on the instance. The idle sweep reads this tag and defers automatic termination until the deadline passes. @@ -253,7 +253,7 @@ A manual stop SHALL take effect immediately regardless of all three. #### Scenario: Setting retention from the CLI -- **WHEN** the user runs `spinloop remote keep 4h` on a running instance +- **WHEN** the user runs `spinloop cloud keep 4h` on a running instance - **THEN** the `Retain-Until` tag is set to 4 hours from now, and the idle sweep defers automatic termination until that time @@ -263,12 +263,12 @@ A user-initiated pause SHALL stop the instance without terminating it, preservin #### Scenario: Pause stops without terminating -- **WHEN** user runs `spinloop remote pause` for a running environment +- **WHEN** user runs `spinloop cloud pause` for a running environment - **THEN** the instance is stopped, not terminated, and the environment's URL is retained #### Scenario: Pause is distinct from stop -- **WHEN** user runs `spinloop remote stop` +- **WHEN** user runs `spinloop cloud stop` - **THEN** the instance is terminated immediately ### Requirement: Engine is stopped before the EC2 instance diff --git a/openspec/specs/endpoint-provisioning/spec.md b/openspec/specs/endpoint-provisioning/spec.md index 06ab99ee..f6f6b354 100644 --- a/openspec/specs/endpoint-provisioning/spec.md +++ b/openspec/specs/endpoint-provisioning/spec.md @@ -2,49 +2,49 @@ ## Purpose -Define how the account-level AWS control plane for remote inference -endpoints is provisioned through `spinloop remote bootstrap`. +Define how the account-level AWS control plane for cloud inference +endpoints is provisioned through `spinloop cloud bootstrap`. ## Requirements ### Requirement: Bootstrap deploys the control plane -The system SHALL provide `spinloop remote bootstrap`, which deploys the -account-level control plane that every remote environment reuses — the EC2 +The system SHALL provide `spinloop cloud bootstrap`, which deploys the +account-level control plane that every cloud environment reuses — the EC2 Image Builder pipelines, the environment-aware lifecycle Lambdas and their IAM, and the shared S3 weights bucket, IAM roles and VPC, and the IAM user that -holds the long-lived control-plane credential (see the Remote Auth +holds the long-lived control-plane credential (see the Cloud Auth specification) together with its policy — by obtaining the CDK project shipped in `remote/` and driving its deploy of the control-plane stack. Bootstrap SHALL NOT start any AMI bake; the bake is a separate -`spinloop remote bake` step. Bootstrap SHALL NOT create any Elastic IP or EC2 +`spinloop cloud bake` step. Bootstrap SHALL NOT create any Elastic IP or EC2 instance, and SHALL NOT register an environment; those belong to -`spinloop remote deploy`. Bootstrap SHALL NOT reimplement the infrastructure; +`spinloop cloud deploy`. Bootstrap SHALL NOT reimplement the infrastructure; it SHALL orchestrate the existing CDK project. On success, bootstrap SHALL -signpost `spinloop remote bake` as the next step, ahead of -`spinloop remote deploy`. +signpost `spinloop cloud bake` as the next step, ahead of +`spinloop cloud deploy`. #### Scenario: A successful bootstrap yields the control plane -- **WHEN** `spinloop remote bootstrap` completes +- **WHEN** `spinloop cloud bootstrap` completes - **THEN** the control-plane stack is deployed — Image Builder pipelines, the lifecycle Lambdas, and the shared bucket/roles/VPC — with no Elastic IP or instance created and no AMI bake started #### Scenario: The control-plane credential user is deployed -- **WHEN** `spinloop remote bootstrap` completes +- **WHEN** `spinloop cloud bootstrap` completes - **THEN** the control-plane IAM user exists with a policy covering the - day-to-day remote commands only — invoking the control URLs, reading the + day-to-day cloud commands only — invoking the control URLs, reading the control-plane log groups, describing the control-plane stack, and managing its own access keys — and no permission to deploy, bake, or otherwise provision AWS resources #### Scenario: Bootstrap signposts the bake -- **WHEN** `spinloop remote bootstrap` completes -- **THEN** its output names `spinloop remote bake` as the next step, ahead of - `spinloop remote deploy` +- **WHEN** `spinloop cloud bootstrap` completes +- **THEN** its output names `spinloop cloud bake` as the next step, ahead of + `spinloop cloud deploy` #### Scenario: Orchestration stops on a failed step @@ -54,7 +54,7 @@ signpost `spinloop remote bake` as the next step, ahead of ### Requirement: The control plane is discoverable The control-plane stack SHALL publish, as CloudFormation stack outputs under a -well-known stack name, the values a later `spinloop remote deploy` needs to create +well-known stack name, the values a later `spinloop cloud deploy` needs to create and drive environments: the lifecycle Lambda URLs, the weights bucket, the shared roles, and the region. Discovery SHALL be from those outputs rather than a file bootstrap writes, so it reflects what is actually deployed and works from any @@ -62,7 +62,7 @@ machine with account access. #### Scenario: Deploy can discover the control plane -- **WHEN** the control-plane stack is deployed and `spinloop remote deploy` runs later +- **WHEN** the control-plane stack is deployed and `spinloop cloud deploy` runs later - **THEN** it reads the Lambda URLs, bucket, roles and region from the stack's outputs, without a local file having to carry them @@ -78,13 +78,13 @@ other than an explicit yes as a decline that makes no changes. #### Scenario: The plan is shown before anything is deployed -- **WHEN** the user runs `spinloop remote bootstrap` +- **WHEN** the user runs `spinloop cloud bootstrap` - **THEN** the account, region, control-plane resources, cost caveat, and commands are printed before any AWS-mutating command runs #### Scenario: Dry run changes nothing -- **WHEN** the user runs `spinloop remote bootstrap --dry-run` +- **WHEN** the user runs `spinloop cloud bootstrap --dry-run` - **THEN** the plan is printed and no package-manager, `cdk`, or AWS-mutating command runs @@ -95,7 +95,7 @@ other than an explicit yes as a decline that makes no changes. ### Requirement: Version-matched CDK sources -Bootstrap and `spinloop remote bake` SHALL obtain the CDK project by +Bootstrap and `spinloop cloud bake` SHALL obtain the CDK project by downloading the `remote/` tree from the project repository at a reference matching the running binary's version, so the infrastructure matches the CLI driving it. A `--ref` flag SHALL override the reference, and a `--dir` flag @@ -165,11 +165,11 @@ be confirmed, without attempting to raise it. ### Requirement: A Node package manager is selected, overridable, and logged -Bootstrap and `spinloop remote bake` SHALL select the Node package manager they +Bootstrap and `spinloop cloud bake` SHALL select the Node package manager they drive the CDK project with. Absent an explicit choice, they SHALL auto-detect by PATH lookup, preferring `pnpm` and falling back to `npm` when `pnpm` is not on the path. The user MAY override -the selection with a `--package-manager` flag or an `SPINLOOP_REMOTE_PACKAGE_MANAGER` +the selection with a `--package-manager` flag or an `SPINLOOP_CLOUD_PACKAGE_MANAGER` environment variable, whose only accepted values are `pnpm` and `npm`; the flag SHALL take precedence over the environment variable, which SHALL take precedence over auto-detection. An unrecognised override value SHALL be rejected with an @@ -196,7 +196,7 @@ yet runs correctly under either manager. #### Scenario: An explicit override is honoured - **WHEN** the user passes `--package-manager npm` (or sets - `SPINLOOP_REMOTE_PACKAGE_MANAGER=npm`) while `pnpm` is also present + `SPINLOOP_CLOUD_PACKAGE_MANAGER=npm`) while `pnpm` is also present - **THEN** bootstrap uses `npm` regardless of auto-detection, and the flag wins if both the flag and the environment variable are set @@ -208,26 +208,26 @@ yet runs correctly under either manager. ### Requirement: AMI bake is a separate command -The system SHALL provide `spinloop remote bake`, which starts an AMI bake for +The system SHALL provide `spinloop cloud bake`, which starts an AMI bake for each runner named as a positional argument — `llamacpp` and `vllm` — defaulting to both when none are named. It SHALL drive the same CDK project that bootstrap orchestrates, with the same version-matched source download into the same ref-keyed default location, the same package-manager selection and override, and `--ref` and `--dir` flags matching bootstrap's. Bake SHALL NOT deploy any stack; when the control-plane stack is not deployed, it SHALL fail -before starting any bake, naming `spinloop remote bootstrap` as the step to run +before starting any bake, naming `spinloop cloud bootstrap` as the step to run first. Bake SHALL block until every requested runner's AMI is available; a `--no-wait` flag SHALL return as soon as the bakes are queued, reporting how to check on them, rather than blocking for the bake duration. #### Scenario: Default bake covers both runners -- **WHEN** the user runs `spinloop remote bake` with no arguments +- **WHEN** the user runs `spinloop cloud bake` with no arguments - **THEN** a bake is started for both `llamacpp` and `vllm` #### Scenario: A single runner is baked -- **WHEN** the user runs `spinloop remote bake llamacpp` +- **WHEN** the user runs `spinloop cloud bake llamacpp` - **THEN** only the `llamacpp` AMI bake is started #### Scenario: An unknown runner is rejected @@ -239,11 +239,11 @@ check on them, rather than blocking for the bake duration. - **WHEN** the control-plane stack is not deployed and bake runs - **THEN** it fails before starting any bake, saying to run - `spinloop remote bootstrap` first + `spinloop cloud bootstrap` first #### Scenario: Bake waits by default -- **WHEN** the user runs `spinloop remote bake` without `--no-wait` +- **WHEN** the user runs `spinloop cloud bake` without `--no-wait` - **THEN** the command blocks until the requested runners' AMIs are available before finishing @@ -273,31 +273,31 @@ Bootstrap SHALL collect the one control-plane setting the CDK has no default for and write it where the CDK reads it: an optional Hugging Face token for the shared secret used when seeding gated weights. Which runner AMIs to bake is not a bootstrap setting — the engine is a per-environment choice made at -`deploy`, and the runners are named by `spinloop remote bake` itself. The +`deploy`, and the runners are named by `spinloop cloud bake` itself. The allowed ingress CIDR is also per-environment and belongs to `deploy`, not here. #### Scenario: Runners are not a bootstrap setting - **WHEN** the user runs bootstrap - **THEN** no runner selection is requested or written, since the runners are - named at `spinloop remote bake` + named at `spinloop cloud bake` #### Scenario: The allowed CIDR is not a bootstrap setting - **WHEN** the user runs bootstrap - **THEN** no ingress CIDR is requested or written, since it is scoped per - environment at `spinloop remote deploy` + environment at `spinloop cloud deploy` ### Requirement: Bootstrap records the deploying CLI's version -`spinloop remote bootstrap` SHALL pass the running binary's version to the control plane deploy, so every control plane Lambda reports it in the `x-spinloop-control-plane-version` response header. A binary built without a version SHALL record `dev`. +`spinloop cloud bootstrap` SHALL pass the running binary's version to the control plane deploy, so every control plane Lambda reports it in the `x-spinloop-control-plane-version` response header. A binary built without a version SHALL record `dev`. #### Scenario: A release build bootstraps -- **WHEN** a CLI at version `1.30.0` runs `spinloop remote bootstrap` +- **WHEN** a CLI at version `1.30.0` runs `spinloop cloud bootstrap` - **THEN** the deployed Lambdas report `1.30.0` as the control plane version #### Scenario: A development build bootstraps -- **WHEN** a CLI built without a version override runs `spinloop remote bootstrap` +- **WHEN** a CLI built without a version override runs `spinloop cloud bootstrap` - **THEN** the deployed Lambdas report `dev` diff --git a/openspec/specs/engine-metrics/spec.md b/openspec/specs/engine-metrics/spec.md index c79a1d10..7a296a95 100644 --- a/openspec/specs/engine-metrics/spec.md +++ b/openspec/specs/engine-metrics/spec.md @@ -12,7 +12,7 @@ worked only for the cloud, and only from outside. Every stat is optional by design: a host with no source for one omits it rather than erroring, which is how a machine without `nvidia-smi` reports engine stats and no GPU figures. The shape is kept value-for-value compatible with what the existing -`spinloop remote metrics` formatters render. +`spinloop cloud metrics` formatters render. ## Requirements @@ -152,7 +152,7 @@ would bury the failures worth seeing. ### Requirement: Rendering-compatible stats shape The collected metrics SHALL be expressible in the same stats shape the -`spinloop remote metrics` formatters render (state, runner, model, GPU, CPU, +`spinloop cloud metrics` formatters render (state, runner, model, GPU, CPU, RAM, token stats, and the engine's last-active time with the idle duration derived from it), so the existing bar, table, and JSON formats display in-process metrics without format-specific changes. diff --git a/openspec/specs/environment-deployment/spec.md b/openspec/specs/environment-deployment/spec.md index 4c5b812c..de248772 100644 --- a/openspec/specs/environment-deployment/spec.md +++ b/openspec/specs/environment-deployment/spec.md @@ -2,24 +2,24 @@ ## Purpose -Define how `spinloop remote deploy` creates a named environment on the +Define how `spinloop cloud deploy` creates a named environment on the account-level control plane: discovering it, provisioning per-environment resources, registering the environment, and guarding against accidental overwrites. ## Requirements ### Requirement: Deploy creates an environment on the control plane -`spinloop remote deploy` SHALL create a named environment on top of the control +`spinloop cloud deploy` SHALL create a named environment on top of the control plane: it SHALL discover it, then provision the environment's own Elastic IP, EC2 instance configuration, per-environment API key, per-environment allowed-ingress rule, and per-environment SSM state (the deploy-config), all tagged by the environment name. It SHALL set what the environment serves from the Spinloop and its preset, and SHALL register -the environment so the other `remote` commands can drive it. Deploying SHALL NOT +the environment so the other `cloud` commands can drive it. Deploying SHALL NOT start the instance. The environment name SHALL come from the command, never from the Spinloop: -`spinloop remote deploy` SHALL take it from its required `--env ` flag, +`spinloop cloud deploy` SHALL take it from its required `--env ` flag, and `spinloop fleet deploy` SHALL take it from the name of the node being deployed, which is the registered environment that node drives. A deploy given no name by either route SHALL fail saying the environment must be named. The @@ -33,7 +33,7 @@ plane to seed, read or write. #### Scenario: Deploying stands up and registers an environment -- **WHEN** `spinloop remote deploy --env prod` runs against a bootstrapped +- **WHEN** `spinloop cloud deploy --env prod` runs against a bootstrapped account - **THEN** the environment's Elastic IP, instance configuration, API key, ingress rule, and SSM state are provisioned, and the environment is @@ -41,7 +41,7 @@ plane to seed, read or write. #### Scenario: A deploy without a name fails -- **WHEN** `spinloop remote deploy` runs with no `--env` flag +- **WHEN** `spinloop cloud deploy` runs with no `--env` flag - **THEN** it fails saying the environment must be named with `--env ` #### Scenario: One Spinloop deploys to two environments @@ -53,7 +53,7 @@ plane to seed, read or write. #### Scenario: A fleet node deploys under its own name -- **WHEN** `spinloop fleet deploy` targets a `kind: remote` node named `qwen` +- **WHEN** `spinloop fleet deploy` targets a `kind: cloud` node named `qwen` - **THEN** the environment created and registered is named `qwen`, from the node's name in the fleet file, and the node's Spinloop is read only for what it serves @@ -68,14 +68,14 @@ plane to seed, read or write. - **WHEN** a deploy succeeds - **THEN** the environment is configured and registered but no instance is - running until `spinloop remote start` + running until `spinloop cloud start` ### Requirement: Discovering the control plane Deploy SHALL discover the control plane from the bootstrap stack's CloudFormation outputs (a well-known stack name) — the lifecycle Lambda URLs, the weights bucket, the shared roles, and the region — rather than from any local file. When the control-plane stack is absent, deploy SHALL fail telling the user to run -`spinloop remote bootstrap` first, rather than attempting to create an environment. +`spinloop cloud bootstrap` first, rather than attempting to create an environment. #### Scenario: The control plane is discovered @@ -86,7 +86,7 @@ file. When the control-plane stack is absent, deploy SHALL fail telling the user #### Scenario: Not bootstrapped - **WHEN** deploy runs against an account with no control-plane stack -- **THEN** it fails saying to run `spinloop remote bootstrap` first, and creates +- **THEN** it fails saying to run `spinloop cloud bootstrap` first, and creates nothing ### Requirement: Per-environment allowed ingress @@ -109,16 +109,16 @@ rule. It SHALL NOT be an account-wide setting. ### Requirement: Registering the environment Deploy SHALL register the environment in the per-user registry defined by the -Remote Environments specification — `~/.config/spinloop/remotes//remote.json`, +Cloud Environments specification — `~/.config/spinloop/clouds//cloud.json`, written owner-only — carrying the shared lifecycle Lambda URLs, the region, the environment's base URL (its Elastic IP), and the environment identifier the shared Lambdas use to select this environment's instance. #### Scenario: The environment is registered and resolvable -- **WHEN** `spinloop remote deploy --env prod` succeeds -- **THEN** `~/.config/spinloop/remotes/prod/remote.json` exists (owner-only) - and `spinloop remote status --env prod` resolves the environment from it +- **WHEN** `spinloop cloud deploy --env prod` succeeds +- **THEN** `~/.config/spinloop/clouds/prod/cloud.json` exists (owner-only) + and `spinloop cloud status --env prod` resolves the environment from it ### Requirement: Refuse to overwrite a live environment @@ -147,7 +147,7 @@ warning and without requiring `--overwrite`. ### Requirement: Deploy accepts an optional spinloop version pin -`spinloop remote deploy` SHALL accept an optional flag pinning the exact spinloop +`spinloop cloud deploy` SHALL accept an optional flag pinning the exact spinloop release the environment's instances install at boot. When the flag is given, deploy SHALL record that version in the environment's stored deploy config so the next fresh boot installs it; when it is absent, deploy SHALL record no pin @@ -157,44 +157,44 @@ whitespace-only value SHALL be treated as if the flag were not given. #### Scenario: A pin is recorded in the deploy config -- **WHEN** `spinloop remote deploy` runs with an spinloop version pin +- **WHEN** `spinloop cloud deploy` runs with an spinloop version pin - **THEN** the environment's stored deploy config carries that version, and the environment's next fresh boot installs exactly that release #### Scenario: No pin leaves the boot on its default -- **WHEN** `spinloop remote deploy` runs without an spinloop version pin +- **WHEN** `spinloop cloud deploy` runs without an spinloop version pin - **THEN** the stored deploy config carries no spinloop version, and the environment's boots install the latest published release #### Scenario: An empty pin value is ignored -- **WHEN** `spinloop remote deploy` is given an spinloop version pin whose value is +- **WHEN** `spinloop cloud deploy` is given an spinloop version pin whose value is empty or whitespace only - **THEN** it is treated as if no pin were given ### Requirement: The deploy plan shows the resolved spinloop version -The plan `spinloop remote deploy` prints — including under `--dry-run`, before +The plan `spinloop cloud deploy` prints — including under `--dry-run`, before any AWS work or send — SHALL state the spinloop version the environment's boots will install: the pinned version when a pin is given, otherwise `latest`. It SHALL appear alongside the runner and model the plan already prints. #### Scenario: A pinned deploy prints the pinned version -- **WHEN** `spinloop remote deploy --dry-run` runs with an spinloop version pin +- **WHEN** `spinloop cloud deploy --dry-run` runs with an spinloop version pin - **THEN** the printed plan names that pinned version as the spinloop the environment will run #### Scenario: An unpinned deploy prints latest -- **WHEN** `spinloop remote deploy --dry-run` runs without an spinloop version pin +- **WHEN** `spinloop cloud deploy --dry-run` runs without an spinloop version pin - **THEN** the printed plan names `latest` as the spinloop the environment will run ### Requirement: Deploy accepts an optional instance type -`spinloop remote deploy` SHALL accept an optional `--instance-type` flag naming +`spinloop cloud deploy` SHALL accept an optional `--instance-type` flag naming the EC2 instance type the environment's instances launch as. When the flag is given, deploy SHALL record that type in the environment's stored deploy config so the environment's next fresh launch uses it; when it is absent, deploy SHALL @@ -209,32 +209,32 @@ of a single start, and never derived from the Spinloop. #### Scenario: A type is recorded in the deploy config -- **WHEN** `spinloop remote deploy` runs with `--instance-type g6e.2xlarge` +- **WHEN** `spinloop cloud deploy` runs with `--instance-type g6e.2xlarge` - **THEN** the environment's stored deploy config carries that type, and the environment's next fresh launch uses it #### Scenario: No type leaves the launch on its default -- **WHEN** `spinloop remote deploy` runs without `--instance-type` +- **WHEN** `spinloop cloud deploy` runs without `--instance-type` - **THEN** the stored deploy config carries no instance type, and the environment's launches use the control plane's default type #### Scenario: An empty type value is ignored -- **WHEN** `spinloop remote deploy` is given an `--instance-type` whose value +- **WHEN** `spinloop cloud deploy` is given an `--instance-type` whose value is empty or whitespace only - **THEN** it is treated as if no type were given #### Scenario: A malformed type is refused before sending -- **WHEN** `spinloop remote deploy` is given an `--instance-type` that is not +- **WHEN** `spinloop cloud deploy` is given an `--instance-type` that is not shaped like an EC2 instance type - **THEN** the command fails, naming the value, and nothing is sent to the control plane ### Requirement: The deploy plan shows the resolved instance type -The plan `spinloop remote deploy` prints — including under `--dry-run`, before +The plan `spinloop cloud deploy` prints — including under `--dry-run`, before any AWS work or send — SHALL state the instance type the environment will launch as: the type named by `--instance-type` when one is given, otherwise a statement that the environment launches as the control plane's default. It @@ -242,20 +242,20 @@ SHALL appear alongside the runner and model the plan already prints. #### Scenario: A typed deploy prints the type -- **WHEN** `spinloop remote deploy --dry-run` runs with `--instance-type +- **WHEN** `spinloop cloud deploy --dry-run` runs with `--instance-type g6e.2xlarge` - **THEN** the printed plan names `g6e.2xlarge` as the instance type the environment will launch as #### Scenario: An untyped deploy prints the default -- **WHEN** `spinloop remote deploy --dry-run` runs without `--instance-type` +- **WHEN** `spinloop cloud deploy --dry-run` runs without `--instance-type` - **THEN** the printed plan says the environment launches as the control plane's default instance type ### Requirement: Externally provided API key -`spinloop remote deploy` SHALL accept an externally provided API key as a +`spinloop cloud deploy` SHALL accept an externally provided API key as a reference to an environment variable, and pass it to the control plane to store as the environment's API key. The value SHALL NOT be written on the command line or in any file the CLI owns: the flag names a variable, and the CLI resolves it @@ -280,21 +280,21 @@ it as 401s. #### Scenario: A deploy stores a supplied key -- **WHEN** `spinloop remote deploy` is given a key variable that is set, for an +- **WHEN** `spinloop cloud deploy` is given a key variable that is set, for an environment whose API-key secret does not yet exist - **THEN** the environment's API-key secret is created holding that value, and the report says the key was applied #### Scenario: A deploy rotates an existing key -- **WHEN** `spinloop remote deploy` is given a key for an environment that already +- **WHEN** `spinloop cloud deploy` is given a key for an environment that already has an API-key secret - **THEN** the secret is set to the new value, the old key is no longer valid, and the report says the key was rotated #### Scenario: A deploy without a key keeps the existing one -- **WHEN** `spinloop remote deploy` runs with no key for an environment that +- **WHEN** `spinloop cloud deploy` runs with no key for an environment that already has an API-key secret - **THEN** the secret is left unchanged and no new key is generated @@ -302,10 +302,10 @@ it as 401s. - **WHEN** a deploy supplies a key - **THEN** the value is not written to the environment's deploy-config, is not - in the registered remote configuration, and is not printed in any reply or in + in the registered cloud configuration, and is not printed in any reply or in the deploy report #### Scenario: A named variable that is unset fails early -- **WHEN** `spinloop remote deploy` names a key variable that is set nowhere +- **WHEN** `spinloop cloud deploy` names a key variable that is set nowhere - **THEN** the deploy fails naming the variable, before anything is sent diff --git a/openspec/specs/fleet-client/spec.md b/openspec/specs/fleet-client/spec.md index cb649231..9a840ebf 100644 --- a/openspec/specs/fleet-client/spec.md +++ b/openspec/specs/fleet-client/spec.md @@ -172,7 +172,7 @@ same node-owned derivation a routed wake already uses (`deployConfigForNode`) — report the resolved source and derived config alongside the node's name, and start the node's engine with that config (`StartWith`) rather than a plain start, exactly as a routed wake tells a -node what to serve. A `kind: remote` node's start is unaffected regardless +node what to serve. A `kind: cloud` node's start is unaffected regardless of whether a source resolves for it: what it serves is fixed at deploy time, not pushed at start time, so it always uses a plain start. @@ -193,10 +193,10 @@ not pushed at start time, so it always uses a plain start. #### Scenario: Start every node - **WHEN** `spinloop fleet start --all` runs against a file mixing `kind: - remote` and `kind: daemon` nodes, and every `kind: daemon` node's Spinloop + cloud` and `kind: daemon` nodes, and every `kind: daemon` node's Spinloop source resolves - **THEN** every node in the file starts — the daemon nodes with their - resolved config, the remote nodes with a plain start + resolved config, the cloud nodes with a plain start #### Scenario: Start with no node names the fleet @@ -228,7 +228,7 @@ not pushed at start time, so it always uses a plain start. #### Scenario: Stop every node - **WHEN** `spinloop fleet stop --all` runs against a file mixing `kind: - remote` and `kind: daemon` nodes + cloud` and `kind: daemon` nodes - **THEN** every node in the file is stopped #### Scenario: Stop with no node names the fleet @@ -264,47 +264,47 @@ not pushed at start time, so it always uses a plain start. as failed naming the three ways a source could have been given, and the command exits non-zero -#### Scenario: Starting a remote node is unaffected by a resolved source +#### Scenario: Starting a cloud node is unaffected by a resolved source -- **WHEN** `spinloop fleet start gpu-env` runs, `gpu-env` is a `kind: remote` +- **WHEN** `spinloop fleet start gpu-env` runs, `gpu-env` is a `kind: cloud` node, and a Spinloop source resolves for it - **THEN** the client starts it with a plain start; the resolved source is - not used, since a `kind: remote` node's `StartWith` always refuses a + not used, since a `kind: cloud` node's `StartWith` always refuses a deploy config -### Requirement: Fleet deploy targets remote nodes +### Requirement: Fleet deploy targets cloud nodes `spinloop fleet deploy ` SHALL deploy the AWS environment for one or -more `kind: remote` nodes in the fleet file, named explicitly. -`spinloop fleet deploy --all` SHALL target every `kind: remote` node in the +more `kind: cloud` nodes in the fleet file, named explicitly. +`spinloop fleet deploy --all` SHALL target every `kind: cloud` node in the file instead. Invoked with neither a node name nor `--all`, it SHALL fail, -listing the fleet's `kind: remote` nodes, and deploy nothing — mutating +listing the fleet's `kind: cloud` nodes, and deploy nothing — mutating however many cloud environments a fleet file lists SHALL NOT happen by default. `--all` combined with one or more node names SHALL fail as ambiguous. An unknown node name SHALL fail the command, naming the known nodes, without deploying anything. A named `kind: daemon` node SHALL fail the command, explaining that `fleet deploy` provisions cloud environments -and that node is not one; `--all` SHALL only ever select `kind: remote` +and that node is not one; `--all` SHALL only ever select `kind: cloud` nodes, so a `kind: daemon` node is never targeted by it and is not reported at all. -#### Scenario: Deploy every remote node +#### Scenario: Deploy every cloud node - **WHEN** `spinloop fleet deploy --all` runs against a file mixing `kind: - remote` and `kind: daemon` nodes -- **THEN** every `kind: remote` node is deployed and no `kind: daemon` node + cloud` and `kind: daemon` nodes +- **THEN** every `kind: cloud` node is deployed and no `kind: daemon` node is touched or mentioned #### Scenario: Deploy named nodes - **WHEN** `spinloop fleet deploy gpu-a gpu-b` runs and both are `kind: - remote` nodes in the file + cloud` nodes in the file - **THEN** only those two are deployed, whatever else the file lists #### Scenario: No target is an error - **WHEN** `spinloop fleet deploy` runs with no node arguments and no `--all` -- **THEN** it fails, listing the fleet's `kind: remote` nodes, and deploys +- **THEN** it fails, listing the fleet's `kind: cloud` nodes, and deploys nothing #### Scenario: Combining --all with node names is an error @@ -331,7 +331,7 @@ source resolves to (see fleet-config's "Node Spinloop source" and "...falls back to name-based lookup" requirements: its `file` field, else an alias registered under its name, else a `/` subdirectory beside the fleet file), deriving the deploy config and registering the resulting environment -exactly as `spinloop remote deploy ` does for that same file — the two +exactly as `spinloop cloud deploy ` does for that same file — the two SHALL NOT be able to disagree about what a given Spinloop file deploys. A targeted node for which no source resolves SHALL fail for that node alone, naming all three ways one could have been given, without touching the other @@ -341,14 +341,14 @@ three supplied it is never left to be inferred. Where a node declares an `instance-type` in the fleet file, the deploy config derived for it SHALL carry that type, so the node's environment launches as -named — the same value a standalone `spinloop remote deploy --instance-type` +named — the same value a standalone `spinloop cloud deploy --instance-type` would record for the environment — and a node declaring none SHALL deploy an environment on the control plane's default type. This keeps `fleet deploy` and a matching standalone deploy in agreement about what a node's environment launches as. Nodes SHALL be deployed independently: one node already registered or live -SHALL require `--overwrite` for that node exactly as a standalone `remote +SHALL require `--overwrite` for that node exactly as a standalone `cloud deploy` does, and refusing it SHALL NOT stop the other targeted nodes from deploying. A node whose deploy fails for any other reason SHALL likewise be reported against that node without aborting the rest. The command SHALL exit @@ -356,7 +356,7 @@ non-zero when any targeted node failed to deploy, having still attempted every other targeted node. `--dry-run` SHALL print the plan for every targeted node without deploying -any of them, exactly as a standalone `remote deploy --dry-run` does for one. +any of them, exactly as a standalone `cloud deploy --dry-run` does for one. `--overwrite` SHALL apply to every targeted node that needs it. #### Scenario: A node deploys from its own Spinloop file @@ -364,27 +364,27 @@ any of them, exactly as a standalone `remote deploy --dry-run` does for one. - **WHEN** `fleet deploy` targets a node declaring `file: ./envs/gpu.Spinloop` - **THEN** that node's environment is created and registered from that file, - the same as `spinloop remote deploy ./envs/gpu.Spinloop` would produce, and + the same as `spinloop cloud deploy ./envs/gpu.Spinloop` would produce, and the resolved path is reported against that node #### Scenario: A node's declared instance type is deployed -- **WHEN** `fleet deploy` targets a `kind: remote` node declaring +- **WHEN** `fleet deploy` targets a `kind: cloud` node declaring `instance-type: g6e.2xlarge` - **THEN** the environment it deploys launches as `g6e.2xlarge`, the same - value a standalone `spinloop remote deploy --instance-type g6e.2xlarge` of + value a standalone `spinloop cloud deploy --instance-type g6e.2xlarge` of the node's source would record #### Scenario: A node with no instance type deploys the default -- **WHEN** `fleet deploy` targets a `kind: remote` node declaring no +- **WHEN** `fleet deploy` targets a `kind: cloud` node declaring no `instance-type` - **THEN** the environment it deploys launches as the control plane's default instance type #### Scenario: A node with no resolvable source fails only that node -- **WHEN** `fleet deploy` targets two remote nodes and one declares no `file` +- **WHEN** `fleet deploy` targets two cloud nodes and one declares no `file` field, has no alias registered under its name, and has no same-named subdirectory beside the fleet file - **THEN** the other node still deploys, and the command reports against the @@ -393,16 +393,16 @@ any of them, exactly as a standalone `remote deploy --dry-run` does for one. #### Scenario: One node's guard does not block the others -- **WHEN** `fleet deploy` targets two remote nodes and one is already +- **WHEN** `fleet deploy` targets two cloud nodes and one is already registered while the other is not, and `--overwrite` is not given - **THEN** the unregistered node deploys, the registered node is refused with - the same message a standalone `remote deploy` gives, and the command exits + the same message a standalone `cloud deploy` gives, and the command exits non-zero #### Scenario: Dry run previews every targeted node - **WHEN** `spinloop fleet deploy --dry-run --all` runs -- **THEN** the plan for every `kind: remote` node in the file is printed and +- **THEN** the plan for every `kind: cloud` node in the file is printed and no environment is created or registered ### Requirement: Fleet deploy reports progress and results legibly @@ -1134,7 +1134,7 @@ the grid on escape rather than closing the dashboard out from under it. The dashboard SHALL refresh the fleet continuously, and SHALL also refresh immediately on the operator's request. The cadence SHALL be by node kind: a local daemon machine SHALL refresh on a short interval — seconds, not the -watch mode's minute — and a `kind: remote` environment SHALL refresh on a +watch mode's minute — and a `kind: cloud` environment SHALL refresh on a much slower cadence, a 60-second interval, one status call a minute, because its status is a signed call through the cloud control plane rather than a local socket, and its state changes on the scale of minutes. A manual refresh @@ -1269,7 +1269,7 @@ heard from yet. - **Not serving**: the node answered its last refresh, that answer is current, no action is in flight for it, and its engine is not serving — its state is `idle`, the daemon has started nothing, `stopped`, a daemon engine that was - stopped, or `undeployed`, a remote environment with no instance at all. + stopped, or `undeployed`, a cloud environment with no instance at all. - **Unknown**: no current status can be determined for the node — it has not yet answered any refresh, its last refresh answered without reporting an engine state, or its newest answer has aged well past its cadence and no longer @@ -1344,9 +1344,9 @@ heard from yet. - **THEN** its panel's status glyph is the faded grey dot, not the green of a serving node -#### Scenario: An undeployed remote environment reads not serving +#### Scenario: An undeployed cloud environment reads not serving -- **WHEN** a remote environment's last completed refresh reports it +- **WHEN** a cloud environment's last completed refresh reports it `undeployed` — it has no instance at all — and no start or stop is in flight for it - **THEN** its panel's status glyph is the faded grey dot, not the green of @@ -1451,7 +1451,7 @@ The dashboard SHALL let the operator set the retention deadline of the node currently selected, from the keyboard, through the same node operation the one-shot keep command uses. Keep applies to a node's instance — the time until which the cloud's idle sweep leaves it alone — so it SHALL be offered only for -nodes that support retention, namely the fleet's remote environments; a node +nodes that support retention, namely the fleet's cloud environments; a node without retention support SHALL take no keep. The keep key SHALL open a duration prompt rather than send anything. The @@ -1494,7 +1494,7 @@ prompt and the same rules the grid applies to the selected node. #### Scenario: Keep opens a prompt, it does not send -- **WHEN** the operator selects a remote environment with no action in flight +- **WHEN** the operator selects a cloud environment with no action in flight and issues keep - **THEN** a duration prompt opens naming the node, pre-filled with a default duration, and nothing has been sent @@ -1507,7 +1507,7 @@ prompt and the same rules the grid applies to the selected node. #### Scenario: A pre-filled keep is one key and a confirm -- **WHEN** the operator issues keep on a remote environment and confirms the +- **WHEN** the operator issues keep on a cloud environment and confirms the prompt without changing its duration - **THEN** the node is kept until now plus the default duration @@ -1538,7 +1538,7 @@ prompt and the same rules the grid applies to the selected node. - **WHEN** the node under the cursor is a local daemon node, or has an action in flight - **THEN** the key help does not name the keep key -- **WHEN** that node is a remote environment with no action in flight +- **WHEN** that node is a cloud environment with no action in flight - **THEN** the key help names the keep key #### Scenario: A busy node is not kept again @@ -1562,7 +1562,7 @@ prompt and the same rules the grid applies to the selected node. #### Scenario: Keep from the detail view -- **WHEN** the operator issues keep from the detail view of a remote +- **WHEN** the operator issues keep from the detail view of a cloud environment - **THEN** the same prompt opens with the same rules, and its outcome is shown as the grid would show it @@ -1585,7 +1585,7 @@ the sweep has already moved past as though the node were still held. #### Scenario: A retained node shows its deadline -- **WHEN** a remote environment's refresh answer carries a retention deadline +- **WHEN** a cloud environment's refresh answer carries a retention deadline in the future - **THEN** its tile shows the deadline, and its detail screen shows the same line @@ -1598,7 +1598,7 @@ the sweep has already moved past as though the node were still held. #### Scenario: An older control plane degrades to no line -- **WHEN** a remote environment's control plane predates the deadline in its +- **WHEN** a cloud environment's control plane predates the deadline in its stats reply - **THEN** its panel shows no deadline line and no error, and the rest of the panel is unaffected diff --git a/openspec/specs/fleet-config/spec.md b/openspec/specs/fleet-config/spec.md index 6a9aad3e..81582055 100644 --- a/openspec/specs/fleet-config/spec.md +++ b/openspec/specs/fleet-config/spec.md @@ -142,7 +142,7 @@ one thing across the group. #### Scenario: A path is not an environment name -- **WHEN** `spinloop fleet status --env ./remote.json` runs +- **WHEN** `spinloop fleet status --env ./cloud.json` runs - **THEN** it fails saying an environment name is a plain identifier with no path @@ -258,15 +258,15 @@ the variable, in the same way a missing daemon token is. ### Requirement: Fleet-wide API key reference A `fleet.yaml` MAY declare a top-level `apiKeyEnv` naming the environment -variable that holds the API key shared by the fleet's remote nodes. The file +variable that holds the API key shared by the fleet's cloud nodes. The file SHALL hold the variable's *name*, never the value — the same discipline as the daemon and engine-token references — and the reference SHALL be resolved exactly the way those are: the process environment first, then the `.env` beside the fleet file. -The reference is the default key for a remote node. A remote node whose own +The reference is the default key for a cloud node. A cloud node whose own entry names no `engineTokenEnv` takes the fleet-wide key. A node's own -`engineTokenEnv` SHALL override the fleet-wide reference, so one remote may +`engineTokenEnv` SHALL override the fleet-wide reference, so one cloud node may carry a distinct key while the rest of the fleet shares one. A daemon node SHALL NOT take the fleet-wide reference: it is gated only by its own `engineTokenEnv`, exactly as it is today. @@ -277,12 +277,12 @@ naming the variable, in the same way a missing engine-token variable is. #### Scenario: A fleet shares one key across its remotes - **WHEN** a `fleet.yaml` declares `apiKeyEnv: SHARED_KEY`, that variable is - set, and it lists two `kind: remote` nodes that name no `engineTokenEnv` + set, and it lists two `kind: cloud` nodes that name no `engineTokenEnv` - **THEN** the value of `SHARED_KEY` is the key both remotes are reached with #### Scenario: A per-node reference overrides the fleet-wide one -- **WHEN** a `fleet.yaml` declares `apiKeyEnv: SHARED_KEY` and one remote node +- **WHEN** a `fleet.yaml` declares `apiKeyEnv: SHARED_KEY` and one cloud node names `engineTokenEnv: SPECIAL_KEY` - **THEN** that node is reached with the value of `SPECIAL_KEY` and the other remotes with the value of `SHARED_KEY` @@ -295,7 +295,7 @@ naming the variable, in the same way a missing engine-token variable is. #### Scenario: An unset fleet-wide variable names itself - **WHEN** a `fleet.yaml` declares an `apiKeyEnv` that is set nowhere, and a - remote node naming no key of its own is reached for its key + cloud node naming no key of its own is reached for its key - **THEN** the failure names that variable, and no agent is launched without a key @@ -309,7 +309,7 @@ naming the variable, in the same way a missing engine-token variable is. A fleet-file node, of either kind, MAY declare a `file` field naming the Spinloop file that describes what it runs — the same file `spinloop fleet -deploy` reads to create a `kind: remote` node's environment, and the same +deploy` reads to create a `kind: cloud` node's environment, and the same file `spinloop fleet start` reads to tell a `kind: daemon` node's engine what to run. The path SHALL resolve relative to the fleet file's directory, the same way other Spinloop-relative paths in the project resolve. The field @@ -318,9 +318,9 @@ SHALL NOT be required to parse a fleet file — every fleet command other than SHALL require it (directly or via the fallbacks below) for the nodes they act on; see fleet-client's "Driving one node" requirement. -#### Scenario: A remote node names its Spinloop file +#### Scenario: A cloud node names its Spinloop file -- **WHEN** a `kind: remote` node declares `file: ./envs/gpu.Spinloop` +- **WHEN** a `kind: cloud` node declares `file: ./envs/gpu.Spinloop` - **THEN** `spinloop fleet deploy` for that node reads the Spinloop at that path, resolved relative to the fleet file's directory, to derive what to deploy @@ -344,14 +344,14 @@ A node declaring no `file` field SHALL have its Spinloop source resolved from its own `name`, tried in order: 1. `name` resolved as a registered `spinloop alias` — the same lookup a bare - argument to `spinloop remote deploy ` already performs. + argument to `spinloop cloud deploy ` already performs. 2. Failing that, a subdirectory named `` beside the fleet file, containing a Spinloop file — the same directory-to-default-file resolution an ordinary Spinloop path argument already gets when it names a directory. A node for which neither resolves SHALL fail the command acting on it — -`fleet deploy` for a `kind: remote` node, `fleet start` for a `kind: daemon` +`fleet deploy` for a `kind: cloud` node, `fleet start` for a `kind: daemon` node — for that node alone, naming all three ways a source could have been given: the `file` field, a `spinloop alias` named after the node, or a `/` subdirectory beside the fleet file. @@ -360,7 +360,7 @@ given: the `file` field, a `spinloop alias` named after the node, or a - **WHEN** a node named `gpu-env` declares no `file` field, and `spinloop alias` has `gpu-env` registered to a Spinloop path -- **THEN** `fleet deploy` (if `gpu-env` is `kind: remote`) or `fleet start` +- **THEN** `fleet deploy` (if `gpu-env` is `kind: cloud`) or `fleet start` (if `kind: daemon`) reads the Spinloop the alias names #### Scenario: Resolved through a named subdirectory @@ -368,7 +368,7 @@ given: the `file` field, a `spinloop alias` named after the node, or a - **WHEN** a node named `dev-1` declares no `file` field, no alias named `dev-1` is registered, and a `dev-1/` directory containing a Spinloop file sits beside the fleet file -- **THEN** `fleet deploy` (if `dev-1` is `kind: remote`) or `fleet start` (if +- **THEN** `fleet deploy` (if `dev-1` is `kind: cloud`) or `fleet start` (if `kind: daemon`) reads the Spinloop from that subdirectory #### Scenario: An alias wins over a same-named subdirectory @@ -378,9 +378,9 @@ given: the `file` field, a `spinloop alias` named after the node, or a file also sits beside the fleet file - **THEN** the alias is used, not the subdirectory -#### Scenario: None of the three resolve for a remote node +#### Scenario: None of the three resolve for a cloud node -- **WHEN** a `kind: remote` node declares no `file` field, no alias is +- **WHEN** a `kind: cloud` node declares no `file` field, no alias is registered under its name, and no same-named subdirectory sits beside the fleet file - **THEN** `fleet deploy` fails for that node, naming the `file` field, the @@ -460,11 +460,11 @@ shape as the fleet-wide setting. When a node names one, it decides whether that node may be woken, taking precedence over the fleet-wide setting for that node alone; a node naming none is governed by the fleet-wide setting as before. This exists because waking is not free the same way on every node: a -remote environment's wake boots and pays for a cloud instance, unlike a local +cloud environment's wake boots and pays for a cloud instance, unlike a local daemon's engine, so an operator may want the fleet's daemons to wake freely -while deciding a remote node's waking on its own terms — opted in under a +while deciding a cloud node's waking on its own terms — opted in under a fleet that otherwise does not wake, or opted out under one that does — -without a second fleet-wide flag governing every remote node in the file +without a second fleet-wide flag governing every cloud node in the file alike. A file declaring nothing at either level SHALL wake, as routing does when no @@ -585,7 +585,7 @@ section SHALL behave exactly as it does today. A `--env ` target SHALL be a fleet holding exactly one node: a cloud node named by the flag, whose registered configuration is the one -`remotes//remote.json` holds. It SHALL be driven, observed and rendered +`clouds//cloud.json` holds. It SHALL be driven, observed and rendered exactly as the same node listed in a fleet file is, so a command's output for one environment does not depend on how that environment was named. diff --git a/openspec/specs/fleet-gateway/spec.md b/openspec/specs/fleet-gateway/spec.md index fd4bfc57..4206ab2b 100644 --- a/openspec/specs/fleet-gateway/spec.md +++ b/openspec/specs/fleet-gateway/spec.md @@ -96,18 +96,18 @@ running and is a wake candidate — waking is allowed for it (its own `wake` setting, or the fleet-wide one when it names none) and it names a model to start with — the list SHALL additionally carry that model, under the same served-name-first naming a wake would start it with: a daemon node's own -Spinloop source describes it, and a remote node's own stats reply carries it +Spinloop source describes it, and a cloud node's own stats reply carries it — the environment's stored deploy config, which the stats reply carries whether the environment is running or stopped (unlike the status reply, which only relays it while running). A running node SHALL contribute nothing but what it reports: a running engine is never displaced, so its source's -model is not a request the gateway would answer from it. An undeployed remote +model is not a request the gateway would answer from it. An undeployed cloud node environment — one whose stats read fails outright, having no deploy config to read — and a node for which waking is not allowed SHALL contribute nothing beyond what is running, and duplicates SHALL be listed once. The model a node would be started with SHALL be resolved at most once in a short window shared by all models requests, so a burst does not re-read every -node's source or re-fetch every remote node's stats. +node's source or re-fetch every cloud node's stats. #### Scenario: Running models are listed @@ -128,16 +128,16 @@ node's source or re-fetch every remote node's stats. request is made - **THEN** the response lists only what the running nodes serve -#### Scenario: A deployed remote environment's model is listed +#### Scenario: A deployed cloud environment's model is listed -- **WHEN** a remote environment is stopped, its stats reply reports what its +- **WHEN** a cloud environment is stopped, its stats reply reports what its stored deploy config would serve, and waking is allowed for it, and a models request is made - **THEN** the response lists that model beside what the running nodes serve -#### Scenario: A stopped remote environment's model is not listed +#### Scenario: A stopped cloud environment's model is not listed -- **WHEN** a remote environment is stopped and has nothing deployed, and a +- **WHEN** a cloud environment is stopped and has nothing deployed, and a models request is made - **THEN** the response does not list it: the gateway has nothing stored to start it with @@ -151,7 +151,7 @@ node's source or re-fetch every remote node's stats. #### Scenario: A burst of models requests resolves each source once - **WHEN** several models requests arrive within the window in which a node's - source or a remote node's status is resolved + source or a cloud node's status is resolved - **THEN** each node's source or status is read once for the burst #### Scenario: Nothing reachable lists nothing @@ -198,7 +198,7 @@ the chosen engine gives is the reply the caller gets. The engine's address SHALL be resolved the way routing resolves it: the node's declared engine override as given, otherwise the node's host with the port and -path the engine reports. Where a node reports its engine's host — a remote +path the engine reports. Where a node reports its engine's host — a cloud node environment, whose control plane publishes the instance's address and which the fleet file names by environment alone — that reported host SHALL be used in place of the node's host, so the request reaches the instance rather than an @@ -310,16 +310,16 @@ model to start with, matching the one the request asks for: the fleet file, resolved the way `spinloop fleet start` resolves it — describing a config whose model or served name is the one the request asks for. -- A remote node names one through its own stats reply, which reads the +- A cloud node names one through its own stats reply, which reads the environment's stored deploy config directly and so carries its model id whether the environment is running or stopped — unlike its status reply, which only relays the deploy config while running, and unlike the stats - reply itself, which carries no served name. An undeployed remote + reply itself, which carries no served name. An undeployed cloud node environment's stats read fails outright, having no deploy config to read; it names nothing and is not a candidate. A node is started with what it names, never with a config invented for the -request: a daemon node is started with the Spinloop source's config; a remote +request: a daemon node is started with the Spinloop source's config; a cloud node node is started as it is — its stored deploy config decides what it serves, and the gateway pushes it nothing new. Candidates whose stored config already names the model SHALL be tried first, since they have the weights, and the @@ -328,20 +328,20 @@ cannot serve — SHALL NOT fail the request while other candidates remain. A daemon engine started this way SHALL be gated with the key the node's fleet entry names, supplied by the gateway: the gateway is the client that starts -the engine, so the key the client sets is the key the engine takes. A remote +the engine, so the key the client sets is the key the engine takes. A cloud node environment's engine is gated by its own key, resolved the same way a request already routed to it resolves one; the gateway does not change it. The wait SHALL be bounded by a wake timeout, defaulting to five minutes and overridable by `--wake-timeout`; exceeding it SHALL fail the request saying the engine did not answer in time, and the started engine SHALL be left -running rather than stopped, so a slow load — or, for a remote node, a slow +running rather than stopped, so a slow load — or, for a cloud node, a slow boot — is not thrown away. When several requests ask for a model nothing is serving at once, the gateway SHALL start at most one engine per node and answer every request from it: the first request's wait is the wait the rest join, regardless of the node's kind. A daemon node's own control API refuses a second concurrent start on -its own, but a remote environment's control plane does not, so the gateway +its own, but a cloud environment's control plane does not, so the gateway SHALL NOT rely on that alone: two requests racing to wake the same node SHALL be coalesced before either reaches the node, not just reconciled after one of them answers. A node another request woke first SHALL be used the @@ -351,7 +351,7 @@ A request for a model nothing is serving, and for which waking is not allowed on any node that names it, SHALL fail without starting anything, naming the nodes and what they could serve, and the command that would start one. A model no node is running and no node names — no daemon source describes it -and no remote node is deployed with it — SHALL fail the same way regardless +and no cloud node is deployed with it — SHALL fail the same way regardless of any wake setting: nothing to wake with, and the failure SHALL say so rather than trying to start a node with nothing. @@ -362,18 +362,18 @@ rather than trying to start a node with nothing. - **THEN** that node is started with its own config, gated with the key its fleet entry names, and the request is answered once the engine answers -#### Scenario: A cold request wakes a deployed remote environment +#### Scenario: A cold request wakes a deployed cloud environment -- **WHEN** no node is running the model a request names, one remote node's +- **WHEN** no node is running the model a request names, one cloud node's stats reply reports it is deployed to serve it, and waking is allowed for it - **THEN** that environment's instance is started, its own stored deploy config decides what it serves, and the request is answered once its engine answers -#### Scenario: An undeployed remote node is not a wake candidate +#### Scenario: An undeployed cloud node is not a wake candidate -- **WHEN** the only node whose name could match a request is a remote +- **WHEN** the only node whose name could match a request is a cloud node environment with nothing deployed - **THEN** it is not started, and the failure says nothing is deployed to serve the model, naming the deployment path @@ -397,10 +397,10 @@ rather than trying to start a node with nothing. - **THEN** that node is started once, and both requests are answered from the same engine -#### Scenario: Concurrent cold requests share one remote wake +#### Scenario: Concurrent cold requests share one cloud node wake - **WHEN** two requests arrive at once for a model nothing is serving, and one - remote node's stats reply reports it is deployed to serve it + cloud node's stats reply reports it is deployed to serve it - **THEN** that environment's instance is started once, not once per request, and both requests are answered once its engine answers @@ -418,9 +418,9 @@ rather than trying to start a node with nothing. - **THEN** nothing is started, and the request fails naming the node whose source describes the model and the command that would start it -#### Scenario: A remote node opted out is not woken though the fleet wakes +#### Scenario: A cloud node opted out is not woken though the fleet wakes -- **WHEN** the fleet file's wake policy is `on`, a stopped remote node +- **WHEN** the fleet file's wake policy is `on`, a stopped cloud node declares its own `wake: off`, and it is the only node that names the model a request asks for - **THEN** it is not started, and the failure names it and says waking is @@ -429,7 +429,7 @@ rather than trying to start a node with nothing. #### Scenario: Nothing can serve the model - **WHEN** no node is running the model a request names, no node's Spinloop - source describes it, and no remote node is deployed with it + source describes it, and no cloud node is deployed with it - **THEN** the request fails, naming each node and why it cannot serve the model, and nothing is started @@ -499,7 +499,7 @@ serving facts — the model it serves when it is running, the name it serves that model under where it reports one, whether its engine has answered, and when it was last active. For a node that is not running, the reply SHALL name the model a request would start it with, where the node names one — a daemon -node's own source, or a remote node's own stats reply — and waking is +node's own source, or a cloud node's own stats reply — and waking is allowed for it (its own `wake` setting, or the fleet's when it names none); a node that names no such model, or for which waking is not allowed, SHALL report none. A node that does not answer SHALL be reported as such in the @@ -531,9 +531,9 @@ gateway holds no copy of either beyond what it already holds. - **THEN** the topology names that model as what a request would start the node with -#### Scenario: A stopped, deployed remote node reports what it would start +#### Scenario: A stopped, deployed cloud node reports what it would start -- **WHEN** a remote node is stopped, its stats reply reports its stored +- **WHEN** a cloud node is stopped, its stats reply reports its stored deploy config, and waking is allowed for it - **THEN** the topology names that config's model as what a request would start it with, the same way a daemon node's is named diff --git a/openspec/specs/fleet-routing/spec.md b/openspec/specs/fleet-routing/spec.md index 527db748..b374348c 100644 --- a/openspec/specs/fleet-routing/spec.md +++ b/openspec/specs/fleet-routing/spec.md @@ -337,7 +337,7 @@ of each node it considered. Routing SHALL resolve the engine key from the variable the node's fleet entry names, supply it to the node when it wakes one, and place it in the launched agent's environment as `OPENAI_API_KEY` — and, for a harness that reads the key -under its own name, under that name too, as the remote path already does. A key +under its own name, under that name too, as the cloud path already does. A key already set in spinloop's environment SHALL win. The client is therefore the one party that holds the key: it decides what the @@ -345,13 +345,13 @@ engine it starts is gated with, and it knows what to give the agent because it set it. A daemon node whose fleet entry names no key SHALL wake an ungated engine, which is correct for a node reached over loopback. -For a remote node the resolution is the same with a fleet-wide default: the key +For a cloud node the resolution is the same with a fleet-wide default: the key is the node's own `engineTokenEnv` when it names one, otherwise the fleet file's -`apiKeyEnv`. A remote's engine is always gated by its API key, so a remote that +`apiKeyEnv`. A cloud node's engine is always gated by its API key, so a cloud node that is selected — running or woken — SHALL be reached with a key. When neither the node nor the fleet names a variable, or the variable it names is set nowhere, routing SHALL fail before the agent launches, naming the node and what to set: a -remote is never reached ungated. +a cloud node is never reached ungated. Routing SHALL NOT ask the daemon for a key, and the daemon SHALL NOT return one: saying a key is required is a fact a router needs, and handing the key out is @@ -376,23 +376,23 @@ authenticate is worse than a message that says so. - **THEN** the engine starts ungated and the agent launches with no key injected for it -#### Scenario: A remote node takes the fleet-wide key +#### Scenario: A cloud node takes the fleet-wide key -- **WHEN** routing selects a remote node whose entry names no `engineTokenEnv`, +- **WHEN** routing selects a cloud node whose entry names no `engineTokenEnv`, and the fleet file declares an `apiKeyEnv` that is set - **THEN** the launched agent's environment carries that value as `OPENAI_API_KEY` -#### Scenario: A remote node's own key overrides the fleet-wide one +#### Scenario: A cloud node's own key overrides the fleet-wide one -- **WHEN** routing selects a remote node that names its own `engineTokenEnv` +- **WHEN** routing selects a cloud node that names its own `engineTokenEnv` and the fleet file also declares an `apiKeyEnv` - **THEN** the node's own variable is the key the agent is given, not the fleet-wide one -#### Scenario: A remote node with no key fails early +#### Scenario: A cloud node with no key fails early -- **WHEN** routing selects a remote node whose entry names no `engineTokenEnv` +- **WHEN** routing selects a cloud node whose entry names no `engineTokenEnv` and the fleet file declares no `apiKeyEnv` - **THEN** the command fails naming the node and what to set, and no agent is launched @@ -494,7 +494,7 @@ address is written, and SHALL also be placed in the launched agent's environment as `OPENAI_BASE_URL`. A variable already set in spinloop's environment SHALL win, as it does on the -remote path — routing fills what is unset, it does not override an explicit +cloud path — routing fills what is unset, it does not override an explicit choice. A Spinloop that pins a `BASEURL` SHALL NOT be routed: the pinned address wins @@ -612,7 +612,7 @@ choose from instead of an empty one. A failure to complete that query (the gateway unreachable, timed out, or answering something unusable) SHALL NOT fail the launch: it SHALL warn and configure the harness with an empty model list, on the same terms a launch already warns and carries on when it cannot refresh -a remote endpoint's key. This model-list population is a capability of +a cloud endpoint's key. This model-list population is a capability of harnesses whose config format holds more than one model per provider; a harness with no such concept is configured with no model or alias. @@ -621,7 +621,7 @@ not by the catalogue's shared generic id: its display name SHALL lead with "Gateway" — not the catalogue engine's own generic label — followed by the gateway's `name` where the fleet file's `gateway` section gives one, or its address otherwise (e.g. "Gateway (dev-2)" or "Gateway (localhost:4000)"), the -same "