From 8f5cbd246431122441b39497fefc7e3c96c4c168 Mon Sep 17 00:00:00 2001 From: spinloop-agent Date: Mon, 5 Oct 2026 20:13:14 +0100 Subject: [PATCH 1/4] docs(gateway): propose a configurable request body limit Refs #240 --- .../gateway-max-request-bytes/.openspec.yaml | 2 ++ .../gateway-max-request-bytes/design.md | 25 ++++++++++++++++ .../gateway-max-request-bytes/proposal.md | 25 ++++++++++++++++ .../specs/fleet-gateway/spec.md | 30 +++++++++++++++++++ .../gateway-max-request-bytes/tasks.md | 15 ++++++++++ 5 files changed, 97 insertions(+) create mode 100644 openspec/changes/gateway-max-request-bytes/.openspec.yaml create mode 100644 openspec/changes/gateway-max-request-bytes/design.md create mode 100644 openspec/changes/gateway-max-request-bytes/proposal.md create mode 100644 openspec/changes/gateway-max-request-bytes/specs/fleet-gateway/spec.md create mode 100644 openspec/changes/gateway-max-request-bytes/tasks.md diff --git a/openspec/changes/gateway-max-request-bytes/.openspec.yaml b/openspec/changes/gateway-max-request-bytes/.openspec.yaml new file mode 100644 index 00000000..e3966d7a --- /dev/null +++ b/openspec/changes/gateway-max-request-bytes/.openspec.yaml @@ -0,0 +1,2 @@ +schema: spec-driven +created: 2026-10-05 diff --git a/openspec/changes/gateway-max-request-bytes/design.md b/openspec/changes/gateway-max-request-bytes/design.md new file mode 100644 index 00000000..ed41a95e --- /dev/null +++ b/openspec/changes/gateway-max-request-bytes/design.md @@ -0,0 +1,25 @@ +## Context + +`requestModel` in `internal/gateway/gateway.go` reads the body with `io.ReadAll(io.LimitReader(r.Body, 1<<20))`, returns the model field and the bytes read, and the proxy forwards those bytes. The 1 MiB figure matches the daemon's control API, whose bodies are small. Gateway bodies carry whole conversations. + +## Goals / Non-Goals + +**Goals:** +- An over-limit body gets a `413` that names the limit and the flag. +- The limit is configurable and the default is above a realistic agent turn. +- Only a fully read body is ever forwarded. + +**Non-Goals:** +- Streaming bodies to engines without buffering them. +- Changing the daemon control API's limit. + +## Decisions + +- Use `http.MaxBytesReader(w, r.Body, limit)`. Its read error is `*http.MaxBytesError`, which `requestModel` returns as a typed error that the handler maps to `413`. `requestModel` takes the `ResponseWriter` and the limit as arguments. +- Any other read error is also returned and refused, so a short read never reaches the parse step or the proxy. +- Default 64 MiB (`DefaultMaxRequestBytes`), held in `Options.MaxRequestBytes`; zero means the default. The flag rejects values below 1 at startup. +- The `413` body uses the gateway's existing error shape (`type: gateway_error`) and names the limit in bytes and `--max-request-bytes`. + +## Risks / Trade-offs + +- The body is held in memory per request, so a higher limit raises the memory a burst of large requests can use. 64 MiB is a default an operator can lower. diff --git a/openspec/changes/gateway-max-request-bytes/proposal.md b/openspec/changes/gateway-max-request-bytes/proposal.md new file mode 100644 index 00000000..7921748b --- /dev/null +++ b/openspec/changes/gateway-max-request-bytes/proposal.md @@ -0,0 +1,25 @@ +## Why + +`spinloop gateway` reads at most 1 MiB of a completion request body with `io.LimitReader`, which cuts a larger body off without saying so. The cut-off JSON fails to parse and the caller gets `400 "the request is not a JSON body: unexpected end of JSON input"`, which points at the caller's JSON when the body was well formed. A long coding-agent session passes 1 MiB routinely, so the agent works for a while and then fails every turn. The same truncated buffer is also what the proxy forwards, so the only thing stopping a truncated prompt reaching an engine is the parse failure. + +## What Changes + +- The gateway reads the request body through `http.MaxBytesReader`, so a body over the limit is an error and not a silent truncation. +- A body over the limit is answered `413` with a message naming the limit and the `--max-request-bytes` flag. +- The default limit rises from 1 MiB to 64 MiB, and `spinloop gateway --max-request-bytes ` sets it. +- A body that was not read in full is never forwarded to an engine. +- `docs/commands/gateway.md` documents the flag and its default. + +## Capabilities + +### New Capabilities + +### Modified Capabilities +- `fleet-gateway`: adds a requirement for the request body limit, its `413` answer, and the flag that sets it. + +## Impact + +- `internal/gateway/gateway.go`: `requestModel`, `Options`, `Handler`. +- `cmd/spinloop/gateway.go`: the new flag, passed to the handler. +- `cmd/spinloop/complete.go`: completion of the new flag if flags are listed there. +- `docs/commands/gateway.md`, `openspec/specs/fleet-gateway/spec.md` (via the delta). diff --git a/openspec/changes/gateway-max-request-bytes/specs/fleet-gateway/spec.md b/openspec/changes/gateway-max-request-bytes/specs/fleet-gateway/spec.md new file mode 100644 index 00000000..c072e1c4 --- /dev/null +++ b/openspec/changes/gateway-max-request-bytes/specs/fleet-gateway/spec.md @@ -0,0 +1,30 @@ +## ADDED Requirements + +### Requirement: Request body limit + +The gateway SHALL read a completion request's body in full or refuse it: a +body larger than the limit SHALL be answered `413` with a message naming the +limit in bytes and the `--max-request-bytes` flag, and SHALL NOT be parsed or +forwarded. The limit SHALL default to 64 MiB and SHALL be set by +`--max-request-bytes`, which SHALL be a positive number of bytes. The gateway +SHALL forward to an engine only a body it read in full, byte for byte. + +#### Scenario: A body under the limit is routed + +- **WHEN** a completion request whose body is just under the limit reaches the gateway +- **THEN** it is routed and its full body is forwarded to the engine + +#### Scenario: A body over the limit is refused with 413 + +- **WHEN** a completion request whose body is larger than the limit reaches the gateway +- **THEN** the gateway answers `413` naming the limit and `--max-request-bytes`, not `400`, and no engine receives anything + +#### Scenario: The limit is set by a flag + +- **WHEN** the gateway is started with `--max-request-bytes 2097152` +- **THEN** a 1.5 MiB body is routed and a 3 MiB body is answered `413` naming 2097152 + +#### Scenario: A non-positive limit fails at startup + +- **WHEN** the gateway is started with `--max-request-bytes 0` +- **THEN** it fails at startup naming the flag, and nothing listens diff --git a/openspec/changes/gateway-max-request-bytes/tasks.md b/openspec/changes/gateway-max-request-bytes/tasks.md new file mode 100644 index 00000000..9f302661 --- /dev/null +++ b/openspec/changes/gateway-max-request-bytes/tasks.md @@ -0,0 +1,15 @@ +## 1. Gateway limit + +- [ ] 1.1 Add `DefaultMaxRequestBytes` and `Options.MaxRequestBytes` in `internal/gateway` +- [ ] 1.2 Read the body with `http.MaxBytesReader` in `requestModel`, returning a distinct error for an over-limit body and for any other short read +- [ ] 1.3 Answer an over-limit body `413` naming the limit and `--max-request-bytes`; never forward a body not read in full + +## 2. Command + +- [ ] 2.1 Add `--max-request-bytes` to `spinloop gateway`, reject values below 1 at startup, pass it to the handler +- [ ] 2.2 Update flag completion if the flag list lives in `cmd/spinloop/complete.go` + +## 3. Tests and docs + +- [ ] 3.1 Tests: just under the limit routes, over the limit gets 413 naming limit and flag, nothing reaches an engine, flag validation +- [ ] 3.2 Document the flag and default in `docs/commands/gateway.md` From 4e341b50cd146b42f20ff7c6df9cce05062053b0 Mon Sep 17 00:00:00 2001 From: spinloop-agent Date: Mon, 5 Oct 2026 20:15:17 +0100 Subject: [PATCH 2/4] fix(gateway): refuse over-limit request bodies with 413 The gateway cut completion bodies off at 1 MiB, so a larger body failed as malformed JSON. The body is now read through http.MaxBytesReader, an over-limit one is answered 413 naming --max-request-bytes, and the default rises to 64 MiB. Fixes #240 --- cmd/spinloop/gateway.go | 15 +++-- cmd/spinloop/gateway_test.go | 17 ++++-- cmd/spinloop/orchestrator_test.go | 8 +-- docs/commands/gateway.md | 1 + internal/gateway/gateway.go | 37 +++++++++-- internal/gateway/gateway_test.go | 61 +++++++++++++++++++ .../gateway-max-request-bytes/tasks.md | 14 ++--- 7 files changed, 126 insertions(+), 27 deletions(-) diff --git a/cmd/spinloop/gateway.go b/cmd/spinloop/gateway.go index 68106ae2..338a101a 100644 --- a/cmd/spinloop/gateway.go +++ b/cmd/spinloop/gateway.go @@ -27,6 +27,7 @@ import ( func gatewayCmd() *cobra.Command { var fleetPath, listen, apiToken, apiTokenFile string var wakeTimeout time.Duration + var maxRequestBytes int64 var loopback bool c := &cobra.Command{ Use: "gateway", @@ -51,7 +52,7 @@ The agent then needs only the gateway's token, as OPENAI_API_KEY.`, SilenceUsage: true, RunE: func(c *cobra.Command, args []string) error { resolve(c) - return runGatewayCommand(fleetPath, listen, apiToken, apiTokenFile, wakeTimeout, loopback, c.Flags()) + return runGatewayCommand(fleetPath, listen, apiToken, apiTokenFile, wakeTimeout, maxRequestBytes, loopback, c.Flags()) }, } fs := c.Flags() @@ -61,6 +62,7 @@ The agent then needs only the gateway's token, as OPENAI_API_KEY.`, fs.StringVar(&apiTokenFile, "api-token-file", "", "read the gateway's bearer token from this file") fs.StringVar(&apiToken, "api-token", "", "the gateway's bearer token") fs.DurationVar(&wakeTimeout, "wake-timeout", 0, "how long to wait for a woken engine to answer") + fs.Int64Var(&maxRequestBytes, "max-request-bytes", gateway.DefaultMaxRequestBytes, "the largest completion request body to accept, in bytes") compRegister(c, "fleet", compFiles) return c } @@ -70,7 +72,7 @@ func cmdGateway(args []string) error { return execCmd(gatewayCmd(), args) } // runGatewayCommand is the body of `spinloop gateway`: the server, and the // signal handling that shuts it down cleanly. -func runGatewayCommand(fleetPath, listen, apiToken, apiTokenFile string, wakeTimeout time.Duration, loopback bool, flags *pflag.FlagSet) error { +func runGatewayCommand(fleetPath, listen, apiToken, apiTokenFile string, wakeTimeout time.Duration, maxRequestBytes int64, loopback bool, flags *pflag.FlagSet) error { // Whether --listen was typed at all, not whether it differs from the // default: --listen :4000 --loopback is still a conflict, and a // compare-against-default check would let it pass. @@ -88,7 +90,10 @@ func runGatewayCommand(fleetPath, listen, apiToken, apiTokenFile string, wakeTim restore = func() { fleet.WakeTimeout = prev } defer restore() } - srv, ln, err := newGatewayServer(fleetPath, listen, apiToken, apiTokenFile) + if maxRequestBytes < 1 { + return fmt.Errorf("--max-request-bytes must be at least 1, got %d", maxRequestBytes) + } + srv, ln, err := newGatewayServer(fleetPath, listen, apiToken, apiTokenFile, maxRequestBytes) if err != nil { return err } @@ -128,7 +133,7 @@ func gatewayListenAddr(listen string, listenExplicit, loopback bool) (string, er // file's token references the way a startup must, opens the listener, and // prints the address a fleet file's gateway section names. Everything that can // fail without serving fails here, before a listener exists. -func newGatewayServer(fleetPath, listen, apiToken, apiTokenFile string) (*http.Server, net.Listener, error) { +func newGatewayServer(fleetPath, listen, apiToken, apiTokenFile string, maxRequestBytes int64) (*http.Server, net.Listener, error) { cfg, err := fleet.Resolve(fleetPath) if err != nil { return nil, nil, err @@ -176,7 +181,7 @@ func newGatewayServer(fleetPath, listen, apiToken, apiTokenFile string) (*http.S return deployConfigForNode(sel, path) } - h := gateway.New(cfg, token, gateway.Options{ConfigFor: cfgFor, Log: logger}) + h := gateway.New(cfg, token, gateway.Options{ConfigFor: cfgFor, Log: logger, MaxRequestBytes: maxRequestBytes}) ln, err := gateway.Listen(listen, token) if err != nil { return nil, nil, err diff --git a/cmd/spinloop/gateway_test.go b/cmd/spinloop/gateway_test.go index a1efda1f..da1d9cea 100644 --- a/cmd/spinloop/gateway_test.go +++ b/cmd/spinloop/gateway_test.go @@ -28,7 +28,7 @@ func TestGatewayStartsAndAnswers(t *testing.T) { var ln net.Listener out := captureStdout(t, func() { var err error - srv, ln, err = newGatewayServer("", "127.0.0.1:0", "", "") + srv, ln, err = newGatewayServer("", "127.0.0.1:0", "", "", 0) if err != nil { t.Fatal(err) } @@ -68,7 +68,7 @@ func TestGatewayStartsAndAnswers(t *testing.T) { // A missing fleet file fails naming the expected path, and nothing listens. func TestGatewayFailsWithoutAFleetFile(t *testing.T) { t.Chdir(t.TempDir()) - _, ln, err := newGatewayServer("", "127.0.0.1:0", "", "") + _, ln, err := newGatewayServer("", "127.0.0.1:0", "", "", 0) if err == nil { t.Fatal("a gateway with no fleet file should fail") } @@ -87,7 +87,7 @@ func TestGatewayFailsOnAnUnsetTokenVariable(t *testing.T) { fleetFileIn(t, dir, "nodes:\n - name: gated\n host: 127.0.0.1\n port: 14242\n tokenEnv: GW_NODE_TOKEN_UNSET\n") t.Chdir(dir) - _, ln, err := newGatewayServer("", "127.0.0.1:0", "", "") + _, ln, err := newGatewayServer("", "127.0.0.1:0", "", "", 0) if err == nil { t.Fatal("an unset token variable should fail the gateway at startup") } @@ -109,7 +109,7 @@ func TestGatewayTokenSourcesConflict(t *testing.T) { tokenFile := filepath.Join(t.TempDir(), "token") mustWrite(t, tokenFile, "from-file\n") - _, _, err := newGatewayServer("", "127.0.0.1:0", "literal", tokenFile) + _, _, err := newGatewayServer("", "127.0.0.1:0", "literal", tokenFile, 0) if err == nil { t.Fatal("two token sources should be a conflict") } @@ -126,7 +126,7 @@ func TestGatewayRefusesTokenlessNonLoopback(t *testing.T) { t.Chdir(dir) t.Setenv("SPINLOOP_API_TOKEN", "") - _, ln, err := newGatewayServer("", "0.0.0.0:0", "", "") + _, ln, err := newGatewayServer("", "0.0.0.0:0", "", "", 0) if err == nil { t.Fatal("a tokenless non-loopback gateway should refuse to start") } @@ -274,3 +274,10 @@ func TestFleetURL(t *testing.T) { } } } + +func TestGatewayRefusesANonPositiveRequestLimit(t *testing.T) { + err := cmdGateway([]string{"--max-request-bytes", "0"}) + if err == nil || !strings.Contains(err.Error(), "--max-request-bytes") { + t.Fatalf("got %v, want an error naming --max-request-bytes", err) + } +} diff --git a/cmd/spinloop/orchestrator_test.go b/cmd/spinloop/orchestrator_test.go index a7137de2..c64695be 100644 --- a/cmd/spinloop/orchestrator_test.go +++ b/cmd/spinloop/orchestrator_test.go @@ -117,7 +117,7 @@ func TestCmdOrchestrator_AFleetFileNamesTheGateway(t *testing.T) { mustWrite(t, filepath.Join(dir, ".env"), "ORCH_FILE_TOKEN=the-token\n") mustWrite(t, "work.yaml", "- id: a\n instructions: do\n dir: .\n tags:\n - gpu=a100\n") - srv, ln, err := newGatewayServer("", "127.0.0.1:0", "", "") + srv, ln, err := newGatewayServer("", "127.0.0.1:0", "", "", 0) if err != nil { t.Fatal(err) } @@ -266,7 +266,7 @@ func TestCmdOrchestrator_WorksAgainstAGatewayOnLoopback(t *testing.T) { t.Chdir(dir) mustWrite(t, "work.yaml", "- id: a\n instructions: do\n dir: .\n tags:\n - gpu=a100\n") - srv, ln, err := newGatewayServer("", "127.0.0.1:0", "", "") + srv, ln, err := newGatewayServer("", "127.0.0.1:0", "", "", 0) if err != nil { t.Fatal(err) } @@ -399,7 +399,7 @@ func TestCmdOrchestrator_LoopbackServesTheWorkListWithoutAToken(t *testing.T) { // The item names a tag no node carries, so nothing is launched. mustWrite(t, "work.yaml", "- id: a\n instructions: do\n dir: .\n tags:\n - gpu=a100\n") - srv, ln, err := newGatewayServer("", "127.0.0.1:0", "", "") + srv, ln, err := newGatewayServer("", "127.0.0.1:0", "", "", 0) if err != nil { t.Fatal(err) } @@ -495,7 +495,7 @@ func TestCmdOrchestrator_StartupShowsARestartsRecoveredState(t *testing.T) { t.Fatal(err) } - srv, ln, err := newGatewayServer("", "127.0.0.1:0", "", "") + srv, ln, err := newGatewayServer("", "127.0.0.1:0", "", "", 0) if err != nil { t.Fatal(err) } diff --git a/docs/commands/gateway.md b/docs/commands/gateway.md index 9f11e055..eb298949 100644 --- a/docs/commands/gateway.md +++ b/docs/commands/gateway.md @@ -196,6 +196,7 @@ on a shared machine wants, and the reason the token is not optional there. | `--api-token-file` | Read the gateway's bearer token from this file | | `--api-token` | The gateway's bearer token | | `--wake-timeout` | How long to wait for a woken engine to answer (default 5m) | +| `--max-request-bytes` | The largest completion request body to accept, in bytes (default 67108864, 64 MiB); a larger one is answered `413` | ## See also diff --git a/internal/gateway/gateway.go b/internal/gateway/gateway.go index ade0c016..fcb0bd3f 100644 --- a/internal/gateway/gateway.go +++ b/internal/gateway/gateway.go @@ -36,6 +36,11 @@ import ( // one address without knowing the machine it lands on. const DefaultListen = ":4000" +// DefaultMaxRequestBytes is the largest completion request body the gateway +// reads when --max-request-bytes is not given. A request carries a whole +// conversation, so the figure sits well above a long agent turn. +const DefaultMaxRequestBytes int64 = 64 << 20 + // LoopbackListen is where `--loopback` binds the gateway: the default port on // loopback, the safe bind a local-only gateway wants — one that Listen's token // check accepts without a token. @@ -68,6 +73,10 @@ type Options struct { Log *slog.Logger // Now is the clock the reading cache ages against; nil uses time.Now. Now func() time.Time + // MaxRequestBytes is the largest completion request body the gateway + // accepts; a larger one is answered 413. Zero or less uses + // DefaultMaxRequestBytes. + MaxRequestBytes int64 } // Handler is the gateway: the fleet it serves, the token its callers present, @@ -80,6 +89,8 @@ type Handler struct { log *slog.Logger now func() time.Time + maxRequestBytes int64 + mu sync.Mutex results []fleet.NodeResult at time.Time @@ -99,7 +110,11 @@ func New(cfg *fleet.Config, token string, opts Options) *Handler { if now == nil { now = time.Now } - return &Handler{cfg: cfg, token: token, cfgFor: opts.ConfigFor, log: log, now: now} + maxBytes := opts.MaxRequestBytes + if maxBytes <= 0 { + maxBytes = DefaultMaxRequestBytes + } + return &Handler{cfg: cfg, token: token, cfgFor: opts.ConfigFor, log: log, now: now, maxRequestBytes: maxBytes} } // Listen opens the gateway's listener, applying the daemon's exposure rule: @@ -430,8 +445,15 @@ func (h *Handler) reading(ctx context.Context) []fleet.NodeResult { // handleCompletion routes a completion request to the node serving its model, // waking one when nothing is and the fleet file allows it. func (h *Handler) handleCompletion(w http.ResponseWriter, r *http.Request) { - model, body, err := requestModel(r) + model, body, err := requestModel(w, r, h.maxRequestBytes) if err != nil { + var tooLarge *http.MaxBytesError + if errors.As(err, &tooLarge) { + writeError(w, http.StatusRequestEntityTooLarge, fmt.Errorf( + "the request body is larger than the gateway's limit of %d bytes: raise it with --max-request-bytes", + tooLarge.Limit)) + return + } writeError(w, http.StatusBadRequest, err) return } @@ -471,10 +493,13 @@ func (h *Handler) handleCompletion(w http.ResponseWriter, r *http.Request) { } // requestModel pulls the model field out of a completion request and returns -// it with the full body, which the proxy must forward unmodified. A body that -// is not a JSON object fails saying so, rather than being routed at a guess. -func requestModel(r *http.Request) (string, []byte, error) { - body, err := io.ReadAll(io.LimitReader(r.Body, 1<<20)) +// it with the full body, which the proxy must forward unmodified. A body +// larger than limit returns an error wrapping *http.MaxBytesError, and any +// other failed read returns an error, so a body that was not read in full is +// never returned. A body that is not a JSON object fails saying so, rather +// than being routed at a guess. +func requestModel(w http.ResponseWriter, r *http.Request, limit int64) (string, []byte, error) { + body, err := io.ReadAll(http.MaxBytesReader(w, r.Body, limit)) if err != nil { return "", nil, fmt.Errorf("reading the request: %w", err) } diff --git a/internal/gateway/gateway_test.go b/internal/gateway/gateway_test.go index 94994272..65b3baac 100644 --- a/internal/gateway/gateway_test.go +++ b/internal/gateway/gateway_test.go @@ -86,6 +86,8 @@ type fakeNode struct { pushedKey string // engineGotAuth is the last authorisation the engine itself saw. engineGotAuth string + // engineGotBody is the last request body the engine itself read. + engineGotBody string // statusHits counts status calls, so a burst's fan-out is countable. statusHits int @@ -135,6 +137,9 @@ func (f *fakeNode) engineHandler() http.Handler { return } body, _ := io.ReadAll(r.Body) + f.mu.Lock() + f.engineGotBody = string(body) + f.mu.Unlock() switch { case strings.Contains(string(body), `"stream":true`): w.Header().Set("Content-Type", "text/event-stream") @@ -1476,3 +1481,59 @@ func TestStartingEngineThatNeverAnswersFailsNamingTheNode(t *testing.T) { t.Errorf("the caller should not be given a dial error: %s", body) } } + +// paddedRequest is a valid completion request of exactly size bytes. +func paddedRequest(size int) string { + const head = `{"model":"org/wanted","messages":[{"role":"user","content":"` + const tail = `"}]}` + return head + strings.Repeat("a", size-len(head)-len(tail)) + tail +} + +func TestRequestAtTheLimitIsForwardedInFull(t *testing.T) { + node := newFakeNode(t, string(daemon.StateRunning), "org/wanted") + h := New(fleetOf(t, []string{"box"}, node), "", Options{MaxRequestBytes: 4096}) + req := paddedRequest(4096) + + resp, body := post(t, h, "", req) + if resp.StatusCode != http.StatusOK { + t.Fatalf("HTTP %d, body %s", resp.StatusCode, body) + } + node.mu.Lock() + defer node.mu.Unlock() + if node.engineGotBody != req { + t.Errorf("the engine read %d bytes, want the full %d", len(node.engineGotBody), len(req)) + } +} + +func TestRequestOverTheLimitIsRefusedWith413(t *testing.T) { + node := newFakeNode(t, string(daemon.StateRunning), "org/wanted") + h := New(fleetOf(t, []string{"box"}, node), "", Options{MaxRequestBytes: 4096}) + + resp, body := post(t, h, "", paddedRequest(4097)) + if resp.StatusCode != http.StatusRequestEntityTooLarge { + t.Fatalf("HTTP %d, want 413; body %s", resp.StatusCode, body) + } + for _, want := range []string{"4096", "--max-request-bytes"} { + if !strings.Contains(body, want) { + t.Errorf("the refusal should name %q: %s", want, body) + } + } + if strings.Contains(body, "not a JSON body") { + t.Errorf("an over-limit body must not read as malformed JSON: %s", body) + } + node.mu.Lock() + defer node.mu.Unlock() + if node.statusHits != 0 || node.engineGotBody != "" { + t.Error("an over-limit request reached a node") + } +} + +func TestDefaultRequestLimitAdmitsWhatOneMiBRefused(t *testing.T) { + node := newFakeNode(t, string(daemon.StateRunning), "org/wanted") + h := New(fleetOf(t, []string{"box"}, node), "", Options{}) + + resp, body := post(t, h, "", paddedRequest(2<<20)) + if resp.StatusCode != http.StatusOK { + t.Fatalf("HTTP %d, body %s", resp.StatusCode, body) + } +} diff --git a/openspec/changes/gateway-max-request-bytes/tasks.md b/openspec/changes/gateway-max-request-bytes/tasks.md index 9f302661..7d78d75a 100644 --- a/openspec/changes/gateway-max-request-bytes/tasks.md +++ b/openspec/changes/gateway-max-request-bytes/tasks.md @@ -1,15 +1,15 @@ ## 1. Gateway limit -- [ ] 1.1 Add `DefaultMaxRequestBytes` and `Options.MaxRequestBytes` in `internal/gateway` -- [ ] 1.2 Read the body with `http.MaxBytesReader` in `requestModel`, returning a distinct error for an over-limit body and for any other short read -- [ ] 1.3 Answer an over-limit body `413` naming the limit and `--max-request-bytes`; never forward a body not read in full +- [x] 1.1 Add `DefaultMaxRequestBytes` and `Options.MaxRequestBytes` in `internal/gateway` +- [x] 1.2 Read the body with `http.MaxBytesReader` in `requestModel`, returning a distinct error for an over-limit body and for any other short read +- [x] 1.3 Answer an over-limit body `413` naming the limit and `--max-request-bytes`; never forward a body not read in full ## 2. Command -- [ ] 2.1 Add `--max-request-bytes` to `spinloop gateway`, reject values below 1 at startup, pass it to the handler -- [ ] 2.2 Update flag completion if the flag list lives in `cmd/spinloop/complete.go` +- [x] 2.1 Add `--max-request-bytes` to `spinloop gateway`, reject values below 1 at startup, pass it to the handler +- [x] 2.2 Flag completion needs no change: the gateway's flags are completed from its Cobra flag set ## 3. Tests and docs -- [ ] 3.1 Tests: just under the limit routes, over the limit gets 413 naming limit and flag, nothing reaches an engine, flag validation -- [ ] 3.2 Document the flag and default in `docs/commands/gateway.md` +- [x] 3.1 Tests: just under the limit routes, over the limit gets 413 naming limit and flag, nothing reaches an engine, flag validation +- [x] 3.2 Document the flag and default in `docs/commands/gateway.md` From c780f6f9f095cc02d6be970f9f515593a2dce51d Mon Sep 17 00:00:00 2001 From: spinloop-agent Date: Mon, 5 Oct 2026 20:17:20 +0100 Subject: [PATCH 3/4] test(gateway): cover partial reads and document the request limit --- docs/commands/gateway.md | 7 ++++++ internal/gateway/gateway_test.go | 41 ++++++++++++++++++++++++++++++++ 2 files changed, 48 insertions(+) diff --git a/docs/commands/gateway.md b/docs/commands/gateway.md index eb298949..d363f619 100644 --- a/docs/commands/gateway.md +++ b/docs/commands/gateway.md @@ -167,6 +167,13 @@ The gateway needs the same environment a machine running `spinloop fleet start` would: the tokens the fleet file names, set in its process environment or in a `.env` beside the fleet file. +### Request size + +A completion request body is read in full or refused. The limit is 64 MiB by +default, set with `--max-request-bytes`. A larger body is answered `413`, naming +the limit and the flag, and nothing is sent to an engine. A long-context model +fed whole conversations may need the limit raised. + ## The gateway's token Callers present the gateway's token as a bearer token on every request — a diff --git a/internal/gateway/gateway_test.go b/internal/gateway/gateway_test.go index 65b3baac..b3208908 100644 --- a/internal/gateway/gateway_test.go +++ b/internal/gateway/gateway_test.go @@ -1537,3 +1537,44 @@ func TestDefaultRequestLimitAdmitsWhatOneMiBRefused(t *testing.T) { t.Fatalf("HTTP %d, body %s", resp.StatusCode, body) } } + +// failingBody yields some bytes and then a read error, the way a client that +// drops its connection part-way through a body does. +type failingBody struct{ sent bool } + +func (b *failingBody) Read(p []byte) (int, error) { + if b.sent { + return 0, fmt.Errorf("connection reset") + } + b.sent = true + return copy(p, `{"model":"org/wanted","messages":[`), nil +} + +func TestPartiallyReadRequestIsNeverForwarded(t *testing.T) { + node := newFakeNode(t, string(daemon.StateRunning), "org/wanted") + h := New(fleetOf(t, []string{"box"}, node), "", Options{}) + req := httptest.NewRequest(http.MethodPost, "http://gw/v1/chat/completions", &failingBody{}) + rec := httptest.NewRecorder() + h.ServeHTTP(rec, req) + + if rec.Code != http.StatusBadRequest { + t.Fatalf("HTTP %d, want 400", rec.Code) + } + if !strings.Contains(rec.Body.String(), "reading the request") { + t.Errorf("the refusal should say the read failed: %s", rec.Body.String()) + } + node.mu.Lock() + defer node.mu.Unlock() + if node.statusHits != 0 || node.engineGotBody != "" { + t.Error("a partly read request reached a node") + } +} + +func TestMalformedJSONStillReadsAsMalformed(t *testing.T) { + node := newFakeNode(t, string(daemon.StateRunning), "org/wanted") + h := New(fleetOf(t, []string{"box"}, node), "", Options{}) + resp, body := post(t, h, "", `{"model":`) + if resp.StatusCode != http.StatusBadRequest || !strings.Contains(body, "not a JSON body") { + t.Fatalf("HTTP %d, body %s; want 400 naming malformed JSON", resp.StatusCode, body) + } +} From f425fca9cfd38ed2d3d8406898532ce76d65a414 Mon Sep 17 00:00:00 2001 From: spinloop-agent Date: Mon, 5 Oct 2026 22:50:30 +0100 Subject: [PATCH 4/4] docs(gateway): archive the request limit change and sync its spec --- .../.openspec.yaml | 0 .../design.md | 0 .../proposal.md | 0 .../specs/fleet-gateway/spec.md | 0 .../tasks.md | 0 openspec/specs/fleet-gateway/spec.md | 29 +++++++++++++++++++ 6 files changed, 29 insertions(+) rename openspec/changes/{gateway-max-request-bytes => archive/2026-10-05-gateway-max-request-bytes}/.openspec.yaml (100%) rename openspec/changes/{gateway-max-request-bytes => archive/2026-10-05-gateway-max-request-bytes}/design.md (100%) rename openspec/changes/{gateway-max-request-bytes => archive/2026-10-05-gateway-max-request-bytes}/proposal.md (100%) rename openspec/changes/{gateway-max-request-bytes => archive/2026-10-05-gateway-max-request-bytes}/specs/fleet-gateway/spec.md (100%) rename openspec/changes/{gateway-max-request-bytes => archive/2026-10-05-gateway-max-request-bytes}/tasks.md (100%) diff --git a/openspec/changes/gateway-max-request-bytes/.openspec.yaml b/openspec/changes/archive/2026-10-05-gateway-max-request-bytes/.openspec.yaml similarity index 100% rename from openspec/changes/gateway-max-request-bytes/.openspec.yaml rename to openspec/changes/archive/2026-10-05-gateway-max-request-bytes/.openspec.yaml diff --git a/openspec/changes/gateway-max-request-bytes/design.md b/openspec/changes/archive/2026-10-05-gateway-max-request-bytes/design.md similarity index 100% rename from openspec/changes/gateway-max-request-bytes/design.md rename to openspec/changes/archive/2026-10-05-gateway-max-request-bytes/design.md diff --git a/openspec/changes/gateway-max-request-bytes/proposal.md b/openspec/changes/archive/2026-10-05-gateway-max-request-bytes/proposal.md similarity index 100% rename from openspec/changes/gateway-max-request-bytes/proposal.md rename to openspec/changes/archive/2026-10-05-gateway-max-request-bytes/proposal.md diff --git a/openspec/changes/gateway-max-request-bytes/specs/fleet-gateway/spec.md b/openspec/changes/archive/2026-10-05-gateway-max-request-bytes/specs/fleet-gateway/spec.md similarity index 100% rename from openspec/changes/gateway-max-request-bytes/specs/fleet-gateway/spec.md rename to openspec/changes/archive/2026-10-05-gateway-max-request-bytes/specs/fleet-gateway/spec.md diff --git a/openspec/changes/gateway-max-request-bytes/tasks.md b/openspec/changes/archive/2026-10-05-gateway-max-request-bytes/tasks.md similarity index 100% rename from openspec/changes/gateway-max-request-bytes/tasks.md rename to openspec/changes/archive/2026-10-05-gateway-max-request-bytes/tasks.md diff --git a/openspec/specs/fleet-gateway/spec.md b/openspec/specs/fleet-gateway/spec.md index b8635e66..fd4bfc57 100644 --- a/openspec/specs/fleet-gateway/spec.md +++ b/openspec/specs/fleet-gateway/spec.md @@ -440,6 +440,35 @@ rather than trying to start a node with nothing. - **THEN** it is not started, and the failure names it and the ways a source could have been given +### Requirement: Request body limit + +The gateway SHALL read a completion request's body in full or refuse it: a +body larger than the limit SHALL be answered `413` with a message naming the +limit in bytes and the `--max-request-bytes` flag, and SHALL NOT be parsed or +forwarded. The limit SHALL default to 64 MiB and SHALL be set by +`--max-request-bytes`, which SHALL be a positive number of bytes. The gateway +SHALL forward to an engine only a body it read in full, byte for byte. + +#### Scenario: A body under the limit is routed + +- **WHEN** a completion request whose body is just under the limit reaches the gateway +- **THEN** it is routed and its full body is forwarded to the engine + +#### Scenario: A body over the limit is refused with 413 + +- **WHEN** a completion request whose body is larger than the limit reaches the gateway +- **THEN** the gateway answers `413` naming the limit and `--max-request-bytes`, not `400`, and no engine receives anything + +#### Scenario: The limit is set by a flag + +- **WHEN** the gateway is started with `--max-request-bytes 2097152` +- **THEN** a 1.5 MiB body is routed and a 3 MiB body is answered `413` naming 2097152 + +#### Scenario: A non-positive limit fails at startup + +- **WHEN** the gateway is started with `--max-request-bytes 0` +- **THEN** it fails at startup naming the flag, and nothing listens + ### Requirement: Paths the gateway does not serve A path other than `/v1/models`, `/v1/chat/completions`, `/v1/completions`, and