Skip to content

bug(relay): target-side close never propagates through forward service, so keep-alive reuse hangs forever #3724

Description

@Marty333

Agent Diagnostic

  • Investigated directly from source (clone of this repo at v0.0.106, cross-checked against v0.1.0 and main). The repo's debug skills were not loaded; the diagnosis comes from live socket inspection and the gateway log.
  • Tested OpenShell v0.0.106 on macOS with the Docker driver. Did not re-run on v0.1.0, but the relevant code is unchanged there (line numbers below are from v0.1.0).
  • Searched existing issues. Closest match is feat(api)!: define lifecycle semantics for bidirectional streams #3056, which proposes explicit close/half-close frames as a design change. This issue is a concrete, reproducible bug in the current EOF handling, with a small fix.
  • Found: when the target closes a connection, the supervisor relay never propagates the close back. The client's local socket stays ESTABLISHED indefinitely, and the next request on it hangs.
  • Workaround that works: forward over ssh -L via openshell ssh-proxy instead of openshell forward service.

Description

A browser dashboard served through openshell forward service loads fine, then hangs on the next click after a few seconds idle.

Cause: the dashboard (a Python server) closes idle keep-alive connections after ~5 s, which is normal. That close never reaches the client side of the forward. The browser reuses what it believes is a live keep-alive connection, writes a request into a dead relay, and waits forever.

Expected: when the target closes, the client's local socket is closed (EOF), so HTTP clients reconnect transparently.

Root cause: crates/openshell-supervisor-process/src/supervisor_session.rs, handle_relay_open (v0.1.0 lines 632–765):

let out_tx_writer = out_tx.clone();                 // L701
let target_to_grpc = tokio::spawn(async move {
    loop {
        match target_r.read(&mut buf).await {
            Ok(0) | Err(_) => break,                // target EOF: drops only the *clone*
            ...
while let Some(next) = inbound.next().await { ... } // main task still holds `out_tx`, blocked here
...
drop(out_tx);                                       // L759: only reached after the *client* side ends

On target EOF, target_to_grpc exits and drops out_tx_writer. But the original out_tx stays alive in the main task, which is blocked waiting on inbound. So the outbound RelayStream never ends, the gateway's bridge_forward_tcp_stream never reads EOF from the relay, and the CLI's forward_one_tcp_connection never ends its response stream and never closes the local socket.

Suggested fix (untested): move out_tx into the target→gRPC task instead of cloning it, so target EOF closes the outbound stream:

-    let out_tx_writer = out_tx.clone();
+    let out_tx_writer = out_tx;
 ...
-    drop(out_tx);
     let _ = target_to_grpc.await;

The rest of the chain appears to handle this correctly. The gateway bridge returns on relay EOF and drops tx. The CLI then calls local_write.shutdown() and aborts its reader, which ends the inbound stream, so the supervisor's inbound loop completes normally.

Related, secondary: MAX_CONNECTIONS_PER_SANDBOX = 20 (crates/openshell-server/src/grpc/sandbox.rs, acquire_ssh_connection_slots) applies to every ForwardTcp connection. It's hard-coded, and a single browser tab uses ~6. Over the cap, ForwardTcp returns RESOURCE_EXHAUSTED and the CLI silently drops the socket. The gateway logs this only at INFO (rpc.grpc.status_code=8). The bug above makes this worse, because zombie connections keep holding slots. Consider making it configurable. Happy to split this into its own issue.

Reproduction Steps

  1. Create a sandbox running any HTTP/1.1 server that closes idle keep-alive connections after a few seconds (e.g. uvicorn's default timeout_keep_alive=5), listening on port P.
  2. openshell forward service <sandbox> --target-port P --target-host 127.0.0.1 --local 127.0.0.1:P
  3. On the host:
    import http.client, time
    c = http.client.HTTPConnection("127.0.0.1", P, timeout=10)
    c.request("GET", "/"); c.getresponse().read()   # 200
    time.sleep(10)                                   # server closes idle conn
    c.request("GET", "/"); c.getresponse().read()   # hangs -> TimeoutError
  4. While it hangs: the host lsof shows the forward's socket ESTABLISHED. Inside the sandbox netns, ss -tan shows the supervisor's socket to the target in CLOSE-WAIT (the target sent FIN, and the supervisor never closed its end).

Same test through ssh -L over openshell ssh-proxy: the second request fails immediately with RemoteDisconnected, which is correct EOF, and clients retry.

Path Reuse after 10 s idle 3×30 parallel GETs
openshell forward service hangs 80/90 (10× RESOURCE_EXHAUSTED)
ssh -L via openshell ssh-proxy clean EOF 90/90
inside sandbox, direct to target n/a 90/90

Environment

Logs

# Host, while the client hangs: forward still holds the connection
openshell 15704 127.0.0.1:18789->127.0.0.1:51292 (ESTABLISHED)

# Sandbox netns: the target has closed, the supervisor never closed its end
CLOSE-WAIT 0 0 127.0.0.1:36508 127.0.0.1:18789
TIME-WAIT  0 0 127.0.0.1:48506 127.0.0.1:19119   # dashboard conns all closed

# Gateway log, parallel burst (secondary cap), status only at INFO
INFO request{path="/openshell.v1.OpenShell/ForwardTcp" rpc.grpc.status_code=8 otel.status_code="ERROR"}: openshell_server::multiplex: response status=200 latency_ms=32

Agent-First Checklist

  • I pointed my agent at the repo and had it investigate this issue
  • I loaded relevant skills — investigated from source and live sockets instead (see diagnostic)
  • I checked the latest OpenShell release and either reproduced the issue there or explained why I cannot upgrade/test it
  • I searched existing issues for possible duplicates or explained why I could not
  • My agent could not resolve this — the diagnostic above explains why

🤖 Generated with Claude Code

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    state:triage-neededOpened without agent diagnostics and needs triage

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions