Skip to content

Bug: Reaper reports success while the session filter never reaches Ryuk (silent leak on Linux/docker-proxy) #1114

Description

@Vaganovski

Describe the bug

Reaper._create_instance() can return a fully "successful" reaper whose session filter never reached Ryuk. Nothing raises, Reaper._instance is set, the ryuk container is running — and yet nothing is ever reaped. When the test process is killed (CI timeout, cancelled job, OOM), every container of that session survives.

Three things combine (testcontainers/core/container.py, 4.14.2):

  1. The wait strategy is registered after the container is started, so it never applies.
Reaper._container = (
    DockerContainer(c.ryuk_image)
    ...
    .start()          # <- started here
)
rc = Reaper._container
rc.waiting_for(LogMessageWaitStrategy(r".* Started!").with_startup_timeout(20))   # <- too late

waiting_for is a builder setter — its own docstring says "Set a wait strategy to be used after container start" — so setting it on an already-started container has no effect. There is no wait for Started! in practice.

  1. docker-proxy accepts on the published port from container creation, before the ryuk process inside has bound to it. The connect loop only retries on ConnectionRefusedError/OSError, so a connection that is accepted by the proxy and then reset is indistinguishable from success:
for _ in range(50):
    try:
        s.connect((container_host, container_port))
        last_connection_exception = None
        break                      # first attempt succeeds against the proxy
    except (ConnectionRefusedError, OSError) as e:
        ...
  1. The filter is sent and the reply is never read. Ryuk answers ACK for every accepted filter line; nothing checks for it:
rs.send(f"label={LABEL_SESSION_ID}={SESSION_ID}\r\n".encode())
Reaper._instance = Reaper()

So the filter goes into a socket that is not (yet) the reaper, the peer resets it, and the library reports success.

Observed result

Ryuk's own log tells the story — no client ever registered:

Pinging Docker...
Docker daemon is available!
Starting on port 8080...
Started!
Timeout waiting for connection
Removed 0 container(s), 0 network(s), 0 volume(s), 0 image(s)

There is no New client connected line, and after the run is killed its containers stay up indefinitely. This is how a CI runner accumulated 40 leftover containers, including a full set of live storages still running hours after the run that created them had ended.

Why this is easy to miss

The failure is silent and platform-dependent. On macOS/Docker Desktop the published port does not accept connections before the process inside binds, so the library wins the race and everything looks correct. On a Linux daemon with docker-proxy it loses.

Measured: 5 out of 5 runs affected on a self-hosted Linux CI runner (Ubuntu in WSL2, Docker CE 29.1.3), 0 out of 5 on macOS with Docker Desktop, same library version and same code.

To Reproduce

On a Linux host with Docker CE:

from testcontainers.core.container import Reaper

Reaper.get_instance()
print("instance:", Reaper._instance)                  # not None — looks fine
print(Reaper._container.get_logs()[0].decode())       # no "New client connected"

# ask the socket whether anyone is on the other end
import socket
s = Reaper._socket
s.settimeout(2)
s.send(b"label=org.testcontainers.session-id=probe\r\n")
print(s.recv(64))                                     # ConnectionResetError instead of b"ACK"

Then start any container in a child process and SIGKILL the process: the container is still there a minute later, and ryuk exits with Removed 0 container(s).

Suggested fix

Two independent halves, either of which closes the hole:

  • apply the wait strategy before connecting — e.g. move waiting_for(...) above .start(), or explicitly wait for the Started! log line after starting;
  • read Ryuk's reply after sending the filter and treat a missing ACK as "not connected to the reaper", retrying the connection. This is the stronger of the two, because it verifies the property that actually matters — the reaper accepted the filter — rather than a proxy for it.

Runtime environment

  • testcontainers-python 4.14.2
  • ryuk 0.8.1 and 0.11.0 (both affected)
  • Docker CE 29.1.3 on Ubuntu (WSL2) — affected; Docker Desktop on macOS — not affected

Related but different

#1093 describes the case where s.connect(...) raises and the failure is visible. This report is the opposite: the connection succeeds, no exception is raised, and the reaper silently does nothing.

Activity

  1. PHcz commented on Sep 25, 2026

    @PHcz

    Docker Desktop on macOS is affected too, just less often, and it can be reproduced on Docker Desktop every time. That may help with #1124, since its author has no docker-proxy host to test end to end.

    What we see on Docker Desktop. A connect() to ryuk's published port made before ryuk calls net.Listen is accepted by Docker Desktop's host-side port forwarder and then closed. connect() and send() succeed, a read would return b"", and ryuk logs no New client connected. Ryuk then runs out its 60 s connection timeout, logs Removed 0 container(s), exits 0 and is auto-removed, exactly as in the report above. On an idle machine ryuk usually listens in time: over 32 bring-ups it logged Started! between 6.8 ms before and 18.4 ms after start() returned, and testcontainers connected 4-30 ms after it. That is probably why the 0/5 above came out that way. Under CPU load it loses the race more often. All of these went through Reaper.get_instance() with ryuk's reply read on Reaper._socket afterwards:

    bring-ups no ACK host 1-min load (12 CPUs)
    ~505 3 (~0.6 %) 14-40
    1508 16 (~1.1 %) 10-44
    499 11 (2.2 %) 13-33
    657 (4 concurrent workers) 24 (3.7 %; 3 of 139 below load 20) 5.6-97

    Deterministic reproduction on Docker Desktop. Give ryuk Docker's minimum CPU quota (nano_cpus=10_000_000, i.e. --cpus 0.01). It then logs Started! 232-2104 ms after start() returns (32 runs), while testcontainers connects 5-52 ms after it. Through testcontainers' own Reaper path, 22 of 22 reaper connections were closed without an ACK. With the Started! wait set before start(), the same slow ryuk was acknowledged every time. Script below.

    Both halves of the suggested fix are needed. With the Started! wait set before start(), the misses dropped to zero in two differentials run side by side: 0 of 484 against 11 of 499, and 0 of 643 against 24 of 657. Most misses had connected before ryuk logged Started! (10 of 11 and 23 of 24), but 1 of 11 and 1 of 24 connected 10.4 ms and 4.1 ms after it and were closed all the same. So the wait removes most of the problem, and reading the ACK is still what guarantees the filter was stored, as #1124 does. One difference from #1124: when we got no ACK, we replaced the reaper with a fresh container (new host port) instead of reconnecting to the same one. Across 1508 + 657 bring-ups, 39 needed a replacement and every one ended with the filter stored. We did not measure retrying against the same ryuk.

    Once acknowledged, the link held: an idle connection stayed ESTABLISHED for 20 minutes under load 5-40, and ryuk pruned 10 s after it was finally closed.

    Reproduction script (Docker Desktop, deterministic)
    """Reaper.get_instance() can return a reaper that never received the session filter.
    
    Usage:
        python repro_ack_not_read.py slow 3     # deterministic: ryuk on a 0.01-CPU quota
        python repro_ack_not_read.py load 300   # stochastic: run under heavy CPU load
    """
    
    import sys
    import time
    import uuid
    
    import testcontainers.core.container as tcc
    from testcontainers.core.config import testcontainers_config as c
    from testcontainers.core.container import DockerContainer, Reaper
    
    MODE = sys.argv[1] if len(sys.argv) > 1 else "slow"
    N = int(sys.argv[2]) if len(sys.argv) > 2 else 3
    
    # Publish ryuk's 8080 on 127.0.0.1 only (host port still kernel-assigned), just to keep the
    # experiment's Docker-socket-holding ryuk off the LAN. Not part of the issue.
    _orig_with_exposed_ports = DockerContainer.with_exposed_ports
    
    
    def _loopback(self, *ports):
        out = _orig_with_exposed_ports(self, *ports)
        if self.image == c.ryuk_image:
            for p in ports:
                self.ports[str(p)] = ("127.0.0.1", None)
        return out
    
    
    DockerContainer.with_exposed_ports = _loopback
    
    if MODE == "slow":
        # Docker's minimum CPU quota: ryuk reaches net.Listen hundreds of ms after start()
        # returns, so testcontainers' connect() always lands first.
        _orig_with_kwargs = DockerContainer.with_kwargs
    
        def _slow(self, **kwargs):
            if self.image == c.ryuk_image:
                kwargs = {**kwargs, "nano_cpus": 10_000_000}
            return _orig_with_kwargs(self, **kwargs)
    
        DockerContainer.with_kwargs = _slow
    
    
    def one():
        tcc.SESSION_ID = f"repro-{uuid.uuid4()}"  # the id Reaper names/filters with
        Reaper._instance = Reaper._socket = Reaper._container = None
        Reaper.get_instance()  # returns normally in every case
        link = Reaper._socket
        link.settimeout(5)
        try:
            answer = link.recv(4)  # b"ACK\n" if ryuk stored the filter; b"" = closed
        except TimeoutError:
            answer = None
        wrapped = Reaper._container.get_wrapped_container()
        deadline = time.monotonic() + 10  # let a slow ryuk get to "Started!" before looking
        while "Started!" not in wrapped.logs().decode(errors="replace") and time.monotonic() < deadline:
            time.sleep(0.1)
        time.sleep(0.5)  # docker log delivery is asynchronous
        log = wrapped.logs().decode(errors="replace")
        Reaper.delete_instance()
        return answer, "New client connected" in log, log
    
    
    not_acked = 0
    for i in range(1, N + 1):
        answer, ryuk_saw_client, log = one()
        if answer != b"ACK\n":
            not_acked += 1
            print(f"#{i}: get_instance() returned, but ryuk answered {answer!r}; "
                  f"ryuk saw a client: {ryuk_saw_client}; ryuk log:\n{log}")
    print(f"{not_acked} of {N} reapers returned by get_instance() never acknowledged the filter")

    Output of slow 3 (iterations 2 and 3 are the same apart from the times):

    #1: get_instance() returned, but ryuk answered b''; ryuk saw a client: False; ryuk log:
    2026/09/25 07:37:39 Pinging Docker...
    2026/09/25 07:37:39 Docker daemon is available!
    2026/09/25 07:37:39 Starting on port 8080...
    2026/09/25 07:37:40 Started!
    ...
    3 of 3 reapers returned by get_instance() never acknowledged the filter
    

    Environment: testcontainers-python 4.15.0 (the Reaper code is the same on main today), ryuk 0.8.1, docker-py 7.2.0, Python 3.13.13; Docker Desktop 4.88.1, Engine 29.7.2 (API 1.55), linuxkit 7.0.12, arm64, 12 CPUs; macOS 27.0 on Apple silicon.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions