Repository navigation
Bug: Reaper reports success while the session filter never reaches Ryuk (silent leak on Linux/docker-proxy) #1114
Description
Activity
Docker Desktop on macOS is affected too, just less often, and it can be reproduced on Docker Desktop every time. That may help with #1124, since its author has no docker-proxy host to test end to end.
What we see on Docker Desktop. A
connect()to ryuk's published port made before ryuk callsnet.Listenis accepted by Docker Desktop's host-side port forwarder and then closed.connect()andsend()succeed, a read would returnb"", and ryuk logs noNew client connected. Ryuk then runs out its 60 s connection timeout, logsRemoved 0 container(s), exits 0 and is auto-removed, exactly as in the report above. On an idle machine ryuk usually listens in time: over 32 bring-ups it loggedStarted!between 6.8 ms before and 18.4 ms afterstart()returned, and testcontainers connected 4-30 ms after it. That is probably why the 0/5 above came out that way. Under CPU load it loses the race more often. All of these went throughReaper.get_instance()with ryuk's reply read onReaper._socketafterwards:bring-ups no ACKhost 1-min load (12 CPUs) ~505 3 (~0.6 %) 14-40 1508 16 (~1.1 %) 10-44 499 11 (2.2 %) 13-33 657 (4 concurrent workers) 24 (3.7 %; 3 of 139 below load 20) 5.6-97 Deterministic reproduction on Docker Desktop. Give ryuk Docker's minimum CPU quota (
nano_cpus=10_000_000, i.e.--cpus 0.01). It then logsStarted!232-2104 ms afterstart()returns (32 runs), while testcontainers connects 5-52 ms after it. Through testcontainers' ownReaperpath, 22 of 22 reaper connections were closed without anACK. With theStarted!wait set beforestart(), the same slow ryuk was acknowledged every time. Script below.Both halves of the suggested fix are needed. With the
Started!wait set beforestart(), the misses dropped to zero in two differentials run side by side: 0 of 484 against 11 of 499, and 0 of 643 against 24 of 657. Most misses had connected before ryuk loggedStarted!(10 of 11 and 23 of 24), but 1 of 11 and 1 of 24 connected 10.4 ms and 4.1 ms after it and were closed all the same. So the wait removes most of the problem, and reading theACKis still what guarantees the filter was stored, as #1124 does. One difference from #1124: when we got noACK, we replaced the reaper with a fresh container (new host port) instead of reconnecting to the same one. Across 1508 + 657 bring-ups, 39 needed a replacement and every one ended with the filter stored. We did not measure retrying against the same ryuk.Once acknowledged, the link held: an idle connection stayed
ESTABLISHEDfor 20 minutes under load 5-40, and ryuk pruned 10 s after it was finally closed.Reproduction script (Docker Desktop, deterministic)
"""Reaper.get_instance() can return a reaper that never received the session filter. Usage: python repro_ack_not_read.py slow 3 # deterministic: ryuk on a 0.01-CPU quota python repro_ack_not_read.py load 300 # stochastic: run under heavy CPU load """ import sys import time import uuid import testcontainers.core.container as tcc from testcontainers.core.config import testcontainers_config as c from testcontainers.core.container import DockerContainer, Reaper MODE = sys.argv[1] if len(sys.argv) > 1 else "slow" N = int(sys.argv[2]) if len(sys.argv) > 2 else 3 # Publish ryuk's 8080 on 127.0.0.1 only (host port still kernel-assigned), just to keep the # experiment's Docker-socket-holding ryuk off the LAN. Not part of the issue. _orig_with_exposed_ports = DockerContainer.with_exposed_ports def _loopback(self, *ports): out = _orig_with_exposed_ports(self, *ports) if self.image == c.ryuk_image: for p in ports: self.ports[str(p)] = ("127.0.0.1", None) return out DockerContainer.with_exposed_ports = _loopback if MODE == "slow": # Docker's minimum CPU quota: ryuk reaches net.Listen hundreds of ms after start() # returns, so testcontainers' connect() always lands first. _orig_with_kwargs = DockerContainer.with_kwargs def _slow(self, **kwargs): if self.image == c.ryuk_image: kwargs = {**kwargs, "nano_cpus": 10_000_000} return _orig_with_kwargs(self, **kwargs) DockerContainer.with_kwargs = _slow def one(): tcc.SESSION_ID = f"repro-{uuid.uuid4()}" # the id Reaper names/filters with Reaper._instance = Reaper._socket = Reaper._container = None Reaper.get_instance() # returns normally in every case link = Reaper._socket link.settimeout(5) try: answer = link.recv(4) # b"ACK\n" if ryuk stored the filter; b"" = closed except TimeoutError: answer = None wrapped = Reaper._container.get_wrapped_container() deadline = time.monotonic() + 10 # let a slow ryuk get to "Started!" before looking while "Started!" not in wrapped.logs().decode(errors="replace") and time.monotonic() < deadline: time.sleep(0.1) time.sleep(0.5) # docker log delivery is asynchronous log = wrapped.logs().decode(errors="replace") Reaper.delete_instance() return answer, "New client connected" in log, log not_acked = 0 for i in range(1, N + 1): answer, ryuk_saw_client, log = one() if answer != b"ACK\n": not_acked += 1 print(f"#{i}: get_instance() returned, but ryuk answered {answer!r}; " f"ryuk saw a client: {ryuk_saw_client}; ryuk log:\n{log}") print(f"{not_acked} of {N} reapers returned by get_instance() never acknowledged the filter")
Output of
slow 3(iterations 2 and 3 are the same apart from the times):#1: get_instance() returned, but ryuk answered b''; ryuk saw a client: False; ryuk log: 2026/09/25 07:37:39 Pinging Docker... 2026/09/25 07:37:39 Docker daemon is available! 2026/09/25 07:37:39 Starting on port 8080... 2026/09/25 07:37:40 Started! ... 3 of 3 reapers returned by get_instance() never acknowledged the filterEnvironment: testcontainers-python 4.15.0 (the
Reapercode is the same onmaintoday), ryuk 0.8.1, docker-py 7.2.0, Python 3.13.13; Docker Desktop 4.88.1, Engine 29.7.2 (API 1.55), linuxkit 7.0.12, arm64, 12 CPUs; macOS 27.0 on Apple silicon.
Describe the bug
Reaper._create_instance()can return a fully "successful" reaper whose session filter never reached Ryuk. Nothing raises,Reaper._instanceis set, the ryuk container is running — and yet nothing is ever reaped. When the test process is killed (CI timeout, cancelled job, OOM), every container of that session survives.Three things combine (
testcontainers/core/container.py, 4.14.2):waiting_foris a builder setter — its own docstring says "Set a wait strategy to be used after container start" — so setting it on an already-started container has no effect. There is no wait forStarted!in practice.docker-proxyaccepts on the published port from container creation, before the ryuk process inside has bound to it. The connect loop only retries onConnectionRefusedError/OSError, so a connection that is accepted by the proxy and then reset is indistinguishable from success:ACKfor every accepted filter line; nothing checks for it:So the filter goes into a socket that is not (yet) the reaper, the peer resets it, and the library reports success.
Observed result
Ryuk's own log tells the story — no client ever registered:
There is no
New client connectedline, and after the run is killed its containers stay up indefinitely. This is how a CI runner accumulated 40 leftover containers, including a full set of live storages still running hours after the run that created them had ended.Why this is easy to miss
The failure is silent and platform-dependent. On macOS/Docker Desktop the published port does not accept connections before the process inside binds, so the library wins the race and everything looks correct. On a Linux daemon with
docker-proxyit loses.Measured: 5 out of 5 runs affected on a self-hosted Linux CI runner (Ubuntu in WSL2, Docker CE 29.1.3), 0 out of 5 on macOS with Docker Desktop, same library version and same code.
To Reproduce
On a Linux host with Docker CE:
Then start any container in a child process and
SIGKILLthe process: the container is still there a minute later, and ryuk exits withRemoved 0 container(s).Suggested fix
Two independent halves, either of which closes the hole:
waiting_for(...)above.start(), or explicitly wait for theStarted!log line after starting;ACKas "not connected to the reaper", retrying the connection. This is the stronger of the two, because it verifies the property that actually matters — the reaper accepted the filter — rather than a proxy for it.Runtime environment
Related but different
#1093 describes the case where
s.connect(...)raises and the failure is visible. This report is the opposite: the connection succeeds, no exception is raised, and the reaper silently does nothing.