Skip to content

Fix FFT cleaner aborting on already-stopped pools - #17

Open
cnnradams wants to merge 1 commit into
mainfrom
fix/fft-pool-cleaner
Open

Fix FFT cleaner aborting on already-stopped pools#17
cnnradams wants to merge 1 commit into
mainfrom
fix/fft-pool-cleaner

Conversation

@cnnradams

Copy link
Copy Markdown
Collaborator

When an FFT inference app has already been stopped, modal app stop exits with status 1. The scheduled cleaner currently raises before deleting that pool's registry entry and aborts the remaining sweep. In a live run, this left ten abandoned inference pools deployed with min_containers=1, even after their model/session records were reaped.

Treat the CLI's explicit “App is already stopped.” response as successful cleanup, while preserving other errors. Isolate each pool's cleanup so a failing entry is logged and does not prevent later pools from being processed; unsuccessful stops keep their registry entries for retry.

Regression coverage checks already-stopped registry removal, normal stop success, propagation of other stop errors, and continuation/retry after a failed stop. Existing active-model and touch-during-stop coverage remains in place.

Validation: uv run pytest -q tests/providers tests/control_plane (143 tests). No deployment or trainer configuration changes.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant