fix(cleanup): scope the #1581 orphan agent-volume sweep to this stack (#3214) - #3231
Conversation
…#3214) Docker volumes are daemon-global while the sweep's ownership check reads one stack's DB, so a second stack on a shared daemon reclaimed every other stack's unattached agent volumes — force-removed, logged at INFO. - Every agent data volume is created through docker_utils. agent_volume_labels, which stamps trinity.instance=<installation_id> (instance_identity.get_instance_id — the durable id the alert label already abbreviates; one source). All six creation sites use it, pinned by a guard test. - The sweep lists candidates on the instance label's key AND value and skips the cycle when the id is unresolvable. - is_reclaimable_agent_volume refuses a volume labelled for another instance, closing the same hole on the retention-purge path (two stacks can both have an agent `alpha`). - Legacy (unlabelled) volumes fail closed: labels are immutable, so the sweep never reclaims them; unowned+unattached ones are named in one WARNING per change of that set for a human. Own agents' volumes are still removed at retention purge, which is ownership-row driven. - Reclaims and removals log at WARNING as unrecoverable, naming volumes. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…mport - test_2669: get_instance_id() mints installation_id on purpose. A volume created before the id exists would carry no owner and never be reclaimed. It gets an _ALLOWED_USES entry with that reason, like the label tier's existing entry. - test_agent_readiness_probe: the hand-written services.docker_utils stub lacked agent_volume_labels, which lifecycle.py now imports. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
Regression details (head_sha: `bed902aee04683fd47ffe85ab3c09ea8fa2b71cf`)Seed Backend unit-suite regression diffPer-XML totals
❌ New failures introduced by HEAD (6)Tests failing under HEAD that did not fail under BASE in any seed:
Legend: [F] = assertion failure, [E] = collection or fixture error. Seed Backend unit-suite regression diffPer-XML totals
❌ New failures introduced by HEAD (6)Tests failing under HEAD that did not fail under BASE in any seed:
Legend: [F] = assertion failure, [E] = collection or fixture error. Seed Backend unit-suite regression diffPer-XML totals
❌ New failures introduced by HEAD (6)Tests failing under HEAD that did not fail under BASE in any seed:
Legend: [F] = assertion failure, [E] = collection or fixture error. Reproduce locally: |
|
merge-train (2026-10-05, evening run): on the train (#3251). Nothing was pushed to this branch. Ruling on legacy volumes (#3214 asked for one): fail-closed is accepted. Unlabelled volumes are never swept, so the sweep only applies to volumes created after the upgrade. Follow-ups from validation, none blocking, for a later PR:
|
Summary
The #1581 orphan agent-volume sweep decided "nobody owns this volume" from one stack's database while listing the whole Docker daemon's volumes. A second stack on a shared daemon therefore force-removed every other stack's unattached agent volumes, up to 100 per cycle, and logged it at INFO as a routine reclaim. On one developer host that was 207 of 238 volumes, per the 09-24 learning.
instance_identity.get_instance_id()returns the fullinstallation_id. This is the same durable, write-once id the alert label already abbreviates, so no second notion of identity is introduced. Any failure returnsNone, and every caller then fails closed.docker_utils.agent_volume_labels(base, platform)is the one label builder. It addstrinity.instance=<id>, and all six creation sites use it:crud.py×3,lifecycle.py×2 anddeploy.py. A guard test fails if any site builds agent-volume labels by hand.list_agent_data_volumes(instance_id)filters on the label key and value at the daemon, so a foreign volume is never a candidate. If this stack's id can't be resolved, the sweep skips the cycle.is_reclaimable_agent_volumerefuses a volume labelled for another instance, and refuses a labelled volume when our own id is unknown. Two stacks can both have an agent calledalpha, so the retention purge ofalphacould hit the other stack's volumes. This is in blast radius and not named in the issue.reliability.md(the sweep's section) and the 09-24 learning, which now points here (AC 6).Changes
services/instance_identity.py:get_instance_idservices/docker_utils.py:AGENT_VOLUME_INSTANCE_LABEL,agent_volume_labels, the scopedlist_agent_data_volumes(instance_id), report-onlylist_all_agent_data_volumes, and instance checks in the guard andremove_agent_volumesservices/cleanup_service.py: the scoped sweep,_report_unlabelled_orphan_volumes, and WARNING logsservices/agent_service/{crud,lifecycle,deploy}.py: the label builder at every creation sitetests/unit/test_3214_volume_instance_scope.py(18 new). Two existing#1581/#1664fakes were updated for the newinstance_idargument.Test Plan
test_3214_volume_instance_scope.py: 18 pass, and all 18 fail without the fix. They cover:-p randomly --randomly-seed=12345.Fixes #3214
🤖 Generated with Claude Code