Skip to content

fix(files): leading slashes no longer skip the deny list and fences; DELETE carries the deny list (abilityai/trinity-enterprise#792) - #3276

Merged
vybe merged 2 commits into
devfrom
feature/792-double-slash-path
Oct 6, 2026
Merged

vybe merged 2 commits into
devfrom
feature/792-double-slash-path

Conversation

@webmixgamer

@webmixgamer webmixgamer commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor

Summary

  • The backend file routes normalised a user path with posixpath.normpath, which keeps exactly two leading slashes. The agent server's Path.resolve() reads them as one. As a result, a //-prefixed path matched none of the path-anchored deny patterns and not the skills.manage fence. _normalize_user_path now collapses any run of leading slashes to one.
  • PUT, POST /mkdir and DELETE /files now send the agent the normalised path they checked, not the raw input. The agent echoes what it receives, so the response path / deleted fields and the agent's 404/409 messages now carry the absolute normalised path. The Files tab does not read them.
  • The agent-side copy of the same policy had the same shape: the PreToolUse file-guardrail.py and read-only-guard.py hooks normalised file_path with os.path.normpath, so a // path skipped their absolute patterns. Both now collapse leading slashes. This was found in PR review.
  • DELETE /files now applies the backend deny list, which it never did before. It also refuses a directory that holds a path-anchored protected path (.ssh, .claude, the home dir), because deleting the directory removes what is inside. The refusal is 403 Cannot delete protected path: {path}.

Changes

  • src/backend/services/agent_service/files.py:
    • the one-line normaliser fix;
    • _deny_anchor / _DENY_ANCHORS, derived from _FILE_WRITE_DENY_PATTERNS so the list is not duplicated;
    • _is_user_deletable_path, wired into delete_agent_file_logic after the access check and the skills.manage fence, and before the container lookup;
    • the normalised path forwarded on all three write routes.
  • tests/unit/test_files_protected_paths.py now imports the shipped module; it used to test an inline copy of the functions. It adds:
    • // rows and a delete table;
    • the pinned anchor set, plus the anchor-derivation rule tested on hypothetical glob shapes;
    • three Hypothesis properties: the normalised form equals an independent segment-stack resolution, extra leading slashes never change a verdict, and DELETE is refused exactly when writing is or the path holds an anchor;
    • route-logic checks;
    • a check through the real FastAPI router on all three routes.
  • docker/base-image/hooks/file-guardrail.py, read-only-guard.py: _normalise collapses leading slashes. No file is added, so the image's COPY set is unchanged.
  • tests/unit/test_ent792_file_guardrail_hook.py (new): runs the shipped hook's main() against the shipped baseline. tests/unit/test_read_only_guard.py: // rows.
  • tests/unit/test_ent596_skill_manager.py: // rows on the skills fence, for classification and for the three routes.
  • tests/test_files_guardrail_bypass.py (live tier): // traversal rows, DELETE refusals against a stopped agent, and a // mkdir row.
  • Docs:
    • feature-flows/file-browser.md and skill-manager-permission.md;
    • requirements/content-files.md 13.1;
    • a learnings fragment;
    • the /cso --diff report.

No migration and no frontend change. The only docker/ change is the two hook functions above.

Behaviour change (Files tab)

These deletes are now refused for every caller, owners included:

  • .ssh, .aws, .gcp and everything inside them;
  • .claude/settings*.json, and the whole .claude folder;
  • .env.* at any depth;
  • .credentials.enc;
  • anything under /proc.

Folders inside .claude (skills/, agents/, projects/) can still be deleted. The UI shows the refusal as the existing Failed to delete: … toast. The Delete button stays enabled; the frontend is unchanged, and #3277 tracks disabling it.

A directory that merely contains a basename-protected file, such as .env, is not refused. Lexical matching cannot see inside a directory, and the agent server's by-name block still applies to the file itself.

Test Plan

  • cd tests && pytest unit/test_files_protected_paths.py unit/test_ent596_skill_manager.py unit/test_ent792_file_guardrail_hook.py unit/test_read_only_guard.py -q: all pass. Together with every unit test file that touches the file routes or the guardrail hooks: 677 passed, random order.
  • Mutations, each restored byte-identical. Each one turns tests red:
    • reverting the normaliser line: 30 red;
    • dropping the DELETE check: 11 red;
    • dropping the holding-directory clause: 22 red;
    • forwarding the raw path: 4 red;
    • cutting anchors at the first glob: 2 red;
    • treating a string prefix of an anchor as a parent folder (a.startswith(head)): 5 red;
    • reverting the file-guardrail.py fix: 7 red; reverting the read-only-guard.py fix: 2 red.
  • /verify-local --skip-agent --skip-unit (rerun after the hook change, still passing): build and import-smoke, boot and health, test preflight, and integration all pass (70 passed, 13 agent- or PG-gated skips, 2 known false-fails deselected).
    • The full --skip-agent run's unit stage failed: 41 tests in four unrelated files that only fail in the full sequential run (test_ent666_objective_join, test_ent477_git_refresh_hook, test_ent477_definitions_endpoint, test_retention_floor).
    • All four pass alone (120 passed) and none touches the changed module.
  • Agent container (trinity-agent-base): //home/developer/…, /proc/self/root/… and /proc/self/cwd/… all resolve into /home/developer.
  • Hooks in the real image (worktree hooks mounted read-only into trinity-agent-base, no network):
    • the file guardrail denies //home/developer/.ssh/authorized_keys, //home/developer/.claude/settings.json and ///opt/trinity/hooks/lib.py (exit 2), and allows //home/developer/notes.md;
    • with read-only mode on, //home/developer/.claude/agents/x.md is denied and //home/developer/content/ok.txt is allowed.
    • Before the fix, the shipped hook allowed the // forms (exit 0).
  • Manual, on a running agent on localhost (as admin, TOKEN from POST /api/token). Verified on 4dc342003; the head ff6c6ec3a changes only tests and tests/registry.json, so the backend is identical:
    • Files tab with hidden files shown. Deleting .ssh, .ssh/<file>, .claude or .env.example gives the toast Failed to delete: Cannot delete protected path: …, and each is still on disk afterwards. An ordinary file deletes normally.
    • curl -X PUT "http://localhost:8000/api/agents/<agent>/files?path=//home/developer/.ssh/x" -H "Authorization: Bearer $TOKEN" -H 'Content-Type: application/json' -d '{"content":"x"}' returns 403 Cannot edit protected path: //home/developer/.ssh/x.
    • mkdir takes its path in the JSON body: curl -X POST "http://localhost:8000/api/agents/<agent>/files/mkdir" -H "Authorization: Bearer $TOKEN" -H 'Content-Type: application/json' -d '{"path":"//home/developer/.ssh/keys"}' returns 403 Cannot create folder in protected path: //home/developer/.ssh/keys.
    • An ordinary file written through ?path=//home/developer/notes.txt lands on notes.txt, which shows the normalised path is what reaches the agent.

Related

#3236 adds _touches_instructions, which calls the same _normalize_user_path, so it picks up this fix when it rebases. Its test_ent164 classification parametrize can then add the // rows.

Fixes abilityai/trinity-enterprise#792

🤖 Generated with Claude Code

…; DELETE carries the deny list (Abilityai/trinity-enterprise#792)

The backend file routes normalised a user path with posixpath.normpath, which
keeps exactly two leading slashes, while the agent server's Path.resolve()
reads them as one. A `//`-prefixed path matched none of the path-anchored deny
patterns and neither capability fence.

- _normalize_user_path collapses any run of leading slashes to one.
- PUT, mkdir and DELETE send the agent the normalised path that was checked,
  not the raw input, so the check and the action read one string.
- DELETE /files applies the deny list too (_is_user_deletable_path), after the
  access check and the skills.manage fence, before the container lookup. A
  directory that holds a path-anchored protected path is refused as well
  (.ssh, .claude, the home dir); anchors are derived from the pattern tuple.
  Refusal: 403 "Cannot delete protected path: {path}".

Tests run against the shipped module: test_files_protected_paths.py no longer
tests an inline copy. Added // rows, a delete table, the pinned anchor set,
three Hypothesis properties, route-logic checks and a real-router check on all
three routes (mkdir takes its path in the JSON body); // rows on the ent#596
fence; live-tier DELETE and // rows in test_files_guardrail_bypass.py.

Mutations, each restored byte-identical: reverting the normaliser line turns
30 tests red; dropping the DELETE check 11; dropping the holding-directory
clause 22; forwarding the raw path 4; cutting anchors at the first glob 2.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ing slashes too (Abilityai/trinity-enterprise#792)

PR review found the agent-side copy of the policy with the same shape: the
PreToolUse file-guardrail.py and read-only-guard.py normalised file_path with
os.path.normpath, which keeps exactly two leading slashes, then matched absolute
patterns such as /home/developer/.ssh/*. In the shipped image a //-prefixed
.ssh or .claude/settings.json path exited 0 where the one-slash form exited 2.
files.py names file-guardrail.py as KEEP IN SYNC.

- Both _normalise functions collapse any run of leading slashes to one.
- tests/unit/test_ent792_file_guardrail_hook.py runs the shipped hook's main()
  against the shipped baseline; test_read_only_guard.py gains // rows.
  Reverting the hook fix turns 7 and 2 tests red; the fixed hooks, mounted into
  trinity-agent-base, deny the // forms in the real image.

Review follow-ups in the same commit:
- a string prefix of an anchor ("/et", "/home/dev") is not a directory above
  it: allowed rows plus a segment-based oracle, so `a.startswith(head)` no
  longer survives (5 red);
- docs: the agent echoes the normalised path (path/deleted fields and its
  404/409 messages); the flow's backend list names .env.*; only the skills.manage
  fence exists on dev; the security report records the hook finding the first
  pass missed and corrects its .claude/skills claim.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@webmixgamer webmixgamer added the ui PR touches the frontend UI — triggers Playwright e2e tests label Oct 6, 2026

@vybe vybe left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

merge-train: batch validated on train #3279 (green).

@vybe
vybe merged commit 19bd6d3 into dev Oct 6, 2026
30 checks passed
vybe pushed a commit that referenced this pull request Oct 6, 2026
…665) — mechanical, per the merge-train note on the PR

#3276 and #3273 landed first and both added registry entries;
tests/registry.json rebuilt from the git stages.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
vybe pushed a commit that referenced this pull request Oct 6, 2026
…) — mechanical

#3276, #3273, #3241 and #3275 landed first; tests/registry.json
rebuilt from the git stages.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
vybe pushed a commit that referenced this pull request Oct 6, 2026
 dev merge

The merge-train's registry rebuild for #3275 took the member branch's
list plus dev's new entries, so dev-side EDITS to existing entries were
lost: #3276's trinity-enterprise#792 text on test_files_guardrail_bypass.py
and test_ent596_skill_manager.py was reverted when #3275 landed
(58c7b80). Recomputed with a per-entry three-way merge from dev before
that landing: both descriptions are restored and #3276/#3273's entries
return to their positions. Metadata only; no test behaviour changes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
vybe added a commit that referenced this pull request Oct 6, 2026
… per subject (#3246, part 1) (#3255)

* feat(operator-queue): platform alert registry and sweep planner (#3246)

C1 of #3246 (one pending row per platform-alert subject). Adds the
stdlib-only leaf services/platform_alerts.py, which both migration tracks
and the service graph will import:

- Kind registry for every in-repo platform emitter, with its id prefix and
  lifetime class: 14 days by default (OPERATOR_PLATFORM_ALERT_LIFETIME_DAYS),
  a fixed 30-day net for edge-triggered kinds with a clear hook, or none for
  the four a person must act on. EXTERNAL_PREFIXES keeps role-drift- out;
  register() takes kinds from outside this module.
- Snooze window after a person ends an alert: 7 days
  (OPERATOR_PLATFORM_ALERT_SNOOZE_DAYS), per the operator's T5 ruling.
- Key belt and subject_for(): "<kind>:<key>", id-shaped keys kept,
  anything else sha256[:16]; event kinds have no subject.
- derive_legacy_subject(): reads every pre-#3246 id shape back to its
  subject, or to "known kind, no subject" when the id does not say which
  condition (git-bloat-, rows missing their context key). Unknown, gate
  and external prefixes are never derived.
- plan_sweep(): pure planner for the upgrade sweep. Keeps the newest
  pending row per (agent, subject) (created_at, then id; NULL oldest; a row
  that already holds a subject keeps it), ends the rest, stamps survivors,
  gives subject-less known kinds a lifetime stamp, stamps person-ended rows
  inside the snooze window, and plans nothing on its own output.

Nothing calls the leaf yet; the migration (C2) and the seam (C4) do.
Tests: tests/unit/test_3246_platform_alerts_leaf.py (78), including a
bare-interpreter load by path. Mutation-checked: survivor min instead of
max, greedy sid regex, no snooze window, and reversed prefix order each go
red.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(operator-queue): subject columns, backlog sweep and one-pending-per-subject index (#3246)

C2 of #3246, both schema tracks in one migration, ordered columns → sweep →
index: the partial unique index can only be created after the sweep has
collapsed the duplicates an installed backlog holds.

- `operator_queue.subject TEXT` + `last_seen_at TEXT` (nullable, no backfill
  beyond the sweep) in schema.py, tables.py and both migration tracks.
- `run_platform_alert_sweep(run)` in db/migrations.py — one function both
  tracks call over their own connection (sqlite3 cursor / SQLAlchemy
  `text()`), so the survivor rule cannot drift. It hands the fetched rows to
  the leaf's pure `plan_sweep` (function-local import; neither track pulls
  the service graph into `init_database()`), stamps the newest pending row
  per derived subject, ends the rest as ONE batch in the ent#611 vocabulary
  (`cancelled` / `platform` / `superseded` / NULL email / `batch_id`) with a
  compare-and-set on `status='pending'`, stamps a lifetime only on known
  kinds whose subject cannot be derived (never merged on a guess), and
  stamps `subject` on person-ended rows inside the 7-day snooze window.
  Agent-raised rows, `gate-` rows and external prefixes never derive.
- `uq_operator_queue_pending_subject` (partial: pending + subject) and
  `idx_operator_queue_agent_subject`, created after the sweep.
- Alembic `0090_platform_alert_subjects` chained after the live head
  `0089_supersede_queue_flood_backlog`.

Test (red first at collection — the migration did not exist):
tests/unit/test_3246_platform_alert_sweep.py drives the registered entry
against the pre-upgrade `schema.py` DDL holding duplicates: survivor and
stamps, one batch, untouched rows, snooze stamp, a second run changes
nothing, the index exists only after the run and bites on a second pending
row, DDL parity with schema.py, both tracks registered.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* feat(operator-queue): locked platform-alert accessors and the platform ending (#3246)

C3 of #3246 — the DB seam the platform-alert service (C4) will stand on.

db/operator_queue.py:
- `create_platform_item(agent, item, *, subject, max_pending_for_type)`:
  find → touch → count → insert in ONE `_lock_agent_for_create` transaction.
  A reading of a subject with a pending row updates it in place by a
  compare-and-set on `status='pending'` (title / question / priority /
  context with `seen_count`+1 / `last_seen_at` / `expires_at`); a lost CAS
  falls through to a fresh row, so a row a person ended is never
  overwritten. The find runs BEFORE the #1677 per-type count, so an update
  is never refused at budget. An `IntegrityError` from the partial unique
  index (a lock that failed open) re-finds and touches the winner. Returns
  `{"outcome": created|updated|refused_at_budget, "row", "changed"}`;
  `changed` is false on a bare repeat reading so nothing broadcasts.
- `find_pending_by_subject`, `find_person_ended_by_subject` (the seam's
  snooze read, newest person ending at or after `since`).
- `end_items_by_platform(ids, *, reason, batch_id)` — the `bulk_cancel_items`
  shape with `disposed_by='platform'`, NULL email, re-selected by batch id
  so only CAS-won rows come back.
- `mark_expired`'s per-id CAS gains `expires_at < now`: a row refreshed
  between the candidate select and the CAS keeps its new deadline.
- `_insert_values` takes keyword-only `subject` / `last_seen_at`;
  `_row_to_item` / `_SELECT_COLS` carry both.

services/ask_service.py: `clear_platform(ids, *, reason, batch_id=None)`
mirroring `expire()` — no Actor, `PLATFORM_ENDING_REASONS =
(condition_cleared, superseded)` enforced, one `platform_cleared` audit row,
one thin `operator_queue_cancelled` trigger per agent, observers get only
the rows this call won. database.py facade for the four accessors.
`test_ent329_operator_resume.py` G2 list gains the new CAS accessor.

#3130 / #3220: `_own_pending_conds`, the flood guard and
`create_bounded_alert_outcome` are untouched; their suites stay green.

Test (red first — the accessors did not exist):
tests/unit/test_3246_platform_alerts_db.py on the unit island's real
SQLite: two readings → one row, bare repeat → unchanged, person-ended row
never overwritten, budget refuses only with nothing to update, the index
refuses a second pending row, `mark_expired` keeps a refreshed row and still
expires an overdue one, the platform ending skips a person-ended row and
hands observers only CAS-won rows.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* test(operator-queue): register the #3246 C2/C3 unit files (#3246)

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* feat(operator-queue): platform alert seam, reserved prefixes and the emitter ratchet (#3246)

C4. `platform_alerts.observe / clear / reconcile` on top of C3's locked
accessors: one pending row per (agent, subject), updated in place; the
kind's lifetime as `expires_at`; the T5 rule (after a person ended a
subject's row, nothing is filed for 7 days unless the priority rose or a
material key changed); never raises. Imports stay function-local, so the
module is still a stdlib-only leaf at import time.

`skills-reconcile-`, `skills-fleet-reinject-`, `retention-guard-` and
`ent615-git-token-scrub-` become reserved prefixes.

Guards (test_3246_platform_alerts_registry.py): two-way prefix parity, the
EMITTERS_NOT_YET_ON_SEAM ratchet (every platform emitter not yet on the
seam, `create_bounded_alert_outcome` included, each naming its follow-up),
the clear-site check, and the caller guard on the #3246 accessors.

Also places C3's `subject` / `last_seen_at` on the machine side of the
ent#715 row allowlist (not person data); test_ent715 re-pinned — it was
red since C3.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(operator-queue): agent circuit breaker reports through the platform alert seam (#3246)

The dormant transition reports through `platform_alerts.observe` (kind
`circuit_dormant`, keyed by agent, 30-day net) and the two places the
code already knows the circuit closed — `record_success` recovering from
dormant, and the admin `reset_circuit` — end the row through
`platform_alerts.clear`. Leaves the #1677 allowlist and the ratchet.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(operator-queue): system agent stale base image reports through the platform alert seam (#3246)

Both #1816 alarms report through `platform_alerts.observe`, keyed by the
system agent: `base_image_stale` (high) and `system_agent_start_failed`
(critical), 30-day net. A newer reading updates the one pending row, so
the N-workers-N-rows boot and the per-bucket start-failure rows collapse
to one row each. Cleared where the code already knows: a running boot
that finds the image current, and a start that adopted the rebuilt image
(`image_drift`), end `base_image_stale`; a successful start (delegated or
the pre-flight plain start) ends `system_agent_start_failed`.

The per-process staleness cooldown stays as emission spacing.
`START_FAILED_ALERT_BUCKET_SECONDS` no longer shapes the id (only its own
constant test reads it). Five test_1816 tests re-pinned from the direct
DB write to the seam call; the cross-process dedup test now proves one
pending row on the real DB.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(operator-queue): subscription headroom reports through the platform alert seam (#3246)

Both ent#434 alerts report through `platform_alerts.observe` on the
`_sub-headroom` host: a subscription keys on the same cleaned id the legacy
`sub-headroom-{sid}-…` ids carried (`subject_key`), so a row the upgrade
sweep re-subjected matches a new reading; the fleet-wide alert keys on
`fleet`. A later reading updates the one pending row in place and a critical
reading replaces the warning (`tier` is the kind's material key), which
retires the documented "a 75% row still reads 75% at 92%" residual. The
reset day moves from the id to the row's context (`window`).

The evaluation pass had no place that named a recovery, so it now ends with
a reconcile (`clear_recovered`): only a subscription MEASURED back under its
threshold (`HAS_HEADROOM`) loses its row — an unassessable member keeps it —
and the fleet row survives until some member is measured with room.

Off the #1677 allowlist and the ratchet. ent#434's id-identity tests are
re-pinned to the subject; test_433_headroom_history asserted nothing about
the id shape and needed no change.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(operator-queue): legacy skills-library adoption reports through the platform alert seam (#3246)

`_record_adoption_failure` reports through `platform_alerts.observe` on the
`_skills-sync` host with ONE subject per URL (`sha256(raw url)[:12]`), shared
by the steady-state refusal and both failure branches. A repeat — any branch —
updates the one pending row (latest message and priority, seen count) rather
than filing another, so one sync never files a second row (#2744's
guarantee, now stated on the subject). Severity per branch is unchanged
(steady state low/info, failures high/error); the URL echo stays scrubbed.

Cleared where the code knows: adoption success and a URL that already names
a source end the URL's row (`_clear_adoption_alert`); every adoption pass
reconciles against the current setting (`_reconcile_adoption_alerts`), so a
removed or changed `skills_library_url` ends the rows it no longer backs.

Off the #1677 allowlist and the ratchet, which now holds only PR-2 entries.
test_2744 re-pinned from request_id to subject (test 4 now asserts the
failure branches keep `high` on the URL's one subject); three ent#346 tests
re-pinned from the direct DB write to the seam call.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(operator-queue): platform-ended alerts and the seen count on the queue surfaces (#3246)

A row the platform ended (disposed_by='platform', reason condition_cleared |
superseded) now reads "Ended by the platform — <why>" on ResolvedCard,
QueueItemDetail, /m and the Workspace ask views (one rule:
utils/operatorQueue.js::queueEnding/queueEndingText) — never a person's answer,
never a timeout, never the raw reason token as an operator note. An expired
alert reads "nobody acted on it". client_portal/asks/service.py::_ending_of
returns "platform" for such a row instead of falling through to "operator".

A pending alert seen more than once shows "seen N times · last seen <when>"
(queueSeenLine) on QueueCard, QueueItemDetail and /m; the Workspace asks panel
never renders platform alerts, so it gets no line.

Canary _PLATFORM_ALARM_SENTINELS gains _skills-sync (L-03 would otherwise read
the seam's rows as orphans). Stale "person | timeout" strings corrected in
dependencies.py, mcp-server types.ts and the get_my_ask description.

Tests: platformAlertSurfaces.spec.js (mounted; 12/15 red with the util change
reverted), test_3246_platform_alert_surfaces.py (4/7 red with the backend
change reverted).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(operator-queue): platform alerts are conditions — one row per subject (#3246)

Requirements §26.14 (OPS-001-PLATFORM-ALERTS), the operating-room flow
(status lifecycle + a Platform alerts section), the platform_alerts.py
catalog line, the subject/last_seen_at columns and the `platform` ending in
the database doc, and a learnings fragment. States plainly that only four
emitter families are on the seam in this PR (headroom, skills adoption,
system agent, circuit breaker); the rest follow in PR 2.

Corrects claims this branch made false: the headroom alert's id as its
state machine (security §20.x, backend catalog), #2744's one-row-per-URL
that a dismissal holds forever (skills), the base-image alarm filing one
item per worker that stays until acknowledged (internal-system-agent), and
disposed_by as an enum of two.

OPERATOR_PLATFORM_ALERT_LIFETIME_DAYS (14) and
OPERATOR_PLATFORM_ALERT_SNOOZE_DAYS (7) in .env.example and the dev, prod
and hosted compose files (hosted is the prod twin, #2280 parity).
tests/registry.json registers test_3246_platform_alert_surfaces.py.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(operator-queue): the upgrade sweep reads raised_by, never an id shape alone (#3246)

The sweep derived a subject from the request id and never read raised_by,
so an agent's own ask under a prefix that was unreserved before this
branch (hosted on the agent's own name) grouped with the platform's
alarm and, if newer, survived while the platform row was ended as
superseded. The SELECT now takes only platform-raised rows, every write
carries the same predicate, and plan_sweep skips any row with raised_by
set even when handed one. The survivor stamp no longer binds a NULL into
COALESCE on the SQLAlchemy track: expires_at is in the statement only
when the planner set one.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* fix(operator-queue): the platform-alert sync broadcast names its host agent (#3246)

platform_alerts._broadcast_sync carried no agent key, so the trigger
reached every logged-in client (ent#467). It is now keyed by the row's
host agent the way the create path's operator_queue_new is, so the scope
filter delivers it only to clients who may see that agent's queue. The
poller's fleet-level operator_queue_sync stays allowlisted: it is one
trigger per cycle that spans agents; this one is always about one row.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* fix(operator-queue): pin the snooze window at the headroom emitter (#3246)

Mutation check: a zero-day snooze window left every emitter test green.
A warning a person just ended must not be re-filed by the next identical
reading, while an escalation to critical still reaches a person.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* fix(operator-queue): the #3246 refresh-vs-CAS test patched a module name, not the dict mark_expired reads (#3246)

`test_a_refresh_between_select_and_cas_keeps_the_row` monkeypatched
`db.operator_queue.select` by module name. Under the full suite a sibling
file evicts or stubs `sys.modules["db.operator_queue"]`, so the name resolves
to a re-imported or stand-in module while the `database` singleton's class
still reads the original module dict — the patch never fires, the refresh
never lands, and `assert state["refreshed"]` fails as `assert False` (CI
seed 12345, one failure in 21,569). Alone, the file is green on every seed.

Same claim, proved deterministically: patch
`type(real_db._operator_queue_ops).mark_expired.__globals__["select"]`, the
dict the running method looks names up in — the suite's established shape
(test_1632 `_patch_engine`, test_ent611 race tests). `mark_expired` itself
is unchanged: its per-id CAS already re-checks `expires_at < now`.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* fix(operator-queue): the platform ending reaches the loop from any thread, not only anyio's (#3246 merge-train)

`spawn_on_loop` handled a running loop or an anyio worker thread and re-raised
from anything else — "not a shape any production caller has". #3246 made it one:
the headroom sweep (`subscription_recovery_service`) and the skills reconcile
(`skills_sync_service`, `routers/skills.py`) end platform alerts through
`asyncio.to_thread`, the loop's DEFAULT executor, which anyio does not own. The
anyio portal raised, `platform_alerts._spawn` / `ask_service._ended` swallowed
it, and every such ending committed with no `platform_cleared` audit row and no
`operator_queue_cancelled` trigger, behind a WARNING traceback. Reproduced: loop
thread OK, `anyio.to_thread` OK, `asyncio.to_thread` → RuntimeError.

Fixed in the helper rather than at the call sites: a captured host loop with a
`call_soon_threadsafe` fallback. `main.py::lifespan` records the loop before any
boot phase, and every on-loop spawn refreshes it, so the next thread-origin
caller — a `threading.Thread` a service starts itself included — inherits the
fix instead of each site having to remember `anyio.to_thread.run_sync`. The
task is still created ON the loop thread (`_schedule_on_loop`'s contract) and
the caller never waits for the work. A closed or non-running captured loop is
not a target (the private-loop footgun below), and with nothing to hop to the
helper still raises loudly — never the silent no-op ent#430 removed.

Pinned by `test_3246_platform_alerts_thread_origin.py`: `clear` / `observe`
driven through `asyncio.to_thread` land the audit row and the trigger with no
"could not schedule" warning; the spawn from the default executor and from a
plain thread; the lifespan capture is first in the boot sequence; the loud
failure with no loop; a closed captured loop is ignored.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* test(1816): a private loop cancels and flushes what the run spawned before closing (#3246 merge-train)

`test_1816_system_agent_adoption` ran `ensure_deployed` on a `new_event_loop()`
and closed it with the #3246 follow-ups (audit, broadcast) still registered in
`operator_resume_service._inflight` — pending tasks bound to a dead loop, or
finished ones whose `_inflight.discard` callback was queued when `stop()`
landed and never ran. The next test to gather `_inflight` on ITS loop failed
deterministically with "The future belongs to a different loop"
(`test_ent329_operator_resume::test_spawn_keeps_a_strong_reference`); CI
shuffles, so it recurred.

`_run` now cancels the leftover tasks (cancelled, not awaited: a follow-up may
wait on a transport this island never provides), gathers them, and runs one
more tick so the done callbacks drain, before closing. Verified in both file
orders.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* test(operator-queue): an open but stopped host loop is not a spawn target (#3246)

Merge-train validation found the `not loop.is_running()` rung of
`_live_host_loop` unpinned: deleting it left all 104 tests green. Such a
loop accepts `call_soon_threadsafe` and parks the task forever. The new
case captures one, expects the loud RuntimeError, and asserts nothing
was queued on it. Deleting the rung now turns it red.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* merge-train: restore #3276's registry descriptions reverted by the #3275 dev merge

The merge-train's registry rebuild for #3275 took the member branch's
list plus dev's new entries, so dev-side EDITS to existing entries were
lost: #3276's trinity-enterprise#792 text on test_files_guardrail_bypass.py
and test_ent596_skill_manager.py was reverted when #3275 landed
(58c7b80). Recomputed with a per-entry three-way merge from dev before
that landing: both descriptions are restored and #3276/#3273's entries
return to their positions. Metadata only; no test behaviour changes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Trinity Agent (trinity) <trinity-agent@ability.ai>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: trinity-ability <309458136+trinity-ability@users.noreply.github.com>
dolho added a commit that referenced this pull request Oct 7, 2026
…al copy that works (ent#164 review)

- require_person_or_capability: after the capability, the target must be
  owned by the holder's owner (owner_username equality, 404 otherwise). The
  reconfigure handlers authorise through can_user_share_agent, which admits
  any admin, so one grant on an admin-owned install reached every agent.
- read-only and guardrails stay person-only when the holder targets itself.
- The 403 tells the agent to raise an ask of type 'question' naming the
  permission (ask_class does not exist on the ask surface) and names the
  grant route, since the Settings toggles (ent#756) are not built.
- skills.manage granted through the generic route audits with the
  skill-manager verb, so both doors share one.
- instructions.manage copy names reset-to-main-preserve-state instead of
  claiming every git reset; the flow doc lists force_reset pull, sync
  pull_first, the instruction writes outside the path list, the other
  owner-tier writes and the grant-time calibration check as out of scope.
- Leading-slash paths (ent#792, fixed on dev by #3276) added to the
  instruction-path classification parametrize.
- Grant route tested against an admin-owned agent key and an admin
  user-scoped key; reach and self bounds tested through real keys.
- MCP descriptions for the schedule, agent, system and reset tools name the
  grant (budget test passes).
- security.md §5a/§6, auth.md and the flow doc updated.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
vybe pushed a commit that referenced this pull request Oct 7, 2026
…ai/trinity-enterprise#164) (#3236)

* feat(auth): an agent changes its own shape only with a grant (ent#164)

Re-scoped 2026-10-02: grants, not per-action approval. Three capabilities
join skills.manage (ent#596) in the closed set:

- schedules.manage: schedule create/update/delete/enable/disable and the
  webhook routes on ANOTHER agent. An agent's own schedules stay free
  (#2996); trigger is not fenced.
- instructions.manage: CLAUDE.md, AGENTS.md and .claude/** except skills
  through the file routes, plus git reset-to-main-preserve-state. Not
  grantable to a calibrating companion (ent#663).
- agents.manage: durable create, delete, deploy-local, systems/deploy,
  the chat model, and the read-only/resources/timeout/
  public-channel-model/guardrails writes (now person-or-grant). Ghost
  spawn and discard are exempt (ent#69).

Without the grant, an agent gets a named 403
(*_management_not_permitted). The 403 tells the agent to raise a
permission-request ask and says an admin grants the permission in
Settings. Every refusal is audited. Autonomy, api-key-setting,
capabilities, capacity and rename stay person-only.

New: GET/PUT /api/agents/{name}/capability-grants[/{capability}], with
PUT limited to an interactive admin and audited; this is the backend for
ent#756. The route census gains a person_or_grant class.

Related to Abilityai/trinity-enterprise#164

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(auth): bound an agents.manage holder to its owner's agents; refusal copy that works (ent#164 review)

- require_person_or_capability: after the capability, the target must be
  owned by the holder's owner (owner_username equality, 404 otherwise). The
  reconfigure handlers authorise through can_user_share_agent, which admits
  any admin, so one grant on an admin-owned install reached every agent.
- read-only and guardrails stay person-only when the holder targets itself.
- The 403 tells the agent to raise an ask of type 'question' naming the
  permission (ask_class does not exist on the ask surface) and names the
  grant route, since the Settings toggles (ent#756) are not built.
- skills.manage granted through the generic route audits with the
  skill-manager verb, so both doors share one.
- instructions.manage copy names reset-to-main-preserve-state instead of
  claiming every git reset; the flow doc lists force_reset pull, sync
  pull_first, the instruction writes outside the path list, the other
  owner-tier writes and the grant-time calibration check as out of scope.
- Leading-slash paths (ent#792, fixed on dev by #3276) added to the
  instruction-path classification parametrize.
- Grant route tested against an admin-owned agent key and an admin
  user-scoped key; reach and self bounds tested through real keys.
- MCP descriptions for the schedule, agent, system and reset tools name the
  grant (budget test passes).
- security.md §5a/§6, auth.md and the flow doc updated.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(auth): a capability holder reaches only its owner's agents on every fenced route (ent#164 validation)

The owner bound landed on the five reconfigure routes only. Every
`capability_fence` route (schedule writes, PUT /model, delete, git
reset) still authorised through `can_user_access_agent`, which admits
any admin, so an admin-owned holder of `schedules.manage` wrote
schedules on every agent on the instance and the agents.manage
refusal text overclaimed. One helper, `_refuse_unless_owners_agent`,
now runs inside the fence after the capability check, so a non-holder's
403 still comes first (#186). Pinned at the unit level across four
fences and through real keys and the real grant table; the schedule
write set is now derived from the router, not only a literal.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: trinity-ability <309458136+trinity-ability@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ui PR touches the frontend UI — triggers Playwright e2e tests

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants