Skip to content

fix(operator-queue): platform alerts are conditions — one pending row per subject (#3246, part 1) - #3255

Open
vybe wants to merge 15 commits into
devfrom
feature/3246-platform-alerts-one-row
Open

vybe wants to merge 15 commits into
devfrom
feature/3246-platform-alerts-one-row

Conversation

@vybe

@vybe vybe commented Oct 5, 2026 •

Copy link
Copy Markdown
Contributor

Fixes #3246 — PR 1 of 2 (the rest is #3254). Independent of the ask-contract stack (#3250 → #3253 → #3247) that came out of the same session. Base is dev. Draft until the operator calls the merge order.

What

Platform alerts describe a condition, and the queue now stores them that way: one pending row per subject, updated in place, ended by the platform when the condition clears.

  • One shared path (services/platform_alerts.py): observe / clear / reconcile. A registry names each alert kind, its id prefix and its lifetime. observe never raises into its caller.
  • One pending row per subject: two nullable columns on the queue table (subject, last_seen_at) and a partial unique index on pending rows per agent and subject; the create runs inside the same per-agent lock as native asks. A newer reading updates the row (latest text, priority, last-seen, a seen count in context).
  • Ended by the platform: ask_service.clear_platform ends a row as cancelled / platform with reason condition_cleared or superseded, through the existing endings vocabulary and as a compare-and-set on pending, so a row a person ended is never overwritten.
  • Lifetimes: polled and per-event kinds expire 14 days after last seen; edge-triggered kinds with a clear hook get a fixed 30 days as a net; the four kinds a person must act on never expire. Two env knobs, in .env.example and all three compose files.
  • After a person ends an alert, the same reading files nothing for 7 days unless its priority rises or a material value changes.
  • Upgrade cleanup (Alembic 0090_platform_alert_subjects, both schema tracks): for each subject it can derive, the newest pending platform row stays and the older ones are ended as one recorded batch; the unique index is created only after that. It considers only rows the platform raised, never guesses a subject, and is idempotent.
  • Four emitter families on the path: subscription headroom (keyed by subscription, so a critical reading replaces the warning), legacy skills-library adoption (one subject per URL), the system agent's stale base image and start failure, the agent circuit breaker. Every other emitter stays on its old path and is listed in a ratchet test; they move in bug: the remaining platform alert emitters still file one row per reading — move them onto the shared alert path (#3246 part 2) #3254.
  • Four previously unreserved platform id prefixes are now reserved, so an agent cannot author rows under them.
  • Surfaces: a platform-ended row reads "Ended by the platform — the condition cleared" (or "superseded by a newer reading"); a repeated pending alert shows "seen N times · last seen …" on the Operations card, its detail and /m.

Rulings carried (orchestrator, on the operator's behalf — plan file)

Review + security

Review in two halves after an all-in-one review timed out with nothing on disk. Correctness + security (claude-fable-5-1, by reading): MERGEABLE AFTER FIXES. Evidence sweep (partial): one new red.

  • C1, fixed in ae988d0e: the upgrade cleanup derived a subject from the row id alone. An agent-raised row under a newly reserved prefix hosted on the agent's own name could be grouped with the platform's alarm and, if newer, survive while the platform alarm was ended. The cleanup now reads raised_by in the query and in the planner; 17 new cases cover agent and gate rows under all four prefixes.
  • C2, fixed in 95cb5cf9: the path's queue-sync WebSocket broadcast named no agent and reached every logged-in client (test_ent467_ws_agent_scope red). It is now keyed by the row's host agent.
  • 443c6b81 merges live dev; 499f2cca fixes a test that was order-dependent across files (see Tests); 34917656 adds the one emitter test a mutation check showed missing (the quiet window at the headroom emitter). All five named mutations now turn tests red.
  • /cso --diff: nothing further supported.

Tests

CI on the tip 499f2cca: all checks green — the full backend suite in CI's shuffled, parallel order (21,570 passed, 39 skipped; regression diff clean against the dev baseline), Postgres Migrations Smoke, frontend-e2e, MCP tests, schema parity, the container and docs guards.

  • CI's first full run found one failure that no engineer-side run had shown: this branch's own test_3246_platform_alerts_db…::test_a_refresh_between_select_and_cas_keeps_the_row. Cause: the test patched db.operator_queue.select by module name, and two sibling files stub that module in sys.modules, so under CI's order the patch landed on a different module object. Fixed in 499f2cca (the test patches the globals the running method reads; claim unchanged; 20 of 20 seeds green). Not a defect in mark_expired.
  • One existing test flaked once on the way: test_ent329_operator_resume::test_spawn_keeps_a_strong_reference ("The future belongs to a different loop"). This branch does not touch it; it passed on the previous head and on the re-run.
  • Engineer-side: the required pytest list (24 files) 851 passed; mcp-server 705 passed; frontend unit 4,633 passed.

Before merge

Handoffs

🤖 Generated with Claude Code

Trinity Agent (trinity) and others added 15 commits October 6, 2026 10:50
C1 of #3246 (one pending row per platform-alert subject). Adds the
stdlib-only leaf services/platform_alerts.py, which both migration tracks
and the service graph will import:

- Kind registry for every in-repo platform emitter, with its id prefix and
  lifetime class: 14 days by default (OPERATOR_PLATFORM_ALERT_LIFETIME_DAYS),
  a fixed 30-day net for edge-triggered kinds with a clear hook, or none for
  the four a person must act on. EXTERNAL_PREFIXES keeps role-drift- out;
  register() takes kinds from outside this module.
- Snooze window after a person ends an alert: 7 days
  (OPERATOR_PLATFORM_ALERT_SNOOZE_DAYS), per the operator's T5 ruling.
- Key belt and subject_for(): "<kind>:<key>", id-shaped keys kept,
  anything else sha256[:16]; event kinds have no subject.
- derive_legacy_subject(): reads every pre-#3246 id shape back to its
  subject, or to "known kind, no subject" when the id does not say which
  condition (git-bloat-, rows missing their context key). Unknown, gate
  and external prefixes are never derived.
- plan_sweep(): pure planner for the upgrade sweep. Keeps the newest
  pending row per (agent, subject) (created_at, then id; NULL oldest; a row
  that already holds a subject keeps it), ends the rest, stamps survivors,
  gives subject-less known kinds a lifetime stamp, stamps person-ended rows
  inside the snooze window, and plans nothing on its own output.

Nothing calls the leaf yet; the migration (C2) and the seam (C4) do.
Tests: tests/unit/test_3246_platform_alerts_leaf.py (78), including a
bare-interpreter load by path. Mutation-checked: survivor min instead of
max, greedy sid regex, no snooze window, and reversed prefix order each go
red.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…per-subject index (#3246)

C2 of #3246, both schema tracks in one migration, ordered columns → sweep →
index: the partial unique index can only be created after the sweep has
collapsed the duplicates an installed backlog holds.

- `operator_queue.subject TEXT` + `last_seen_at TEXT` (nullable, no backfill
  beyond the sweep) in schema.py, tables.py and both migration tracks.
- `run_platform_alert_sweep(run)` in db/migrations.py — one function both
  tracks call over their own connection (sqlite3 cursor / SQLAlchemy
  `text()`), so the survivor rule cannot drift. It hands the fetched rows to
  the leaf's pure `plan_sweep` (function-local import; neither track pulls
  the service graph into `init_database()`), stamps the newest pending row
  per derived subject, ends the rest as ONE batch in the ent#611 vocabulary
  (`cancelled` / `platform` / `superseded` / NULL email / `batch_id`) with a
  compare-and-set on `status='pending'`, stamps a lifetime only on known
  kinds whose subject cannot be derived (never merged on a guess), and
  stamps `subject` on person-ended rows inside the 7-day snooze window.
  Agent-raised rows, `gate-` rows and external prefixes never derive.
- `uq_operator_queue_pending_subject` (partial: pending + subject) and
  `idx_operator_queue_agent_subject`, created after the sweep.
- Alembic `0090_platform_alert_subjects` chained after the live head
  `0089_supersede_queue_flood_backlog`.

Test (red first at collection — the migration did not exist):
tests/unit/test_3246_platform_alert_sweep.py drives the registered entry
against the pre-upgrade `schema.py` DDL holding duplicates: survivor and
stamps, one batch, untouched rows, snooze stamp, a second run changes
nothing, the index exists only after the run and bites on a second pending
row, DDL parity with schema.py, both tracks registered.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…m ending (#3246)

C3 of #3246 — the DB seam the platform-alert service (C4) will stand on.

db/operator_queue.py:
- `create_platform_item(agent, item, *, subject, max_pending_for_type)`:
  find → touch → count → insert in ONE `_lock_agent_for_create` transaction.
  A reading of a subject with a pending row updates it in place by a
  compare-and-set on `status='pending'` (title / question / priority /
  context with `seen_count`+1 / `last_seen_at` / `expires_at`); a lost CAS
  falls through to a fresh row, so a row a person ended is never
  overwritten. The find runs BEFORE the #1677 per-type count, so an update
  is never refused at budget. An `IntegrityError` from the partial unique
  index (a lock that failed open) re-finds and touches the winner. Returns
  `{"outcome": created|updated|refused_at_budget, "row", "changed"}`;
  `changed` is false on a bare repeat reading so nothing broadcasts.
- `find_pending_by_subject`, `find_person_ended_by_subject` (the seam's
  snooze read, newest person ending at or after `since`).
- `end_items_by_platform(ids, *, reason, batch_id)` — the `bulk_cancel_items`
  shape with `disposed_by='platform'`, NULL email, re-selected by batch id
  so only CAS-won rows come back.
- `mark_expired`'s per-id CAS gains `expires_at < now`: a row refreshed
  between the candidate select and the CAS keeps its new deadline.
- `_insert_values` takes keyword-only `subject` / `last_seen_at`;
  `_row_to_item` / `_SELECT_COLS` carry both.

services/ask_service.py: `clear_platform(ids, *, reason, batch_id=None)`
mirroring `expire()` — no Actor, `PLATFORM_ENDING_REASONS =
(condition_cleared, superseded)` enforced, one `platform_cleared` audit row,
one thin `operator_queue_cancelled` trigger per agent, observers get only
the rows this call won. database.py facade for the four accessors.
`test_ent329_operator_resume.py` G2 list gains the new CAS accessor.

#3130 / #3220: `_own_pending_conds`, the flood guard and
`create_bounded_alert_outcome` are untouched; their suites stay green.

Test (red first — the accessors did not exist):
tests/unit/test_3246_platform_alerts_db.py on the unit island's real
SQLite: two readings → one row, bare repeat → unchanged, person-ended row
never overwritten, budget refuses only with nothing to update, the index
refuses a second pending row, `mark_expired` keeps a refreshed row and still
expires an overdue one, the platform ending skips a person-ended row and
hands observers only CAS-won rows.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…emitter ratchet (#3246)

C4. `platform_alerts.observe / clear / reconcile` on top of C3's locked
accessors: one pending row per (agent, subject), updated in place; the
kind's lifetime as `expires_at`; the T5 rule (after a person ended a
subject's row, nothing is filed for 7 days unless the priority rose or a
material key changed); never raises. Imports stay function-local, so the
module is still a stdlib-only leaf at import time.

`skills-reconcile-`, `skills-fleet-reinject-`, `retention-guard-` and
`ent615-git-token-scrub-` become reserved prefixes.

Guards (test_3246_platform_alerts_registry.py): two-way prefix parity, the
EMITTERS_NOT_YET_ON_SEAM ratchet (every platform emitter not yet on the
seam, `create_bounded_alert_outcome` included, each naming its follow-up),
the clear-site check, and the caller guard on the #3246 accessors.

Also places C3's `subject` / `last_seen_at` on the machine side of the
ent#715 row allowlist (not person data); test_ent715 re-pinned — it was
red since C3.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…orm alert seam (#3246)

The dormant transition reports through `platform_alerts.observe` (kind
`circuit_dormant`, keyed by agent, 30-day net) and the two places the
code already knows the circuit closed — `record_success` recovering from
dormant, and the admin `reset_circuit` — end the row through
`platform_alerts.clear`. Leaves the #1677 allowlist and the ratchet.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…he platform alert seam (#3246)

Both #1816 alarms report through `platform_alerts.observe`, keyed by the
system agent: `base_image_stale` (high) and `system_agent_start_failed`
(critical), 30-day net. A newer reading updates the one pending row, so
the N-workers-N-rows boot and the per-bucket start-failure rows collapse
to one row each. Cleared where the code already knows: a running boot
that finds the image current, and a start that adopted the rebuilt image
(`image_drift`), end `base_image_stale`; a successful start (delegated or
the pre-flight plain start) ends `system_agent_start_failed`.

The per-process staleness cooldown stays as emission spacing.
`START_FAILED_ALERT_BUCKET_SECONDS` no longer shapes the id (only its own
constant test reads it). Five test_1816 tests re-pinned from the direct
DB write to the seam call; the cross-process dedup test now proves one
pending row on the real DB.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…orm alert seam (#3246)

Both ent#434 alerts report through `platform_alerts.observe` on the
`_sub-headroom` host: a subscription keys on the same cleaned id the legacy
`sub-headroom-{sid}-…` ids carried (`subject_key`), so a row the upgrade
sweep re-subjected matches a new reading; the fleet-wide alert keys on
`fleet`. A later reading updates the one pending row in place and a critical
reading replaces the warning (`tier` is the kind's material key), which
retires the documented "a 75% row still reads 75% at 92%" residual. The
reset day moves from the id to the row's context (`window`).

The evaluation pass had no place that named a recovery, so it now ends with
a reconcile (`clear_recovered`): only a subscription MEASURED back under its
threshold (`HAS_HEADROOM`) loses its row — an unassessable member keeps it —
and the fleet row survives until some member is measured with room.

Off the #1677 allowlist and the ratchet. ent#434's id-identity tests are
re-pinned to the subject; test_433_headroom_history asserted nothing about
the id shape and needed no change.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…the platform alert seam (#3246)

`_record_adoption_failure` reports through `platform_alerts.observe` on the
`_skills-sync` host with ONE subject per URL (`sha256(raw url)[:12]`), shared
by the steady-state refusal and both failure branches. A repeat — any branch —
updates the one pending row (latest message and priority, seen count) rather
than filing another, so one sync never files a second row (#2744's
guarantee, now stated on the subject). Severity per branch is unchanged
(steady state low/info, failures high/error); the URL echo stays scrubbed.

Cleared where the code knows: adoption success and a URL that already names
a source end the URL's row (`_clear_adoption_alert`); every adoption pass
reconciles against the current setting (`_reconcile_adoption_alerts`), so a
removed or changed `skills_library_url` ends the rows it no longer backs.

Off the #1677 allowlist and the ratchet, which now holds only PR-2 entries.
test_2744 re-pinned from request_id to subject (test 4 now asserts the
failure branches keep `high` on the URL's one subject); three ent#346 tests
re-pinned from the direct DB write to the seam call.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… queue surfaces (#3246)

A row the platform ended (disposed_by='platform', reason condition_cleared |
superseded) now reads "Ended by the platform — <why>" on ResolvedCard,
QueueItemDetail, /m and the Workspace ask views (one rule:
utils/operatorQueue.js::queueEnding/queueEndingText) — never a person's answer,
never a timeout, never the raw reason token as an operator note. An expired
alert reads "nobody acted on it". client_portal/asks/service.py::_ending_of
returns "platform" for such a row instead of falling through to "operator".

A pending alert seen more than once shows "seen N times · last seen <when>"
(queueSeenLine) on QueueCard, QueueItemDetail and /m; the Workspace asks panel
never renders platform alerts, so it gets no line.

Canary _PLATFORM_ALARM_SENTINELS gains _skills-sync (L-03 would otherwise read
the seam's rows as orphans). Stale "person | timeout" strings corrected in
dependencies.py, mcp-server types.ts and the get_my_ask description.

Tests: platformAlertSurfaces.spec.js (mounted; 12/15 red with the util change
reverted), test_3246_platform_alert_surfaces.py (4/7 red with the backend
change reverted).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…bject (#3246)

Requirements §26.14 (OPS-001-PLATFORM-ALERTS), the operating-room flow
(status lifecycle + a Platform alerts section), the platform_alerts.py
catalog line, the subject/last_seen_at columns and the `platform` ending in
the database doc, and a learnings fragment. States plainly that only four
emitter families are on the seam in this PR (headroom, skills adoption,
system agent, circuit breaker); the rest follow in PR 2.

Corrects claims this branch made false: the headroom alert's id as its
state machine (security §20.x, backend catalog), #2744's one-row-per-URL
that a dismissal holds forever (skills), the base-image alarm filing one
item per worker that stays until acknowledged (internal-system-agent), and
disposed_by as an enum of two.

OPERATOR_PLATFORM_ALERT_LIFETIME_DAYS (14) and
OPERATOR_PLATFORM_ALERT_SNOOZE_DAYS (7) in .env.example and the dev, prod
and hosted compose files (hosted is the prod twin, #2280 parity).
tests/registry.json registers test_3246_platform_alert_surfaces.py.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…hape alone (#3246)

The sweep derived a subject from the request id and never read raised_by,
so an agent's own ask under a prefix that was unreserved before this
branch (hosted on the agent's own name) grouped with the platform's
alarm and, if newer, survived while the platform row was ended as
superseded. The SELECT now takes only platform-raised rows, every write
carries the same predicate, and plan_sweep skips any row with raised_by
set even when handed one. The survivor stamp no longer binds a NULL into
COALESCE on the SQLAlchemy track: expires_at is in the statement only
when the planner set one.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
… agent (#3246)

platform_alerts._broadcast_sync carried no agent key, so the trigger
reached every logged-in client (ent#467). It is now keyed by the row's
host agent the way the create path's operator_queue_new is, so the scope
filter delivers it only to clients who may see that agent's queue. The
poller's fleet-level operator_queue_sync stays allowlisted: it is one
trigger per cycle that spans agents; this one is always about one row.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…3246)

Mutation check: a zero-day snooze window left every emitter test green.
A warning a person just ended must not be re-filed by the next identical
reading, while an escalation to critical still reaches a person.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…ame, not the dict mark_expired reads (#3246)

`test_a_refresh_between_select_and_cas_keeps_the_row` monkeypatched
`db.operator_queue.select` by module name. Under the full suite a sibling
file evicts or stubs `sys.modules["db.operator_queue"]`, so the name resolves
to a re-imported or stand-in module while the `database` singleton's class
still reads the original module dict — the patch never fires, the refresh
never lands, and `assert state["refreshed"]` fails as `assert False` (CI
seed 12345, one failure in 21,569). Alone, the file is green on every seed.

Same claim, proved deterministically: patch
`type(real_db._operator_queue_ops).mark_expired.__globals__["select"]`, the
dict the running method looks names up in — the suite's established shape
(test_1632 `_patch_engine`, test_ent611 race tests). `mark_expired` itself
is unchanged: its per-id CAS already re-checks `expires_at < now`.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@vybe
vybe force-pushed the feature/3246-platform-alerts-one-row branch from 499f2cc to 2d9f665 Compare October 6, 2026 10:09
@vybe
vybe marked this pull request as ready for review October 6, 2026 10:11
@vybe

vybe commented Oct 6, 2026

Copy link
Copy Markdown
Contributor Author

merge-train (2026-10-06): ejected — rides the next train once fixed. Design, SQL (lock-first upsert + partial unique backstop), both migration tracks, Alembic heads, compose defaults and every gate check out; 3/3 mutations red on the seam. One real defect:

spawn_on_loop raises from every thread-origin caller this PR introduces, so the platform ending loses its audit row and live trigger. subscription_recovery_service.py:418 (await asyncio.to_thread(alerts.clear_recovered, …)) and :424, skills_sync_service.py:183, routers/skills.py:289 and :1058 run the seam on the loop's default executor. operator_resume_service.spawn_on_loop (:559-579) handles only a running loop or an anyio worker thread and re-raises otherwise ("not a shape any production caller has" — now false). ask_service._ended (:993-997) and platform_alerts._spawn (:559-564) swallow it, so every 5-minute headroom sweep that clears a recovered subscription, and every skills auto-sync reconcile, commits the ending with no platform_cleared audit row, no operator_queue_cancelled trigger, and a WARNING traceback. Reproduced: loop thread OK, anyio.to_thread OK, asyncio.to_thread → RuntimeError. It escaped because test_1816 monkeypatches spawn_on_loop to a no-op and test_ent434 stubs observe.

Fix: anyio.to_thread.run_sync at those call sites, or a captured-loop call_soon_threadsafe fallback in spawn_on_loop; plus a thread-origin clear/observe test asserting the audit row lands.

Also: test_1816_system_agent_adoption.py followed by test_ent329_operator_resume.py::test_spawn_keeps_a_strong_reference fails deterministically (The future belongs to a different loop, from the new_event_loop() at test_1816:115); CI shuffles, so it will recur. And the PR body cites commits that no longer exist on the branch after the rebase.

Merge order after the fix: this PR leads the next train; #3256 and #3145 re-parent their 0090 onto it (disjoint tables, mechanical).

@vybe vybe added the status-needs-fix PR has an unaddressed review/validation finding; cleared by the author's next push (#2815) label Oct 6, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

status-needs-fix PR has an unaddressed review/validation finding; cleared by the author's next push (#2815) ui PR touches the frontend UI — triggers Playwright e2e tests

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant