Skip to content

DO NOT MERGE — merge train: 3169,3162,3155,3147,3158,3156 - #3172

Closed
trinity-ability wants to merge 17 commits into
devfrom
train/20261001-1834
Closed

trinity-ability wants to merge 17 commits into
devfrom
train/20261001-1834

Conversation

@trinity-ability

Copy link
Copy Markdown
Contributor

Integration surface for #3169, #3162, #3155, #3147, #3158, #3156. Never merged; members merge individually once green.

🤖 Generated with Claude Code

dolho and others added 17 commits October 1, 2026 15:36
An external Workspace client opening an agent's page received other people's
runs of that agent: their run ids, triggers, start and end times, and
durations. The rows were never scoped to the viewer. Message, cost and model
were already projected away, but the timing still leaked (10 of 12 rows in the
ent#610 review walk).

For a non-platform viewer, every executions read on the page is now scoped in
SQL by one predicate, db/query_helpers.viewer_scope(). The viewer sees their
own turns (source_user_email), the runs those turns spawned (the inherited
source_channel_client), and the agent's scheduled runs, except a run of a
schedule that delivers to one other person (deliver_to_workspace_email, #498,
a seat's brief). Emails compare lower-cased, and a missing viewer admits shared
schedule runs only, so the scope fails closed.

The scope applies before the LIMIT, as the #2423 loop exclusion does, so
another person's busy day cannot push the viewer's own turns off the list. It
also covers every analytics aggregate, first_try_stats and "last active", so a
client's counts and rates are over exactly the rows they can see. first_try_stats
moves from text() to Core to share the predicate. The platform view and every
other caller of these accessors are unchanged (new keywords, defaults off).

Tests: test_3139_client_recent_work_scope.py, 12 tests on real SQLite through
the real accessors and page module. They cover two clients on one agent, a
spawned run, case-insensitive email matching, scope before the LIMIT, last
active, stats agreeing with the rows, rates, seat-delivered briefs, a missing
viewer, and an unchanged platform view. Verified by mutation: with the fix
reverted 9 of 12 fail; with a scope that admits every row, 7 fail; without the
seat-brief exclusion, 3 fail. The #2423 stubs take the new keywords. The
seat-brief exclusion came from the /cso --diff pass (one LOW finding).

Fixes #3139

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…#3140)

Two Workspace URLs offered a live composer for something the URL did not name:

- /workspace/c/<an id that is not yours>: the shell fell back to the first
  roster agent, so the page showed an empty chat of one of your own agents
  under someone else's chat id.
- /workspace/a/<an agent not shared with you>: landOnAgent opened a fresh chat
  with no roster check. The band then said "You don't have access" above
  "Start a conversation below." and a live composer, with a "Try again" that
  cannot help.

The chat case is decided in two places. The shell's thread list is the
viewer's whole set (no server-side LIMIT), so when it loaded cleanly an id it
lacks, and that the shell is not holding as a just-adopted new chat, renders
"This chat isn't available" with a way back. The conversation is never mounted
and its history is never read. When the list failed, the conversation still
tries, and a 404 on its history read emits `thread-missing`. The shell takes
that only while the URL still names the id, so a late answer cannot blank the
chat the person moved to.

An off-roster /workspace/a/<name> now uses the existing "You don't have
access" stage (the ?agent= rule) and gains a way back. The roster must have
loaded cleanly first, so a roster error keeps its own copy.

usePortalAgentPage exposes `denied` for a 403/404, and the band and the details
panel withhold "Try again" for it, since a retry gets the same answer. A 5xx
keeps its retry.

Tests (mounted, #2918), 14 in all:
- portalUnavailableTargets.mount.spec.js: the shell, both URLs, own chat and
  agent still open, in-app navigation, the failed-list fallback, a late 404, and
  the composable's denied flag.
- portalAgentBandDenied.mount.spec.js: no retry when refused.
- portalThreadMissing.mount.spec.js: a 404 emits, a 5xx does not.

Reverting each source file turns its tests red: Portal.vue 4, the composable
3, the band 2, the conversation 1. Full suite 4,512 passed; the design-token
check and the build pass. Looked at on a branch build against the local stack,
in light and dark: both states render with no composer and no page errors, the
way back works, and an own agent and chat still open with their composer.

Fixes #3140

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
After the first-run overlay closed on a fresh install, the Timeline showed no
agents for about 30 seconds, although setup had already seeded four. Three gaps
let that happen; each is fixed here.

1. Closing the overlay refetched nothing. Dashboard now watches the overlay's
   v-model and calls networkStore.fetchAgents() on true -> false, as
   onCreateModalClose already does for the create modal. That covers finish,
   skip and confirm-close. Setup's seeding does not go through the crud create
   path, so this refetch is what fixes the reported case.
2. stores/network.js had no `agent_created` handler. It now has one. The event
   triggers a refetch of the REST list, which is the access-controlled,
   tag-filtered source of the full row, rather than inserting the broadcast's
   partial row (the #918 thin-trigger rule). A burst (setup seeds four) is
   coalesced into one refetch within AGENT_CREATED_COALESCE_MS (300 ms).
3. The backend sent `agent_created` and `agent_deleted` with `event` but no
   `type`, and the dashboard dispatcher keys on `type`. Both now send `type`
   too, like agent_started and agent_stopped. That also revives the dead
   `agent_deleted` branch, which now reads the name from `data.data` where the
   payload carries it. The ent#467 /ws scoping reads `event` first, so it is
   unaffected (its tests pass).

Tests:
- test_3109_agent_ws_type.py: the real _broadcast_agent_created payload, plus
  an AST guard that every agent lifecycle broadcast in the backend carries a
  matching `type`.
- networkAgentCreatedWs.spec.js: the store's own WebSocket fed the backend's
  envelope. Four creates cause one refetch, the payload is never inserted, and
  a delete removes the agent.
- dashboardFirstRunRefetch.mount.spec.js: Dashboard mounted; closing the
  overlay refetches and opening it does not.

Reverting each source file turns its tests red: crud.py 2, agents.py 1,
network.js 3, Dashboard.vue 1. Full frontend suite 4,499 passed; the design
token check and the build pass; the related backend suites (ent#467 scoping,
ent#107 seed, capacity, backlog) pass.

Fixes #3109

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…atches on (#3109)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Throwaway pre-release so the AWS Marketplace AMI (#3004) can be built from
published images. A later 1.0.0 candidate supersedes it.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
SQLite reached end-of-support on 2026-09-01. Every one-click channel runs
docker-compose.hosted.yml through start.sh --hosted (or the compose file
alone), so the default changes there.

- docker-compose.hosted.yml: `postgres` service (postgres:16-alpine,
  postgres-data volume, pg_isready healthcheck, *default-logging,
  no-new-privileges, platform network only). Backend and scheduler wait
  for it healthy. POSTGRES_PASSWORD is required to render.
  DATABASE_URL=${DATABASE_URL-<bundled url>}: unset uses the bundled
  server, set-but-empty keeps SQLite.
- start.sh ensure_hosted_database: generates POSTGRES_PASSWORD (refuses if
  the postgres-data volume exists without one); with no DATABASE_URL in
  the shell or .env, writes the bundled URL on a fresh install and
  `DATABASE_URL=` when trinity.db exists, then prints the
  SQLITE_TO_POSTGRES.md pointer while the install stays on SQLite. Runs
  after the data-switch guard and before compose pull.
- Packer provision pre-pulls postgres:16-alpine (DO and AWS).
- Parity test allowlists the hosted-only service, volume, depends_on and
  DATABASE_URL lines; CI compose render exports POSTGRES_PASSWORD.
- Docs: HOST-022, SQLITE_TO_POSTGRES.md hosted section, DEPLOYMENT.md,
  .env.example, DO and AWS listings, DO deploy guide, topology table.

Live ACs (droplet boot, 4 GB RSS, nightly .dump on a droplet, mp-submit)
remain on the issue.

Refs #3159

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…er mounts and event delivery (Abilityai/trinity-enterprise#739)

An agent_permissions edge gates peer calls, shared-folder consume mounts and
event subscriptions. Calls re-read it on every call. The other two read it only
when the access was first set up, so withdrawing the edge stopped the next call
but not the files or the events.

Shared folders: check_shared_folder_mounts_match, the drift check behind the
start-time recreate, compared consume mounts in one direction only (permitted
but missing). It now also treats a /home/developer/shared-in/* mount that is
not in the currently permitted set as drift, which includes consume being
disabled while mounts remain. This mirrors the two-way expose check beside it.
The agent's next start rebuilds the container without the folder. A running
agent keeps it until then, and the docs say so.

Events: trigger_subscription re-reads the subscriber -> source edge on every
delivery, using the create-time rule (self-subscription needs no edge). It reads
the subscription's own source_agent, not the event's (a human emit carries a
username there). An unreadable grant skips the delivery, failing closed. The
subscription is kept, so granting the edge again resumes it. Agent rename
cascades both tables together, so a rename cannot desynchronise them.

Docs: agent-permissions.md -> Enforcement states when a withdrawn permission
takes effect for each control. system-manifest.md records that the per-agent
guardrails (disallowed_tools, extra_path_deny, extra_bash_deny) are not
manifest keys; they earn an unknown-key warning and configure nothing.

Tests: test_ent739_grant_withdrawal.py has 6 mount cases and 4 delivery cases.
Reverting helpers.py turns 2 red; reverting event_dispatch_service.py turns 3
red. Four suites whose subscription stubs carry no edge (#1578, #2973 budget,
ent#614, #3104) hold the grant with a named fixture, since they test what a
delivery carries, not whether it is permitted.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…/trinity-enterprise#729)

Tandem R45 ("Accept corrections"): a record_metrics point that repeats an
existing (metric, ts, dims) with a DIFFERENT value now updates the stored row
instead of being silently dropped as a duplicate. An identical value is still
a dedup that writes nothing.

- db/metric_points.insert_points: ON CONFLICT DO UPDATE ... WHERE the value IS
  DISTINCT FROM the stored one; value, execution_id, recorded_at and
  revision + 1 follow the correcting write; ts, created_at and dims never move.
  RETURNING revision gives a portable three-way count (PointWriteCounts:
  recorded / deduplicated / corrected). The store refuses two rows with one
  identity (PG raises, SQLite would apply both) and writes in (ts, key) order
  so overlapping batches cannot deadlock on PostgreSQL.
- Schema, both tracks: revision BIGINT NOT NULL DEFAULT 0 and recorded_at TEXT
  (nullable, no backfill: NULL = written before ent#729). SQLite migration
  metric_points_restatement + Alembic 0086_metric_points_restatement.
- The 201 MetricPointsResult and the MCP record_metrics result gain
  `corrected` (default 0 so stored idempotency snapshots still replay; `?? 0`
  against an older backend). One INFO line per correcting batch, counts only.
- Freshness, the daily cap and every read are unchanged: last_point_at is the
  newest ts, "used today" counts rows created, and get_metrics / the objective
  join read a corrected store exactly like a born-correct one (parity test).
- Contract text amended: service docstring, the tool description, the agent
  guide, user docs, requirements §48.2/§48.3/§48.6 + new §48.9 (incl. the
  stated limits: last write wins, lossy restatement, an identical earlier batch
  replays over a correction), architecture database/api/mcp entries and the
  custom-metrics flow.

Tests: tests/unit/test_ent729_metric_restatement.py (real rows, SQLite and
requires_postgres PostgreSQL, incl. the Alembic upgrade 0085 -> head over an
existing row), route + MCP receipt tests. Nine mutations of the applying lines
(WHERE, revision + 1, sort, duplicate guard, recorded_at, execution_id,
created_at, null-safety, the route's corrected) each turn a test red.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…enterprise#729)

From an independent backend / test-quality / contract review of #3169:

- Route: a store-shape refusal (the store's duplicate-identity ValueError)
  and an unreadable store result both answer the non-retryable 500
  `metric_store_rejected_batch` THROUGH `_reject`, so the batch claim is
  released. Before, the first reached the generic retryable 503 and the
  second escaped `_reject` and wedged the caller's key for 24 h. Both are
  unreachable today; the documented "every non-2xx exit past the claim
  releases it" is true again. Both new route tests fail against the old route.
- Tests: the SQLite migration test now runs the REGISTERED `MIGRATIONS`
  entry (deleting the registration used to pass every suite, schema parity
  included); the dims, daily-cap, write-order and store-read-key tests are
  rebuilt so the five mutants that survived them now go red; the Alembic test
  also asserts `NOT NULL` and `DEFAULT 0`; the rollback-insert test runs on
  PostgreSQL too.
- MCP tool + agent guide: the `execution_id` description states that an
  identical earlier batch in one turn replays over a correction.
- Docs: §48.9 replay-limit wording made exact; private-planning references
  dropped from public text; registry descriptions and the cso report's
  closed coverage gap updated; learnings fragment on testing a migration
  through its registered entry.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…t check (Abilityai/trinity-enterprise#739)

The two-way drift check caught every exception from volume_get as 'volume
missing', so a transient daemon fault dropped a present mount out of
'expected' and a start on a running agent recreated it (the #2196 class).
Narrow the skip to docker.errors.NotFound, matching the mount builder in
lifecycle.py; any other error keeps the present mount. Mechanical, per the
merge-train note on the PR.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…vailable' (#3140)

The sessionId watcher set unavailableChatId when the trusted list lacked the
id, but its 'known' branch never cleared it, so a chat created in another tab
after this one loaded its list stayed behind the unavailable screen until the
user navigated away. Clear it when the id turns up. Mechanical, per the
merge-train note on the PR; new mount test fails without the line.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ui PR touches the frontend UI — triggers Playwright e2e tests

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants