Skip to content

DO NOT MERGE — merge train: 3226,3225,3222,3224,3220,3181 - #3228

Closed
trinity-ability wants to merge 28 commits into
devfrom
train/20261005-1214
Closed

trinity-ability wants to merge 28 commits into
devfrom
train/20261005-1214

Conversation

@trinity-ability

Copy link
Copy Markdown
Contributor

Integration surface for #3226, #3225, #3222, #3224, #3220, #3181. Never merged; members merge individually once green.

🤖 Generated with Claude Code

obasilakis and others added 28 commits September 23, 2026 15:58
Vultr's Marketplace takes a Vendor Data script instead of a snapshot: Vultr
runs it once, as root, through cloud-init on a clean Ubuntu 24.04. So the
listing's artifact is a script, and everything the DigitalOcean 1-Click bakes
at build time happens at first boot instead — ~10 minutes against ~2, in
exchange for nothing to rebuild, re-review and re-publish per release.

`start.sh --provision` gains `--cloud vultr`: one arm reading Vultr's
plain-text metadata endpoint, and the cloud guard. That read doubles as the
"am I on a cloud VM" check and runs before any package is installed, so it
depends on nothing but curl — `/v1.json` would need a JSON parser. Verified
live: the endpoint returns the public IPv4, `network-type` reads `public`,
and a stock instance has one interface.

Every apt call in the machine phase now goes through `provision_apt`, which
carries `-o DPkg::Lock::Timeout=600`. Imageless is the first lane whose apt
runs at first boot, where Ubuntu's apt-daily and unattended-upgrades hold the
dpkg lock; a non-interactive apt-get that does not wait exits 100 under
`set -e`. The Packer bakery's `cloud-init status --wait` remedy is unavailable
to a script cloud-init is itself running. The DigitalOcean doc installer runs
the same code and gets the fix with it.

The script resolves the release rather than carrying one. The copy that runs
lives in Vultr's portal, so a literal here would be a pin nothing can enforce,
and a VERSION test would guard the wrong copy while forcing an RC tag into it
mid-cycle. `releases/latest` excludes pre-releases; a failed resolve aborts,
because a successful install of the wrong release is invisible to everyone.

No admin is provisioned (trinity-enterprise#580): the first visitor claims the
instance at /setup, so no password ever enters Vultr's metadata service, which
serves app variables to every host process for the life of the machine
(trinity-enterprise#622 item 2).

Without a snapshot there is nowhere to bake a status surface, so the script
installs its own login banner before anything can fail and writes the failure
marker from an EXIT trap. The banner separates installing, ready and failed —
a URL that refuses connections for ten minutes otherwise reads as broken.

Refs #2282

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…tempt

A live Vultr instance (vc2-4c-8gb, Frankfurt, Ubuntu 24.04) went from first-boot
start to serving in 3m43s, not the ~10 minutes estimated from the image sizes.
The listing copy sets expectations for someone staring at a refused connection,
so it carries the measurement rather than a guess.

`/var/log/trinity-install.log` is appended across attempts, so a FAILED line
from an earlier try sits above a later success — which misread the first live
run until the state markers settled it. The runbook now says to read from the
last first-boot header, or to trust `/etc/trinity/{ready,firstboot-failed}`,
which the script clears on entry.

Refs #2282

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The Cornelius template declared `fork_to_own: required` on 2026-09-11, closing
a real bind: it pushes a personal knowledge base to `Brain/`, and without that
declaration every agent created from it bound to the shared public upstream.

The first-run seeder cannot satisfy it. It runs at first boot with no user
present and no token to fork with, so since that date every fresh install has
ended with three agents instead of four, an ERROR in the backend log and a
high-priority "Cornelius seed failed" alert in the operator queue. Verified on
a clean marketplace install.

What the gate prevents is a PUSH. trinity-enterprise#705 already pins the
seeder pull-only, so for it that push is unreachable and the gate is guarding
something that cannot happen. `_apply_fork_to_own` now stands aside for that
one caller — and re-derives the pinned-pull-only shape itself rather than
taking the call site's word, so the flag alone cannot open the gate for a
create that could push.

Also corrects three copies of stale guidance that sent people to Settings for
the Claude key. Connecting Claude is a step in the first-run setup, and it is
the step that blocks.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Provisioning went through `services.agent_service.crud` directly, which takes
no `ws_manager`, and the `agent_created` broadcast is a no-op without one. So
Cornelius was the one seeded agent an already-open browser was never told
about. The fleet seeder uses the `routers/agents.py` facade for exactly this
reason and says so in a comment; this takes the same door.

Docs described a local bundled template with no git origin. That bundle was
deleted in #1656 — Cornelius clones the public upstream and, since
trinity-enterprise#705, tracks it pull-only. Both files now say that, and
record why the fork-to-own gate has a seeder-shaped door in it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Dismiss (Abilityai/trinity-enterprise#748): one click ends a pending
question or approval as disposition=dismissed (status stays cancelled)
through the ask ending sink — audit, thin broadcast, and the filer's
wake with its own framing. 5s client-side Undo window; nothing is sent
until it lapses. Addressee only (person gate, uniform 404); a dismiss
that loses a race, or lands past the deadline, is a no-op.

Discuss (Abilityai/trinity-enterprise#747): opens one chat per ask with
the asking agent (CAS link in a new platform-only context key), titled
after the ask, with the ask as its tile; a turn-context provider puts
the ask (id, kind, live status, options) in every turn there. The ask
stays one pending row; a question can be answered from the composer
with "Send as answer" (written as response).

Also renders agent markdown in the ask context's "Your recent answers"
list, the surface #3115 missed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The Inbox turned Discuss's open-thread into a bare /workspace/c/<id>
push, and the #3140 guard read the just-created chat (not yet in the
thread list) as not the viewer's. Route it through the shell's
openThread, which adopts the id, and refresh the list. The chat tile
now forwards open-thread too — its Discuss did nothing before.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…#747)

An answer that woke an opted-in agent left the woken run's output in the
execution history only — the person saw the agent start and never saw
what it did. The resume run now carries the Workspace destination (the
ask's discussion chat, else its attached chat; only when the answerer is
the addressee), so the ent#457 completion report posts the result there
as an agent message. The open chat watches for it for up to 5 minutes,
since it has no history poll.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…-discuss-dismiss

# Conflicts:
#	src/frontend/src/components/portal/PortalConversation.vue
…rite errors

F1: an answered ask's resume run reports into the ask's DISCUSSION chat
only. The attached-chat fallback (Main for a background ask) widened every
answered ask's audience to the client — verbatim final reply, shared files,
inherited by delegated children. The run is now told its reply goes to
the person. Store watch narrowed to match. Tests pin turn_audience and
delegation inheritance for the stamped run.

F2: the discussion-link write tolerates only SQLite's busy snapshot
(OperationalError). Any other failure, or a busy write with no racing
link while the ask is still pending, is a retryable 503 instead of
"409 This ask is already pending".

F3: the discussion's system notice no longer quotes the agent-written title.
F4: options are truncated before JSON encoding, keeping the closing quote.
F5: Dismiss refuses alerts (422), like Discuss and the card.
F7: real TestClient tests through get_portal_principal (agent key 403)
and the live in-process limiter (429).
Send as answer stands down while a reply chip or attachment is in the
composer (#3168 collision). tables.py disposition comment lists dismissed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The #1632 flood alert was a raw platform create filed against the agent
whose depth it reported, the depth read counted it, and a sustained
over-cap condition minted a fresh timestamped alert every cooldown
window. Measured on one install: 386 of 435 rows were flood alerts, and
four file-seam requests were held with nothing telling the agent.

- The agent's own budget excludes platform-minted rows. Both depth caps
  (file seam, native ask_operator queue_full) read one predicate,
  _own_pending_conds, with _RESERVED_ID_PREFIXES excluded.
- The flood alert is edge-triggered: one per over-cap episode, re-armed
  only when a cycle holds nothing; the cooldown is now only the minimum
  spacing between episodes.
- It is a budgeted platform alert (#1677): type queue_flood through
  create_bounded_alert, as the backstop for failover. The UI labels it
  "Heads up" with the alert pill, as before.
- A held file is told why: a file-level platform.ingestion block
  (queue_full / rate_limited / invalid_id, max_pending, since) rides the
  guarded write-back, written only on change, removed when nothing is
  held; a refused write is retried only after the file changes.
- Migration supersede_queue_flood_backlog (SQLite + Alembic 0089) keeps
  the newest pending flood row per agent and cancels the rest
  (disposed_by 'platform', reason 'superseded').

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ow-ups (#3130)

- _write_responses_to_agent's ingestion argument defaults to "unchanged",
  never "remove": a caller that omits it must not strip a hold or force a
  write. Pinned by a test that fails with the old default.
- tables.py: disposed_by now also records 'platform' (the backlog
  migration's superseded flood alarms).
- Learnings fragment: an alarm must not count toward the limit it reports.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Scheduler dispatch-error path (_dispatch_and_record_outcome): a row the
backend already handed to the durable queue was read as finalized, so the
scheduler published schedule_execution_completed(status=queued), never
polled, and lost the real terminal plus retry/validation. Its FAILED write
was also unconditional and could land on a row a pull worker had claimed,
publishing a false failure that made the run retry-eligible. The FAILED
write is now a CAS on (running, claim_token IS NULL); a row it cannot fail
is polled to its real terminal through the same background poll the
accepted path uses.

Dispatch breaker on pilots: the half-open probe took a push slot, running
the turn outside the pilot's pool (#1982). It is now enqueued, and the pull
sink records the breaker verdict on its CAS-won branch (success, or auth),
so a pilot's breaker trips on pulled auth failures and closes on the probe
it pulled. Before this the pull sink recorded no verdict at all.

Tests: tests/unit/test_2514_scheduler_pull_gaps.py (15), including the
breaker-before-enqueue rule on the pull path.

Fixes #2514

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…geless

# Conflicts:
#	docs/memory/requirements/infrastructure.md
#	scripts/deploy/start.sh
The panel's list watcher only re-selected when the open canvas vanished
from a refresh, so a canvas the agent created while another was open
landed in the selector and stayed hidden.

- canvasesAppeared (canvasUtils.js) names the id a refresh brought that
  the previous list lacked: nothing on the first load or when the whole
  list was replaced (another agent's list under the same panel); the
  most recently updated when several arrive; never rows[0], so a pinned
  canvas sorting ahead does not steal it.
- CanvasPanel follows it through select() (canvas-selected emitted, the
  selectSeq race guard kept). While manage mode, a search or the share
  dialog is active the switch is deferred and happens when it ends; a
  canvas the reader picks meanwhile cancels it. A rewrite of a different
  existing canvas does not pull focus. Rewrite-in-place and the
  deleted-canvas fallback are unchanged.
- Flow doc: the selection rules in agent-canvas.md.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
operating-room.md: an answered ask's result returns to its discussion
chat only; the attached chat is not a destination (F1), so an
undiscussed answer's run stays owner-only.
backend.md: turn_context now has one OSS provider, the discussed-ask line.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ollow-up)

The dev merge 806b538 in #3022 resolved its conflicts on the #3021 files
by taking the branch side, which reverted #3021's (c3ba98a) review-driven
wording back to "never discards its own changes" / "never discards local
work". That is false: _safe_to_reset only spares registered executions, so
Files API, web terminal and docker exec writes made while a pull integrates
can still be lost.

Restores #3021's wording in the four places (Git sync settings panel copy,
_integrate_remote and _run_pull_once docstrings, the auto_sync pull-loop
comment, requirements/github.md) and keeps everything #3022 legitimately
added. Adds a mounted vitest assertion so the panel cannot regain the
claim. No behaviour change.

Related to #3022, #3021

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ncels a waiting follow (#3218)

Review C1: Agent Detail mounts CanvasPanel with canvases=[] before its list
loads. The empty branch set knownIds to an empty Set, so every row of the
first real list counted as an arrival and the most recently updated canvas
opened instead of the pinned-first row. The empty branch now resets knownIds
to null, so the next list is a first load.

Review I2: a match canvasAutoSelect selects during a search now goes through
pick(), clearing a pending follow, so clearing the query does not move the
reader off the canvas they searched for.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…he episode (#3130)

Review follow-ups on #3220:

- I1: create_bounded_alert gains an outcome-returning twin
  (create_bounded_alert_outcome; the bool API is unchanged). The flood
  emitter releases its episode on a FAILED count read or create so the
  alarm is retried once per cooldown window, instead of one transient
  failure suppressing it until the over-cap condition clears. A budget
  refusal still consumes the episode: the budget emits its own episode
  alert, and re-knocking every window adds only churn.
- I2: QueueItemDetail gives queue_flood the alert badge and renders the
  shared queueTypeLabel instead of the raw type string.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… — mechanical, per the merge-train note on the PR

test_the_cloud_allowlist_accepts_aws pinned the pre-vultr allowlist line
(digitalocean|aws) and message; this PR extends both to include vultr.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@vybe vybe added the ui PR touches the frontend UI — triggers Playwright e2e tests label Oct 5, 2026
@vybe

vybe commented Oct 5, 2026

Copy link
Copy Markdown
Contributor

Train complete: members merged individually (3221, 3223, 3226, 3225, 3222, 3224, 3220). 3181 waits on re-review.

@vybe vybe closed this Oct 5, 2026
@vybe
vybe deleted the train/20261005-1214 branch October 5, 2026 12:42
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ui PR touches the frontend UI — triggers Playwright e2e tests

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants