Skip to content

feat(deploy): list Trinity on Vultr as an imageless marketplace app (#2282) - #3224

Merged
vybe merged 8 commits into
devfrom
feature/2282-vultr-imageless
Oct 5, 2026
Merged

vybe merged 8 commits into
devfrom
feature/2282-vultr-imageless

Conversation

@obasilakis

Copy link
Copy Markdown
Contributor

Vultr's Marketplace takes an imageless app: a Vendor Data script Vultr runs once, as root, through cloud-init on a clean Ubuntu 24.04. No snapshot, so nothing to rebuild or resubmit per release.

What this adds

  • scripts/deploy/vultr-vendor-data.sh (new) — the file pasted into the vendor portal. Resolves the latest release, clones it, exports TRINITY_IMAGE_TAG, sets ADMIN_PASSWORD_SOURCE=browser, installs a login banner, writes /etc/trinity/{ready,firstboot-failed} markers, then calls start.sh --provision --cloud vultr --provenance vultr-marketplace --hosted --unattended.
  • scripts/deploy/vultr-marketplace.md (new) — App Instructions text, listing copy, runbook.
  • start.sh — a vultr arm on provision_metadata_ip, vultr accepted by the cloud guard, and provision_apt() putting -o DPkg::Lock::Timeout=600 on every apt call in the machine phase.
  • Tests — test_2282_vultr_vendor_data.py (20), and test_2380_provision_single_source.py extended to scan the new script as a fourth caller.
  • Docs — PROV-017, amended PROV-011/012, and two feature flows.

Two decisions worth reviewing

Unpinned, not pinned. The script resolves releases/latest at boot rather than carrying a version literal. The copy that actually runs lives in Vultr's portal, so a repo literal is a pin nothing can enforce, and a VERSION test would guard the wrong copy while forcing an RC tag into it mid-cycle. releases/latest excludes pre-releases. A failed resolve aborts rather than falling back — a successful install of the wrong release is invisible. This deviates from AC #5, which asked for a pinned tag bumped per release; recorded on the issue.

No admin password, no app variables. trinity-enterprise#580 put marketplace installs on the browser admin claim and names Vultr. So ADMIN_PASSWORD_SOURCE=browser, nothing generated, and no password enters Vultr's metadata service — which serves app variables to every host process for the machine's life. This overrides AC #2 as written; also recorded on the issue. Vultr does support auto-generated password variables, so reinstating it later is a few lines.

The metadata read uses the plain-text path rather than /v1.json because it doubles as the "am I on a cloud VM" guard and runs before any package is installed, so it may depend on nothing but curl.

Live test

Instance vc2-4c-8gb (4 vCPU / 8 GB), Ubuntu 24.04, Frankfurt, script delivered as cloud-init user data; destroyed afterwards.

  • First boot start → serving: 2m20s on the final run (3m43s on an earlier one).
  • /v1/interfaces/0/ipv4/address returns the public IPv4; network-type reads public; a stock instance has one interface.
  • HTTPS 200 on the bare IP, chain verifies, Let's Encrypt short-lived certificate; http:// 308s to it.
  • .env: TRINITY_INSTALL_SOURCE=vultr-marketplace, tag persisted, ADMIN_PASSWORD empty, ADMIN_PASSWORD_SOURCE=browser.
  • 7 containers healthy; browser admin claim worked end to end; Settings reports Vultr Marketplace.
  • Four agents seeded — the acme trio plus Cornelius — and an empty operator queue.
  • DOCKER-USER → TRINITY-FW present with the 169.254.0.0/16 DROP: the host reads the metadata service, a container on trinity-agent-network times out.
  • Banner correct in all three states (installing / ready-unclaimed / ready-claimed); the EXIT trap wrote the failure marker on a failed run.

Not verified

  • Creating an agent on a box with no model credential (not Vultr-specific).
  • A VPC-attached instance — whether interface 0 stays the public one.
  • IPv6-only plans; the endpoint returns nothing, and the listing states IPv4 is required.

Sequencing

The listing cannot be published until a release contains this change. The script resolves releases/latest and clones that tag, so until a release carries --cloud vultr, a marketplace boot fails with unsupported --cloud 'vultr'. The remaining ACs — vendor application, Vultr's review, the README badge — are vendor-portal work, so this closes nothing.

Dev gained AWS provisioning while this was in flight; the merge keeps both arms, and the cloud guard now reads digitalocean|vultr|aws.

Refs #2282

🤖 Generated with Claude Code

obasilakis and others added 8 commits September 23, 2026 15:58
Vultr's Marketplace takes a Vendor Data script instead of a snapshot: Vultr
runs it once, as root, through cloud-init on a clean Ubuntu 24.04. So the
listing's artifact is a script, and everything the DigitalOcean 1-Click bakes
at build time happens at first boot instead — ~10 minutes against ~2, in
exchange for nothing to rebuild, re-review and re-publish per release.

`start.sh --provision` gains `--cloud vultr`: one arm reading Vultr's
plain-text metadata endpoint, and the cloud guard. That read doubles as the
"am I on a cloud VM" check and runs before any package is installed, so it
depends on nothing but curl — `/v1.json` would need a JSON parser. Verified
live: the endpoint returns the public IPv4, `network-type` reads `public`,
and a stock instance has one interface.

Every apt call in the machine phase now goes through `provision_apt`, which
carries `-o DPkg::Lock::Timeout=600`. Imageless is the first lane whose apt
runs at first boot, where Ubuntu's apt-daily and unattended-upgrades hold the
dpkg lock; a non-interactive apt-get that does not wait exits 100 under
`set -e`. The Packer bakery's `cloud-init status --wait` remedy is unavailable
to a script cloud-init is itself running. The DigitalOcean doc installer runs
the same code and gets the fix with it.

The script resolves the release rather than carrying one. The copy that runs
lives in Vultr's portal, so a literal here would be a pin nothing can enforce,
and a VERSION test would guard the wrong copy while forcing an RC tag into it
mid-cycle. `releases/latest` excludes pre-releases; a failed resolve aborts,
because a successful install of the wrong release is invisible to everyone.

No admin is provisioned (trinity-enterprise#580): the first visitor claims the
instance at /setup, so no password ever enters Vultr's metadata service, which
serves app variables to every host process for the life of the machine
(trinity-enterprise#622 item 2).

Without a snapshot there is nowhere to bake a status surface, so the script
installs its own login banner before anything can fail and writes the failure
marker from an EXIT trap. The banner separates installing, ready and failed —
a URL that refuses connections for ten minutes otherwise reads as broken.

Refs #2282

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…tempt

A live Vultr instance (vc2-4c-8gb, Frankfurt, Ubuntu 24.04) went from first-boot
start to serving in 3m43s, not the ~10 minutes estimated from the image sizes.
The listing copy sets expectations for someone staring at a refused connection,
so it carries the measurement rather than a guess.

`/var/log/trinity-install.log` is appended across attempts, so a FAILED line
from an earlier try sits above a later success — which misread the first live
run until the state markers settled it. The runbook now says to read from the
last first-boot header, or to trust `/etc/trinity/{ready,firstboot-failed}`,
which the script clears on entry.

Refs #2282

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The Cornelius template declared `fork_to_own: required` on 2026-09-11, closing
a real bind: it pushes a personal knowledge base to `Brain/`, and without that
declaration every agent created from it bound to the shared public upstream.

The first-run seeder cannot satisfy it. It runs at first boot with no user
present and no token to fork with, so since that date every fresh install has
ended with three agents instead of four, an ERROR in the backend log and a
high-priority "Cornelius seed failed" alert in the operator queue. Verified on
a clean marketplace install.

What the gate prevents is a PUSH. trinity-enterprise#705 already pins the
seeder pull-only, so for it that push is unreachable and the gate is guarding
something that cannot happen. `_apply_fork_to_own` now stands aside for that
one caller — and re-derives the pinned-pull-only shape itself rather than
taking the call site's word, so the flag alone cannot open the gate for a
create that could push.

Also corrects three copies of stale guidance that sent people to Settings for
the Claude key. Connecting Claude is a step in the first-run setup, and it is
the step that blocks.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Provisioning went through `services.agent_service.crud` directly, which takes
no `ws_manager`, and the `agent_created` broadcast is a no-op without one. So
Cornelius was the one seeded agent an already-open browser was never told
about. The fleet seeder uses the `routers/agents.py` facade for exactly this
reason and says so in a comment; this takes the same door.

Docs described a local bundled template with no git origin. That bundle was
deleted in #1656 — Cornelius clones the public upstream and, since
trinity-enterprise#705, tracks it pull-only. Both files now say that, and
record why the fork-to-own gate has a seeder-shaped door in it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…geless

# Conflicts:
#	docs/memory/requirements/infrastructure.md
#	scripts/deploy/start.sh
… — mechanical, per the merge-train note on the PR

test_the_cloud_allowlist_accepts_aws pinned the pre-vultr allowlist line
(digitalocean|aws) and message; this PR extends both to include vultr.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@vybe

vybe commented Oct 5, 2026

Copy link
Copy Markdown
Contributor

merge-train: pushed two commits to feature/2282-vultr-imageless so this rides today's train.

  • 3adbbd17a: merged origin/dev (clean).

  • 5a85599c8: updated the two pins in tests/unit/test_3004_aws_provision.py (test_the_cloud_allowlist_accepts_aws, the one regression diff failure) to the allowlist this PR introduces:

    • digitalocean\|aws\) ;; → digitalocean\|vultr\|aws\) ;;
    • "supported: digitalocean, aws" → "supported: digitalocean, vultr, aws"

    Locally, test_3004, test_2282 and test_2380_provision_single_source give 73 passed and 5 skipped. The skips need util-linux flock.

Non-blocking findings from validation, for a follow-up:

  • scripts/deploy/vultr-vendor-data.sh:67-69: the retry hint runs a bare start.sh, so a manual retry after an apt-phase failure loses ADMIN_PASSWORD_SOURCE=browser and the resolved tag. It also never clears firstboot-failed or writes ready, so the banner keeps saying the first boot failed.
  • Doc nits: the script header says ~10 min and PROV-017 says ~4. hosted-install.md lists <digitalocean|vultr> without aws.
  • Coverage: the vultr metadata-IP read and provision_apt are only pinned as source text. test_3004's _run harness could execute provision_metadata_ip with PROVISION_CLOUD=vultr and a stubbed curl.

@vybe vybe left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

merge-train: batch validated on train/20261005-1214 (#3228, all gates green)

@vybe
vybe merged commit 2067cd2 into dev Oct 5, 2026
25 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants