|
| 1 | +# Deploy runbook (sandbox / production relay) |
| 2 | + |
| 3 | +Step-by-step procedure for updating a running nostream host after a new |
| 4 | +`ghcr.io/cameri/nostream:main` image is available. |
| 5 | + |
| 6 | +| Doc | Purpose | |
| 7 | +|-----|---------| |
| 8 | +| [`deploy/README.md`](../deploy/README.md) | Bootstrap, compose stacks, health endpoints, HAProxy setup | |
| 9 | +| [`docs/DEPLOYMENT.md`](DEPLOYMENT.md) | How CI publishes container images | |
| 10 | + |
| 11 | +## Image delivery vs relay recreate |
| 12 | + |
| 13 | +```text |
| 14 | +GitHub push to main |
| 15 | + → CI publishes ghcr.io/cameri/nostream:main |
| 16 | + → (optional) webhook pull stack loads the image on the host |
| 17 | + → operator runs this runbook (migrate + recreate relay(s)) |
| 18 | + → Cloudflare Tunnel / reverse proxy keeps pointing at 127.0.0.1:8008 |
| 19 | +``` |
| 20 | + |
| 21 | +The webhook pull stack **loads images only**. It does **not** run migrations or |
| 22 | +recreate relay containers. This runbook closes that gap. |
| 23 | + |
| 24 | +## Choose your stack |
| 25 | + |
| 26 | +| Stack | Compose file | Update script | Downtime | |
| 27 | +|-------|--------------|---------------|----------| |
| 28 | +| Single relay (default bootstrap) | `docker-compose.yml` | [`deploy/recreate-relay.sh`](../deploy/recreate-relay.sh) | Brief while `nostream` recreates | |
| 29 | +| HAProxy blue/green | `docker-compose.haproxy.yml` | [`deploy/rolling-relay-recreate.sh`](../deploy/rolling-relay-recreate.sh) | None when peer relay is healthy | |
| 30 | + |
| 31 | +Do **not** run both stacks at once — both bind `127.0.0.1:8008`. Initial HAProxy |
| 32 | +install and trusted-proxy settings are in |
| 33 | +[`deploy/README.md`](../deploy/README.md#zero-downtime-updates-haproxy-bluegreen). |
| 34 | + |
| 35 | +During any deploy, wait for `/readyz` to return `200` before treating the relay |
| 36 | +as good. Endpoint behavior, fan-out checks, drain on SIGTERM, and probe timeouts |
| 37 | +are documented under |
| 38 | +[Health checks](../deploy/README.md#health-checks) in `deploy/README.md`. |
| 39 | + |
| 40 | +## Prerequisites |
| 41 | + |
| 42 | +- Host bootstrapped (`deploy/bootstrap.sh /opt/nostream`) |
| 43 | +- New relay image on the host (`docker pull` or `docker load`) |
| 44 | +- Shell access (`cd /opt/nostream`) |
| 45 | +- Optional: record the previous image ID for rollback |
| 46 | + |
| 47 | +```bash |
| 48 | +docker image inspect ghcr.io/cameri/nostream:main --format '{{.Id}}' | tee /tmp/nostream-pre-deploy-image-id |
| 49 | +``` |
| 50 | + |
| 51 | +## Shared steps (both stacks) |
| 52 | + |
| 53 | +### 1. Refresh release-managed files (when release notes say so) |
| 54 | + |
| 55 | +```bash |
| 56 | +./deploy/bootstrap.sh /opt/nostream |
| 57 | +``` |
| 58 | + |
| 59 | +Existing `.env` and `.nostr/settings.yaml` are preserved. |
| 60 | + |
| 61 | +### 2. Load the new image (skip if webhook pull already did) |
| 62 | + |
| 63 | +```bash |
| 64 | +docker pull ghcr.io/cameri/nostream:main |
| 65 | +``` |
| 66 | + |
| 67 | +On hosts that cannot reach GHCR over IPv4, use `docker save` / `docker load`. See |
| 68 | +[`deploy/README.md`](../deploy/README.md#image-delivery-on-restricted-networks). |
| 69 | + |
| 70 | +## Single-relay deploy |
| 71 | + |
| 72 | +From a repository checkout (for scripts) or on the host: |
| 73 | + |
| 74 | +```bash |
| 75 | +chmod +x deploy/recreate-relay.sh |
| 76 | +./deploy/recreate-relay.sh /opt/nostream |
| 77 | +``` |
| 78 | + |
| 79 | +Equivalent manual steps: |
| 80 | + |
| 81 | +```bash |
| 82 | +cd /opt/nostream |
| 83 | +docker compose up --no-deps --force-recreate --abort-on-container-exit nostream-migrate |
| 84 | +docker compose up -d --force-recreate --no-deps nostream |
| 85 | +``` |
| 86 | + |
| 87 | +## HAProxy blue/green deploy |
| 88 | + |
| 89 | +After the new image is loaded: |
| 90 | + |
| 91 | +```bash |
| 92 | +cd /opt/nostream |
| 93 | +docker compose -f docker-compose.haproxy.yml run --rm nostream-migrate |
| 94 | +./rolling-relay-recreate.sh |
| 95 | +``` |
| 96 | + |
| 97 | +The rolling script replaces blue and green one at a time; the peer must stay up |
| 98 | +and `/readyz` healthy. See |
| 99 | +[`deploy/README.md`](../deploy/README.md#zero-downtime-updates-haproxy-bluegreen). |
| 100 | + |
| 101 | +## Verify |
| 102 | + |
| 103 | +```bash |
| 104 | +cd /opt/nostream |
| 105 | +docker compose ps # or: docker compose -f docker-compose.haproxy.yml ps |
| 106 | +curl -sf http://127.0.0.1:8008/readyz && echo |
| 107 | +curl -sf http://127.0.0.1:8008/healthz |
| 108 | +curl -s -H 'Accept: application/nostr+json' http://127.0.0.1:8008/ |
| 109 | +``` |
| 110 | + |
| 111 | +Expected: migrate exited `0`, relay(s) running, `/readyz` HTTP `200` with |
| 112 | +database and redis `"ok": true`, NIP-11 JSON at `/`. |
| 113 | + |
| 114 | +If admin is enabled, also check `/admin/health` through your normal auth path. |
| 115 | + |
| 116 | +### Watch logs (first few minutes) |
| 117 | + |
| 118 | +```bash |
| 119 | +docker compose logs -f --tail=100 nostream |
| 120 | +# HAProxy stack: nostream-blue and/or nostream-green |
| 121 | +``` |
| 122 | + |
| 123 | +## Rollback |
| 124 | + |
| 125 | +When the new relay fails readiness or behaves incorrectly: |
| 126 | + |
| 127 | +### 1. Stop the bad relay(s) |
| 128 | + |
| 129 | +Single stack: |
| 130 | + |
| 131 | +```bash |
| 132 | +cd /opt/nostream |
| 133 | +docker compose stop nostream |
| 134 | +``` |
| 135 | + |
| 136 | +HAProxy stack: stop the relay you just replaced; leave the healthy peer running. |
| 137 | + |
| 138 | +### 2. Restore the previous image |
| 139 | + |
| 140 | +```bash |
| 141 | +PREVIOUS_IMAGE="$(cat /tmp/nostream-pre-deploy-image-id)" |
| 142 | +docker tag "$PREVIOUS_IMAGE" ghcr.io/cameri/nostream:main |
| 143 | +``` |
| 144 | + |
| 145 | +Or `docker load -i /path/to/nostream-main-backup.tar.gz`. |
| 146 | + |
| 147 | +### 3. Recreate on the old image |
| 148 | + |
| 149 | +Use the same script as your stack (`recreate-relay.sh` or `rolling-relay-recreate.sh`). |
| 150 | + |
| 151 | +Do **not** re-run migrations against a downgraded image unless you know the |
| 152 | +schema is backward compatible. |
| 153 | + |
| 154 | +## Troubleshooting |
| 155 | + |
| 156 | +### `/readyz` stays `503` |
| 157 | + |
| 158 | +- Check Postgres and Redis: `docker compose ps` |
| 159 | +- Inspect relay logs: `docker compose logs nostream --tail=200` (or blue/green) |
| 160 | +- Confirm `.env` credentials match the running database and cache |
| 161 | +- On HAProxy backends, confirm `relayBroadcast.ok` in `/readyz` when fan-out is enabled |
| 162 | + |
| 163 | +### `nostream-migrate` exits non-zero |
| 164 | + |
| 165 | +- Read migrate logs: `docker compose logs nostream-migrate` |
| 166 | +- Do not recreate relay(s) until migrate succeeds |
| 167 | +- Escalate if a migration is destructive — deploy compatible image order for expand/contract migrations |
| 168 | + |
| 169 | +### Image pull fails (GHCR / IPv6) |
| 170 | + |
| 171 | +- Use `docker save` / `docker load` from a machine that can reach GHCR |
| 172 | +- Keep `pull_policy: never` on nostream services when using pre-loaded images |
| 173 | + |
| 174 | +### Relay up but clients cannot connect |
| 175 | + |
| 176 | +- Confirm tunnel or reverse proxy still targets `127.0.0.1:8008` |
| 177 | +- Webhook image delivery alone does not replace this runbook |
| 178 | + |
| 179 | +## Automation (out of scope here) |
| 180 | + |
| 181 | +Wiring migrate + recreate into the webhook pipeline after this manual procedure |
| 182 | +is proven on sandbox is a separate follow-up. |
0 commit comments