-
Notifications
You must be signed in to change notification settings - Fork 80
Expand file tree
/
Copy pathdocker-compose.cloud.yml
More file actions
379 lines (369 loc) · 17.3 KB
/
Copy pathdocker-compose.cloud.yml
File metadata and controls
379 lines (369 loc) · 17.3 KB
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
# =============================================================================
# AnythingMCP Cloud — Managed SaaS Deployment
# =============================================================================
# Usage: deploy/cloud/release.sh (via .github/workflows/deploy-cloud.yml).
# A bare `docker compose up -d` starts only the infrastructure: the app
# services are blue/green and belong to release.sh (see below).
# Domain: cloud.anythingmcp.com
# =============================================================================
name: amcp-cloud
# ── Backend and frontend: the app, in two colours ─────────────────────────
# Until 2026-09-20 both processes ran in one container under start.sh's
# liveness loop, which shut the container down whenever either process
# died. That afternoon the backend exhausted its heap four times, and each
# time the restart took the frontend and every tenant's MCP endpoint with
# it. Now a backend crash restarts the backend. The image is the same;
# `./start.sh backend` and `./start.sh frontend` exec one process each as
# PID 1, so `docker stop` and the backend's own heap guard reach Node
# directly and an exit becomes a restart through `restart: unless-stopped`.
#
# Each app service exists twice below, as <app>-blue and <app>-green, from
# these two definitions. One colour serves; a release starts the other next
# to it, checks it, moves Caddy over and only then stops the old one — see
# deploy/cloud/release.sh. The colours sit behind compose profiles, so a
# plain `docker compose up -d` never starts or recreates them: only
# release.sh does, one colour at a time. Do not `up` them by hand.
#
# Whichever colour is live is renamed to amcp-cloud-backend and
# amcp-cloud-frontend once the old one is gone, so `docker exec`, `docker
# logs` and scripts/ops keep working with the names they always used. The
# compose service (backend-blue / backend-green) is in the container's
# com.docker.compose.service label; Caddy addresses the service name.
x-backend: &backend
image: helpcodeai/anythingmcp:latest
command: ["./start.sh", "backend"]
expose:
- "4000"
# Hard ceiling for the kernel. It must sit far enough above the V8 cap
# that the backend's own guard (process-vitals.ts) acts first, and the
# margin is larger than it looks: on 21 Sep RSS ran 1.6 GB above the JS
# heap — native buffers, and V8's snapshot builder — so a 5 GB limit over
# a 4 GB heap was reached with the heap at 60 %, and the kernel killed
# the process nineteen times with no log line. 6 GB over 4 GB, with the
# guard now watching RSS against this limit as well (HEAP_RSS_EXIT_PERCENT),
# turns that into a logged, graceful restart. On the 7.9 GB droplet the
# rest — frontend, postgres, redis, motis, caddy — sits near 1 GB today.
# If NODE_MAX_OLD_SPACE_MB is raised, raise this too.
#
# During a release two backends run at once for a minute or two, and two
# of these limits (12 GB) do not fit the droplet. They do not have to:
# the limit is a ceiling for one runaway process, not a reservation, and
# what the overlap costs is the NEW backend's boot footprint (~0.7 GB
# steady, ~1.5 GB at the most while it loads the registry) on top of what
# the host already uses. release.sh checks that room exists before it
# starts the new colour (RELEASE_MIN_AVAILABLE_MB), aborts the new colour
# if the host runs short while it boots (RELEASE_MEM_ABORT_MB), and falls
# back to stop-then-start when the room is not there. The numbers and the
# reasoning are in release.sh's header.
mem_limit: ${BACKEND_MEM_LIMIT:-6g}
environment:
- NODE_ENV=production
- PORT=4000
# V8 old-space cap for the backend, read by start.sh. Left unset it
# defaults to 2048, which suits a 4 GB host; the tool registry outgrew
# that and crash-looped the instance on 10 Sep. Raise it in the server
# .env when the host has the RAM to back it — roughly half of total is a
# safe ceiling, since postgres, redis, motis and caddy share the box.
- NODE_MAX_OLD_SPACE_MB=${NODE_MAX_OLD_SPACE_MB:-2048}
# Heap guard (packages/backend/src/common/process-vitals.ts; every var
# documented in .env.example). The snapshot stays OFF here: on a heap
# this size V8's snapshot builder is what pushed RSS into the cgroup
# limit on 21 Sep. Enable it deliberately, at a low percentage, when a
# snapshot is wanted — the guard refuses it anyway unless the room is
# there.
- HEAP_SNAPSHOT_PERCENT=${HEAP_SNAPSHOT_PERCENT:-0}
- HEAP_EXIT_PERCENT=${HEAP_EXIT_PERCENT:-90}
- HEAP_RSS_EXIT_PERCENT=${HEAP_RSS_EXIT_PERCENT:-90}
- HEAP_WARN_PERCENT=${HEAP_WARN_PERCENT:-75}
- HEALTH_HEAP_PERCENT=${HEALTH_HEAP_PERCENT:-90}
- DEPLOYMENT_MODE=cloud
# Sentry (packages/backend/src/instrument.ts). Unset DSN = off. The
# release is baked into the image by the publish workflow.
# SENTRY_VERIFY_TOKEN enables GET /health/sentry-verify for a one-off
# end-to-end check; leave it empty otherwise.
- SENTRY_DSN=${SENTRY_DSN_BACKEND:-}
- SENTRY_ENVIRONMENT=${SENTRY_ENVIRONMENT:-production}
- SENTRY_TRACES_SAMPLE_RATE=${SENTRY_TRACES_SAMPLE_RATE:-0.05}
- SENTRY_VERIFY_TOKEN=${SENTRY_VERIFY_TOKEN:-}
- NEXT_PUBLIC_API_URL=https://${DOMAIN}
- DATABASE_URL=postgresql://amcp:${POSTGRES_PASSWORD}@postgres:5432/anythingmcp
- REDIS_URL=redis://redis:6379
- JWT_SECRET=${JWT_SECRET}
- ENCRYPTION_KEY=${ENCRYPTION_KEY}
# Shared secret for the onboarding-reminders GitHub Actions cron.
# Must match the repo secret ONBOARDING_CRON_SECRET. Unset = cron
# endpoint refuses all calls (self-host default).
- CRON_SECRET=${CRON_SECRET:-}
# Shared secret presented to anythingmcp.com as x-amcp-service-token when
# this backend asks for a licence on a user's behalf. The licence API
# rate-limits anonymous callers per IP, and every cloud trial request
# leaves from this one droplet, so without it the whole cloud shared a
# bucket of three trials an hour. Must match LICENSE_SERVICE_TOKEN on the
# website. Unset = public limit (self-host default).
- LICENSE_SERVICE_TOKEN=${LICENSE_SERVICE_TOKEN:-}
- CORS_ORIGIN=https://${DOMAIN}
- SERVER_URL=https://${DOMAIN}
- FRONTEND_URL=https://${DOMAIN}
- MCP_AUTH_MODE=${MCP_AUTH_MODE:-oauth2}
# MCP Streamable HTTP response framing. `true` = single-shot
# `application/json` responses; `false` = SSE-framed `text/event-stream`.
# Microsoft Copilot Studio's MCP client cannot deserialize SSE-framed
# responses ("response could not be deserialized as JSON") and requires
# plain JSON, so the cloud defaults to JSON. This is spec-compliant and
# every compliant client (incl. Claude) handles it; combined with the
# default stateless mode (MCP_STATEFUL_SESSIONS unset) it also avoids the
# `-32001 Session not found` errors Copilot hits against stateful servers.
- MCP_STREAMABLE_JSON_RESPONSE=${MCP_STREAMABLE_JSON_RESPONSE:-true}
- ALLOW_OPEN_REGISTRATION=${ALLOW_OPEN_REGISTRATION:-true}
- LICENSE_API_URL=${LICENSE_API_URL:-https://anythingmcp.com}
# Operator analytics — cloud build only. Leave unset on self-hosted.
- GTM_ID=${GTM_ID:-}
- COOKIE_DOMAIN=${COOKIE_DOMAIN:-}
# Proxy / web-unblocker (e.g. Zyte API proxy mode). Tools with
# use_proxy=true route through this; unset = feature off everywhere.
- CONNECTOR_PROXY_URL=${CONNECTOR_PROXY_URL:-}
- PROXY_RATE_LIMIT_DEFAULT=${PROXY_RATE_LIMIT_DEFAULT:-100}
# Internal MOTIS instance for the Deutsche Bahn connector. When set, the
# adapter's MOTIS_URL is filled in automatically at import (and hidden
# from the install form), so cloud users never see or choose it.
# Self-hosters leave it unset and enter their own MOTIS URL.
- MOTIS_INTERNAL_URL=${MOTIS_INTERNAL_URL:-http://motis:8080}
# Allow the SSRF guard to reach the internal MOTIS host (it resolves to
# a private docker IP, which the guard blocks by default). Harmless on
# self-host (no such host on their network).
- SSRF_ALLOWED_HOSTS=${SSRF_ALLOWED_HOSTS:-motis}
# Knowledge Graph — AI (LLM) features. Off unless KG_LLM_ENABLED=true AND a
# provider key are set; the per-workspace toggle still gates actual usage.
- KG_LLM_ENABLED=${KG_LLM_ENABLED:-false}
- KG_LLM_PROVIDER=${KG_LLM_PROVIDER:-}
- KG_LLM_MODEL=${KG_LLM_MODEL:-}
- ANTHROPIC_API_KEY=${ANTHROPIC_API_KEY:-}
- OPENAI_API_KEY=${OPENAI_API_KEY:-}
# Domain-verification token for the OpenAI plugin directory; served at
# /.well-known/openai-apps-challenge. Unset = route answers 404.
- OPENAI_APPS_CHALLENGE_TOKEN=${OPENAI_APPS_CHALLENGE_TOKEN:-}
- KG_LLM_REDACT_INTENTS=${KG_LLM_REDACT_INTENTS:-true}
# System (operator) transactional-email fallback — used when a workspace
# hasn't set its own SMTP (e.g. Resend). Read only by the email sender,
# never exposed to workspace admins. Set these in the droplet .env.
- SMTP_HOST=${SMTP_HOST:-}
- SMTP_PORT=${SMTP_PORT:-587}
- SMTP_USER=${SMTP_USER:-}
- SMTP_PASS=${SMTP_PASS:-}
- SMTP_FROM=${SMTP_FROM:-}
- SMTP_SECURE=${SMTP_SECURE:-false}
volumes:
# One heap snapshot per process, written by the heap guard when the heap
# crosses HEAP_SNAPSHOT_PERCENT. A volume so it outlives the container
# that wrote it — the whole point is to read it after the restart.
- diagnostics:/app/diagnostics
depends_on:
postgres:
condition: service_healthy
redis:
condition: service_healthy
healthcheck:
test: ["CMD", "wget", "--no-verbose", "--tries=1", "--spider", "http://localhost:4000/health"]
interval: 10s
timeout: 5s
retries: 5
# Cold start loads the full connector/tool catalog (~18.5k tools, ~90s).
# start_period must comfortably exceed that: during it, failing probes keep
# the container in "starting" (not "unhealthy"), so `docker compose --wait`
# in deploy-cloud.yml keeps waiting and the first healthy probe (~90s) turns
# the deploy green. Too short (was 30s) → the container flipped to unhealthy
# mid-boot and --wait aborted, failing a deploy that was actually coming up
# fine (false negative). A genuinely dead app never becomes healthy, so once
# start_period elapses it still goes unhealthy and the deploy fails for real.
#
# /health now also fails at HEALTH_HEAP_PERCENT of the heap limit, so the
# container reads unhealthy during a decline instead of right up to the
# abort. Docker does not restart an unhealthy container by itself; the
# backend's own guard exits, and this status is what a deploy and an
# outside probe see in the meantime.
start_period: 180s
restart: unless-stopped
logging:
driver: json-file
options:
max-size: "50m"
max-file: "5"
x-frontend: &frontend
image: helpcodeai/anythingmcp:latest
command: ["./start.sh", "frontend"]
expose:
- "3000"
mem_limit: ${FRONTEND_MEM_LIMIT:-1g}
# Only what the Next.js server reads at runtime — found by grepping
# `process.env.` under packages/frontend/src, not by copying the backend's
# list. Its backend rewrites are baked at build time (see the note in
# deploy/cloud/Caddyfile — Caddy routes every backend path itself), so no
# database URL and no secrets here: a process that does not need them
# should not hold them.
#
# GTM_ID and COOKIE_DOMAIN are read server-side to render the Tag Manager
# snippet and the consent banner; leave them out and both silently vanish.
environment:
- NODE_ENV=production
- NEXT_PUBLIC_API_URL=https://${DOMAIN}
- GTM_ID=${GTM_ID:-}
- COOKIE_DOMAIN=${COOKIE_DOMAIN:-}
# Sentry: server runtime, and handed to the browser by the root layout
# (components/sentry-config.tsx). Unset DSN = off, nothing on the page.
- SENTRY_DSN=${SENTRY_DSN_FRONTEND:-}
- SENTRY_ENVIRONMENT=${SENTRY_ENVIRONMENT:-production}
- SENTRY_TRACES_SAMPLE_RATE=${SENTRY_FRONTEND_TRACES_SAMPLE_RATE:-0}
- SENTRY_BROWSER_TRACES_SAMPLE_RATE=${SENTRY_BROWSER_TRACES_SAMPLE_RATE:-0.1}
- SENTRY_VERIFY_TOKEN=${SENTRY_VERIFY_TOKEN:-}
healthcheck:
# A static file under public/, served by the standalone server without
# touching the backend: this probe answers for the frontend alone.
test: ["CMD", "wget", "--no-verbose", "--tries=1", "--spider", "http://localhost:3000/logo.svg"]
interval: 10s
timeout: 5s
retries: 5
start_period: 60s
restart: unless-stopped
logging:
driver: json-file
options:
max-size: "50m"
max-file: "5"
services:
# Caddy Reverse Proxy — automatic HTTPS via Let's Encrypt
caddy:
image: caddy:2-alpine
container_name: amcp-cloud-caddy
ports:
- "80:80"
- "443:443"
environment:
- DOMAIN=${DOMAIN}
volumes:
- ./Caddyfile:/etc/caddy/Caddyfile:ro
- caddy_data:/data
- caddy_config:/config
# No depends_on the app: which app containers exist, and which colour
# Caddy sends traffic to, is release.sh's business (the colour is a word
# in /config/amcp-colour, in the caddy_config volume; see the Caddyfile).
# Caddy starting before them only means a few 502s while the host boots.
restart: unless-stopped
# Docker's json-file driver is unbounded by default. On 2026-09-18 MOTIS's
# debug log reached 32 GB and filled the droplet's 77 GB disk: every
# container went unhealthy, Postgres could not write, and a deploy failed
# at the scp step because nothing could be written to /tmp. Caps here, on
# every service, so one chatty container cannot do that again.
logging:
driver: json-file
options:
max-size: "50m"
max-file: "5"
backend-blue:
<<: *backend
container_name: amcp-cloud-backend-blue
profiles: [blue]
backend-green:
<<: *backend
container_name: amcp-cloud-backend-green
profiles: [green]
frontend-blue:
<<: *frontend
container_name: amcp-cloud-frontend-blue
profiles: [blue]
frontend-green:
<<: *frontend
container_name: amcp-cloud-frontend-green
profiles: [green]
# PostgreSQL Database
postgres:
image: postgres:17-alpine
container_name: amcp-cloud-postgres
volumes:
- postgres_data:/var/lib/postgresql/data
environment:
- POSTGRES_USER=amcp
- POSTGRES_PASSWORD=${POSTGRES_PASSWORD}
- POSTGRES_DB=anythingmcp
healthcheck:
test: ["CMD-SHELL", "pg_isready -U amcp -d anythingmcp"]
interval: 5s
timeout: 3s
retries: 5
restart: unless-stopped
logging:
driver: json-file
options:
max-size: "50m"
max-file: "5"
# Redis Cache — caching & rate limiting
redis:
image: redis:7-alpine
container_name: amcp-cloud-redis
volumes:
- redis_data:/data
healthcheck:
test: ["CMD", "redis-cli", "ping"]
interval: 5s
timeout: 3s
retries: 5
restart: unless-stopped
logging:
driver: json-file
options:
max-size: "50m"
max-file: "5"
# MOTIS — routing engine behind the Deutsche Bahn connector, fed by the open
# gtfs.de timetable (CC BY 4.0) plus its GTFS-RT feed for live delays.
# Replaces db-rest, which scraped bahn.de and died when Deutsche Bahn began
# blocking datacenter IPs (see deploy/motis/README.md).
# INTERNAL ONLY: deliberately NO `ports:` mapping → reachable solely as
# http://motis:8080 from the app over the amcp-cloud_default network, never
# from the internet. Do NOT add a `ports:` entry or a Caddy route.
motis:
# Custom image: official MOTIS + an entrypoint that downloads and imports
# the feeds and refreshes them weekly (deploy/motis). The tag is pinned so
# the image persists across deploys and is only rebuilt when it changes.
build:
context: ./deploy/motis
# The -N suffix is a build revision of OUR image, not of MOTIS: bump it
# whenever deploy/motis/{Dockerfile,entrypoint.sh,config.yml} changes.
# The deploy only builds when the pinned tag is missing locally, so an
# edit with the same tag ships nothing — -2 carries `--log-level info`.
image: anythingmcp-motis:2.11.3-2
container_name: amcp-cloud-motis
expose:
- "8080"
volumes:
- motis_data:/data
environment:
- MOTIS_REFRESH_DAYS=${MOTIS_REFRESH_DAYS:-7}
# info, not MOTIS's default debug — see deploy/motis/entrypoint.sh.
# Set to debug temporarily when diagnosing a routing problem, then
# put it back: debug writes ~1.5 GB an hour.
- MOTIS_LOG_LEVEL=${MOTIS_LOG_LEVEL:-info}
healthcheck:
# 127.0.0.1, not localhost: MOTIS binds IPv4 only and localhost resolves
# to ::1 first inside the container, so the probe was refused every time.
# /metrics, not /: MOTIS answers 404 on the root, which --spider treats as
# a failure. Between the two, the container reported unhealthy for its
# whole life while serving traffic perfectly.
test: ["CMD", "wget", "--quiet", "--tries=1", "--spider", "http://127.0.0.1:8080/metrics"]
interval: 30s
timeout: 5s
retries: 3
# First boot downloads ~12 MB of GTFS and imports it (seconds, not
# minutes); the margin is for a slow gtfs.de.
start_period: 180s
restart: unless-stopped
logging:
driver: json-file
options:
max-size: "50m"
max-file: "5"
volumes:
postgres_data:
redis_data:
motis_data:
caddy_data:
caddy_config:
diagnostics: