diff --git a/docs/before-you-arrive.md b/docs/before-you-arrive.md index f3ce0c1..cd9deed 100644 --- a/docs/before-you-arrive.md +++ b/docs/before-you-arrive.md @@ -16,8 +16,10 @@ Everyone also needs: - `git` and a terminal that runs `bash`. On Windows that is WSL or Git Bash, on every path: the checks in the labs are `bash` commands. - [`ork`](https://github.com/orca-ae/orca-cli) (the Orca CLI), v0.6.0 or newer: - `brew install orca-ae/tap/ork`. All three paths use it for the first MCP login - in Lab 3, and for the checks in every lab. Check + `brew install orca-ae/tap/ork`, or a + [release archive](https://github.com/orca-ae/orca-cli/releases) unpacked onto + your `PATH`. All three paths use it for the first MCP login in Lab 3, and for + the checks in every lab. Check `ork agent vaults credentials create --help` for `--oauth-issuer` and `--oauth-allow-issuer-mismatch`: a build with those flags handles StreamNative's same-domain proxy issuer aliases and dynamic client @@ -45,6 +47,11 @@ source .venv/bin/activate # Git Bash on Windows: source .venv/Scripts/act pip install -r requirements.txt ``` +`python3 --version` has to say 3.11 or newer. On macOS, Apple's own `python3` +is 3.9, and with it `pip` stops at +`No matching distribution found for runorca`. Install a newer Python and name it +in the second line, for example `python3.13 -m venv .venv`. + **TypeScript** ```bash diff --git a/labs/cloud/00-set-up.md b/labs/cloud/00-set-up.md index 6caa20e..81b57f7 100644 --- a/labs/cloud/00-set-up.md +++ b/labs/cloud/00-set-up.md @@ -45,6 +45,11 @@ source .venv/bin/activate # Git Bash on Windows: source .venv/Scripts/act pip install -r requirements.txt ``` +`python3 --version` has to say 3.11 or newer. On macOS, Apple's own `python3` +is 3.9, and with it `pip` stops at +`No matching distribution found for runorca`. Install a newer Python and name it +in the second line, for example `python3.13 -m venv .venv`. + **TypeScript** (Node.js 20 or newer) ```bash diff --git a/labs/cloud/troubleshooting.md b/labs/cloud/troubleshooting.md index f326055..cdd81dd 100644 --- a/labs/cloud/troubleshooting.md +++ b/labs/cloud/troubleshooting.md @@ -13,6 +13,7 @@ Still stuck after two tries? Raise your hand. | Symptom | Fix | |---|---| +| `pip install -r requirements.txt`: `No matching distribution found for runorca==0.3.0` | The `python3` that made your virtual environment is older than the course needs: on macOS, Apple's own is 3.9. Install Python 3.11 or newer, then make the environment again with it. In `python/`: `rm -rf .venv`, then the install commands from Lab 0 with that Python's name in place of `python3`, for example `python3.13 -m venv .venv`. | | Doctor: `Agent Engine HTTP 401/403` | The key was rejected. A key created before its permissions must be re-created: ask a facilitator. | | Doctor: `Kafka ... authentication` | `SN_SERVICE_ACCOUNT` must be the full principal, `@.auth.streamnative.cloud`; `SN_API_KEY` is the raw key. | | Doctor: `Kafka security.login_events: not found` | The topic is not there yet. Create it and load it: Lab 0, step 3. | diff --git a/labs/local/00-set-up.md b/labs/local/00-set-up.md index 6a84b32..db2d27a 100644 --- a/labs/local/00-set-up.md +++ b/labs/local/00-set-up.md @@ -10,7 +10,9 @@ part of the stack answers and the topic holds 246 logins. - **Docker** is running, with Compose v2 (`docker compose version`). - You have [`ork`](https://github.com/orca-ae/orca-cli) v0.6.0 or newer - (`brew install orca-ae/tap/ork`), and [`jq`](https://jqlang.org/download/). + (`brew install orca-ae/tap/ork`, or a + [release archive](https://github.com/orca-ae/orca-cli/releases) unpacked onto + your `PATH`), and [`jq`](https://jqlang.org/download/). - You have an **Anthropic API key**. - You cloned this repository, and installed **one** path: @@ -18,6 +20,11 @@ part of the stack answers and the topic holds 246 logins. |---|---| | `cd python && python3 -m venv .venv && source .venv/bin/activate && pip install -r requirements.txt` | `cd typescript && npm install` | + For Python, `python3 --version` has to say 3.11 or newer. On macOS, Apple's + own `python3` is 3.9, and with it `pip` stops at + `No matching distribution found for runorca`. Install a newer Python and name + it in the command, for example `python3.13 -m venv .venv`. + The CLI path needs only `ork` and `jq` for the labs, plus one of the two above for the doctor, the seeder, and the injector. - Keep two terminals open: one in your path's folder (`cli/`, `python/` or @@ -92,9 +99,9 @@ engine needs, usually 8080: see [Troubleshooting](troubleshooting.md). ### Check -The engine is up, has a provider key, and can reach the MCP server. The script -ends with these lines, and `--check` prints them again without starting -anything. +The engine is up, has a provider key that the model provider accepts, and can +reach the MCP server. The script ends with these lines, and `--check` prints +them again without starting anything. ```bash local/engine.sh --check @@ -103,12 +110,17 @@ local/engine.sh --check ```text PASS the AI Gateway is running PASS the gateway has a provider key +PASS the model provider accepts that key PASS the gateway allows the MCP host risingwave-mcp PASS the MCP server answers at http://risingwave-mcp:8000/mcp on the engine's network The Agent Engine is up and can reach your MCP server. ``` +A line that says `FAIL` prints its fix under it. While the third line fails, +your agent cannot answer. Usually the provider refused the key you exported: +export one that works and run `local/engine.sh` again. + ## Step 3: Write your `.env` Every script in this repository reads `.env` in the repository root. diff --git a/labs/local/01-hello-agent.md b/labs/local/01-hello-agent.md index 263d3e8..0487f72 100644 --- a/labs/local/01-hello-agent.md +++ b/labs/local/01-hello-agent.md @@ -35,8 +35,8 @@ As for what I can see right now — nothing live just yet! The next step connect ``` This is the first time your stack calls the model. If you see -`[error] ... (retrying)` lines and no reply, the model provider rejected the key -the engine started with: press Ctrl-C and see +`[error] ... (retrying)` lines and no reply, the model provider is refusing the +request: press Ctrl-C and see [Troubleshooting](troubleshooting.md#the-model-does-not-answer). ### Check diff --git a/labs/local/troubleshooting.md b/labs/local/troubleshooting.md index 674c71a..aca51f3 100644 --- a/labs/local/troubleshooting.md +++ b/labs/local/troubleshooting.md @@ -3,7 +3,7 @@ Two commands tell you what is wrong. Run both first. ```bash -local/engine.sh --check # the Agent Engine, and its link to the MCP server +local/engine.sh --check # the Agent Engine, its provider key, and its link to the MCP server (cd python && .venv/bin/python doctor.py) # or: npm --prefix typescript run doctor ``` @@ -15,7 +15,9 @@ Each failed line prints its fix. |---|---| | `docker compose ... up` fails with a port already in use | Another program has one of the ports the stack publishes on `127.0.0.1`: 29092 (Kafka), 18081 (schema registry), 4566 and 5691 (RisingWave), 8000 (MCP). Stop that program, run `local/down.sh`, then run the `up` command again. Running it again without `local/down.sh` is not enough: Docker brings the container whose port was taken back without its network. The ports are fixed: the broker tells its clients to come back to `127.0.0.1:29092`, and `local/write-env.sh` writes these ports into `.env`. | | `local/engine.sh` fails because a port is taken (`Bind for 0.0.0.0:8080 failed: port is already allocated`) | Another program has a port the engine publishes on `127.0.0.1`: 8080 (the registry) or 18082. If it is a container, `docker ps` shows which: look for `:8080->` under PORTS. Stop that program and run `local/engine.sh` again. If the port is 8080 and you want to keep that program running, move the registry instead: `export ORCA_LOCAL_REGISTRY_PORT=18080`, run `local/engine.sh` again, then `local/write-env.sh` so `.env` has the new address. Export it in every terminal you run `local/engine.sh` from: a run without it goes back to 8080. | +| `pip install -r requirements.txt`: `No matching distribution found for runorca==0.3.0` | The `python3` that made your virtual environment is older than the course needs: on macOS, Apple's own is 3.9. Install Python 3.11 or newer, then make the environment again with it. In `python/`: `rm -rf .venv`, then the install command from Lab 0 with that Python's name in place of `python3`, for example `python3.13 -m venv .venv`. | | `local/engine.sh`: `ANTHROPIC_API_KEY is not set in this shell` | `export ANTHROPIC_API_KEY=` in the terminal where you run the script. The engine reads the key only when it starts. | +| `local/engine.sh`: `FAIL the model provider accepts that key` | The fix under that line says why. `the provider answered 401` (or `403`): the provider refuses the key the engine started with. `export ANTHROPIC_API_KEY=`, then `local/engine.sh`. `the engine cannot reach https://api.anthropic.com`: the engine has no way out to the provider. Check your network, then `local/engine.sh --check`. | | `local/engine.sh`: `bootstrap refused: an organization already exists` | The engine's volumes exist but its keys in `.lab/ork` are gone. Start over: `local/down.sh --reset`, then Lab 0. | | CLI path: `.venv/bin/python: No such file or directory` | The Python path is not installed. If you installed the TypeScript path, use the `npm` command the lab gives beside the Python one. Otherwise install one of the two: Lab 0, "Before you start". | | Doctor: `Agent Engine HTTP 401` | The key in `.env` is not the running engine's key. Run `local/write-env.sh`. If it still fails, the engine's volumes and keys are out of step: `local/down.sh --reset`, then Lab 0. | @@ -52,9 +54,11 @@ You do not have to wait for it: press Ctrl-C. That stops your script, not the engine. The engine keeps retrying that turn for the rest of the three minutes, and a turn you start meanwhile waits behind it. The usual cause is the key: -1. Check that the key works at all, for example in the - [Anthropic console](https://console.anthropic.com/). -2. Export the working key and restart the engine, which reads the key only at +1. Ask the provider about it: `local/engine.sh --check`. If + `the model provider accepts that key` says `FAIL`, the fix under it says + why. Most often the provider refuses the key the engine started with, for + example one that was revoked since Lab 0. +2. Export a key that works and restart the engine, which reads the key only at start: ```bash @@ -64,8 +68,8 @@ and a turn you start meanwhile waits behind it. The usual cause is the key: 3. Run the lab script again. -If the key is fine, check that `ORCA_MODEL` in `.env` is a model your key can -use. +If that line says `PASS`, the key is fine: check that `ORCA_MODEL` in `.env` is +a model your key can use. ## Start over diff --git a/local/down.sh b/local/down.sh index 5e9fe6b..0bc280b 100755 --- a/local/down.sh +++ b/local/down.sh @@ -29,9 +29,14 @@ if [ -n "$mcp" ]; then fi # The engine. `down` by project name also removes the gateway, and needs no file. +# Without --reset, `rm --volumes` goes first: `down` alone keeps the unnamed +# volumes that some images declare, and the next start makes new ones beside +# them. Your data is in named volumes, which `rm --volumes` does not touch. It +# is quiet and may fail: `down` stops whatever it left. if $reset; then docker compose --progress quiet --project-name "$project" down -v --remove-orphans else + docker compose --progress quiet --project-name "$project" rm --stop --force --volumes >/dev/null 2>&1 || true docker compose --progress quiet --project-name "$project" down --remove-orphans fi @@ -39,6 +44,7 @@ fi if $reset; then streaming down -v --remove-orphans else + streaming rm --stop --force --volumes >/dev/null 2>&1 || true streaming down --remove-orphans fi diff --git a/local/engine.sh b/local/engine.sh index 0fce8bb..589fa84 100755 --- a/local/engine.sh +++ b/local/engine.sh @@ -4,7 +4,7 @@ # # export ANTHROPIC_API_KEY=... # the engine reads your provider key when it starts # local/engine.sh # start (or restart) the engine, then link it -# local/engine.sh --check # only check the link +# local/engine.sh --check # only run the checks # # What it runs, in order: # @@ -21,6 +21,10 @@ # restarts the gateway, and attaches the MCP server's container to the # engine's Docker network under that name. # +# 3. The checks, which `--check` runs alone. One of them asks the model +# provider whether it accepts your key. The harness makes that request, so +# the key stays in the engine; it lists models, which costs nothing. +# # Safe to run again at any time. Run it again after anything restarts the engine. set -euo pipefail # shellcheck source=lib.sh @@ -38,6 +42,22 @@ curl_on() { # curl_on docker run --rm --network "$1" --entrypoint curl "$CURL_IMAGE" "${@:2}" } +# Ask the model provider about the engine's key: the HTTP status of a request +# that lists models, or "unreachable". `ork local start` hands the harness and +# the gateway the same key, and of the two images only the harness's has an HTTP +# client (node), so the harness asks. +provider_status() { + local harness + harness=$(engine_container harness) + [ -n "$harness" ] || return 0 + docker exec -e PROVIDER_URL="$PROVIDER_URL" "$harness" node -e ' + fetch(process.env.PROVIDER_URL + "/v1/models?limit=1", { + headers: { "x-api-key": process.env.ANTHROPIC_API_KEY, "anthropic-version": "2023-06-01" }, + signal: AbortSignal.timeout(15000), + }).then((response) => console.log(response.status), () => console.log("unreachable")); + ' 2>/dev/null || true +} + start_engine() { command -v ork >/dev/null || die "ork (the Orca CLI) is not installed. See labs/local/00-set-up.md." [ -n "${ANTHROPIC_API_KEY:-}" ] || @@ -49,10 +69,12 @@ start_engine() { # Replace what is not running. A container whose port could not be bound # (another program had it) stays cut off from its network: Docker starts it # with loopback only from then on, even once the port is free (seen with - # Docker Engine 29.2). The engine's data is in volumes, so nothing is lost. + # Docker Engine 29.2). The engine's data is in named volumes, so nothing is + # lost: `--volumes` takes only the unnamed ones an image declares, which would + # otherwise be left behind at every start. local name while IFS= read -r name; do - [ -z "$name" ] || docker rm "$name" >/dev/null + [ -z "$name" ] || docker rm --volumes "$name" >/dev/null done < <(engine_stopped_containers) ork local --data-dir "$ORK_DIR" start --with-gateway } @@ -82,7 +104,7 @@ link() { } check() { - local gateway network mcp + local gateway network mcp fix gateway=$(engine_container ai-gateway) mcp=$(mcp_container) @@ -94,6 +116,14 @@ check() { if docker inspect -f '{{range .Config.Env}}{{println .}}{{end}}' "$gateway" | grep -q '^ANTHROPIC_API_KEY=.'; then pass "the gateway has a provider key" + # A key can be there and still be refused. Without this, the first sign is + # an agent that retries its first turn for three minutes. + fix=$(provider_fix "$(provider_status)") + if [ -z "$fix" ]; then + pass "the model provider accepts that key" + else + fail "the model provider accepts that key" "$fix" + fi else fail "the gateway has a provider key" "export ANTHROPIC_API_KEY=, then local/engine.sh" fi diff --git a/local/lib.sh b/local/lib.sh index b473131..01fe3e5 100644 --- a/local/lib.sh +++ b/local/lib.sh @@ -11,6 +11,8 @@ MCP_HOST=risingwave-mcp # The object store's image ships a curl. Both stacks already use the image. # shellcheck disable=SC2034 # read by engine.sh CURL_IMAGE=rustfs/rustfs:1.0.0 +# Where the gateway that `ork local` configures sends your agent's model calls. +PROVIDER_URL=https://api.anthropic.com die() { printf '\n%s\n' "$1" >&2 @@ -82,3 +84,21 @@ This script was written for ork v0.6.0; your ork may write a different gateway c cat "$file.tmp" >"$file" rm -f "$file.tmp" } + +# What the model provider's answer about the engine's key calls for: nothing +# when it accepts the key, otherwise the fix. +provider_fix() { # provider_fix + case "$1" in + 200) ;; + 401 | 403) + printf 'the provider answered %s. export ANTHROPIC_API_KEY=, then local/engine.sh' "$1" + ;; + unreachable) + printf 'the engine cannot reach %s. Check your network, then local/engine.sh --check' "$PROVIDER_URL" + ;; + [0-9][0-9][0-9]) + printf 'the provider answered %s, which is not about your key. Try again: local/engine.sh --check' "$1" + ;; + *) printf 'the harness could not ask the provider. Start the engine again: local/engine.sh' ;; + esac +} diff --git a/local/tests/run.sh b/local/tests/run.sh index f7660d5..d2e3da0 100755 --- a/local/tests/run.sh +++ b/local/tests/run.sh @@ -14,7 +14,8 @@ WORK=$(cd "$WORK" && pwd) # as the scripts see it: a TMPDIR ending in "/" leav trap 'rm -rf "$WORK"' EXIT # A `docker` that knows which host port the registry is published on, which -# containers exist, and remembers what it was asked. +# containers exist, what the gateway holds and what the model provider answers, +# and remembers what it was asked. mkdir -p "$WORK/bin" cat >"$WORK/bin/docker" <<'EOF' #!/usr/bin/env bash @@ -43,7 +44,22 @@ case "$1" in [ -f "$FAKE_DOCKER_DIR/registry-port" ] && echo "127.0.0.1:$(cat "$FAKE_DOCKER_DIR/registry-port")" ;; rm) [ $# -ge 2 ] || exit 1 ;; # like docker, it wants at least one container - compose) ;; + inspect) + # The gateway's environment, or the network it is on, whichever is asked for. + case "$*" in + *.Config.Env*) cat "$FAKE_DOCKER_DIR/gateway-env" ;; + *) echo ork-local_default ;; + esac + ;; + exec) cat "$FAKE_DOCKER_DIR/provider-status" 2>/dev/null ;; # the harness asks the provider + run) ;; # curl on the engine's network: the MCP server answers + compose) + # `rm` can be made to fail, noisily: stopping must not depend on it. + if [[ " $* " == *" rm "* && -f "$FAKE_DOCKER_DIR/compose-rm-fails" ]]; then + echo "No stopped containers" >&2 + exit 1 + fi + ;; *) exit 1 ;; esac EOF @@ -216,6 +232,7 @@ test_engine_replaces_its_containers_that_are_not_running() { start_engine check "removes a container that never started, before the engine starts" removed_before_the_start ork-registry-1 check "removes a container that has stopped, before the engine starts" removed_before_the_start ork-migrate-1 + check "takes each one's unnamed volumes with it" grep -qx -- "rm --volumes ork-migrate-1" "$FAKE_DOCKER_DIR/calls" check "leaves a running container alone" not removed ork-harness-1 check "leaves other projects' containers alone" not removed someone-elses-container check "then starts the engine" started @@ -230,6 +247,128 @@ test_engine_starts_when_nothing_has_stopped() { check "and starts the engine" started } +# ------------------------------------------------------------------ stopping -- + +asked() { grep -qE -- "$1" "$FAKE_DOCKER_DIR/calls"; } # asked : docker was asked this + +test_stopping_keeps_your_data_and_leaves_no_unnamed_volumes() { + # `down` alone keeps the unnamed volumes some images declare. Every stop and + # start left four more behind. + fresh_repo down + engine_started + run down.sh + check "stopping succeeds" [ "$STATUS" -eq 0 ] + check "the engine's containers go with their unnamed volumes" \ + asked '^compose .*--project-name ork-local-[0-9a-f]{8} rm --stop --force --volumes$' + check "so do the streaming stack's" \ + asked '^compose .*--project-name hello-data-agent .*rm --stop --force --volumes$' + check "both stacks are brought down" [ "$(grep -cE '^compose .* down --remove-orphans$' "$FAKE_DOCKER_DIR/calls")" -eq 2 ] + check "no named volume is deleted" not asked ' down -v' + check "says your data is kept" out_has "Your data is kept" +} + +test_stopping_does_not_depend_on_removing_unnamed_volumes() { + fresh_repo down-rm-fails + engine_started + touch "$FAKE_DOCKER_DIR/compose-rm-fails" + run down.sh + check "stopping succeeds when that step fails" [ "$STATUS" -eq 0 ] + check "both stacks are still brought down" [ "$(grep -cE '^compose .* down --remove-orphans$' "$FAKE_DOCKER_DIR/calls")" -eq 2 ] + check "and that step's messages are not shown" [ ! -s "$R/err" ] +} + +test_reset_deletes_the_volumes_with_the_stacks() { + fresh_repo down-reset + engine_started + printf 'TUTORIAL_STACK=local\n' >"$R/.env" + run down.sh --reset + check "reset succeeds" [ "$STATUS" -eq 0 ] + check "the engine's volumes are deleted" asked '^compose .*--project-name ork-local-[0-9a-f]{8} down -v --remove-orphans$' + check "the streaming stack's too" asked '^compose .*--project-name hello-data-agent .*down -v --remove-orphans$' + check "the engine's keys go with its volumes" [ ! -e "$R/.lab/ork" ] + check "and the local .env" [ ! -e "$R/.env" ] +} + +# ------------------------------------------------------ checking the engine -- + +PROVIDER_KEY=sk-ant-test-key-0123456789 + +engine_running() { # an engine that is up and linked, and holds a provider key + engine_started + printf 'PATH=/usr/bin\nANTHROPIC_API_KEY=%s\n' "$PROVIDER_KEY" >"$FAKE_DOCKER_DIR/gateway-env" + printf " allowed_private_hosts: ['risingwave-mcp']\n" >"$R/.lab/ork/gateway.yaml" +} + +provider_answers() { # what the harness hears back when it asks about the key + printf '%s\n' "$1" >"$FAKE_DOCKER_DIR/provider-status" +} + +test_check_passes_when_the_provider_accepts_the_key() { + fresh_repo provider-accepts + engine_running + provider_answers 200 + run engine.sh --check + check "the check succeeds" [ "$STATUS" -eq 0 ] + check "says the provider accepts the key" out_has "PASS the model provider accepts that key" + check "and that the engine is up" out_has "The Agent Engine is up" + check "the key is on no command line" not grep -qF "$PROVIDER_KEY" "$FAKE_DOCKER_DIR/calls" + check "and in nothing the script prints" not grep -qF "$PROVIDER_KEY" "$R/out" "$R/err" +} + +test_check_fails_when_the_provider_refuses_the_key() { + # Before this check, a refused key first showed in Lab 1, as three minutes of retries. + fresh_repo provider-refuses + engine_running + provider_answers 401 + run engine.sh --check + check "exits 1 on a refused key" [ "$STATUS" -eq 1 ] + check "names the check that failed" out_has "FAIL the model provider accepts that key" + check "says what the provider answered" out_has "the provider answered 401" + check "says how to give the engine another key" out_has "export ANTHROPIC_API_KEY=, then local/engine.sh" + check "still runs the checks that follow" out_has "PASS the gateway allows the MCP host risingwave-mcp" +} + +test_check_tells_a_network_problem_from_a_refused_key() { + fresh_repo provider-unreachable + engine_running + provider_answers unreachable + run engine.sh --check + check "exits 1 when the provider cannot be reached" [ "$STATUS" -eq 1 ] + check "blames the network" out_has "the engine cannot reach https://api.anthropic.com" + check "does not say to change the key" not out_has "export ANTHROPIC_API_KEY" +} + +test_check_does_not_blame_the_key_for_a_provider_outage() { + fresh_repo provider-outage + engine_running + provider_answers 529 + run engine.sh --check + check "exits 1 while the provider is failing" [ "$STATUS" -eq 1 ] + check "says the answer is not about the key" out_has "the provider answered 529, which is not about your key" + check "does not say to change the key" not out_has "export ANTHROPIC_API_KEY" +} + +test_check_says_so_when_the_harness_cannot_ask() { + fresh_repo provider-not-asked + engine_running # and no answer: `docker exec` fails, as on a harness that is down + run engine.sh --check + check "exits 1 when nothing could ask" [ "$STATUS" -eq 1 ] + check "names the check that failed" out_has "FAIL the model provider accepts that key" + check "says that nothing could ask" out_has "the harness could not ask the provider" + check "and to start the engine again" out_has "Start the engine again: local/engine.sh" +} + +test_check_asks_the_provider_nothing_without_a_key() { + fresh_repo provider-no-key + engine_running + printf 'PATH=/usr/bin\nANTHROPIC_API_KEY=\n' >"$FAKE_DOCKER_DIR/gateway-env" + provider_answers 401 + run engine.sh --check + check "fails on the missing key" out_has "FAIL the gateway has a provider key" + check "prints one failure for it, not two" not out_has "the model provider accepts that key" + check "and makes no request" not grep -q '^exec ' "$FAKE_DOCKER_DIR/calls" +} + # ------------------------------------------------------- the gateway patch -- gateway_yaml() { # the two lines of `ork local`'s gateway.yaml that matter here