Skip to content

fix(agent): an environment hard stop persists nothing, and recall follows the status lineage - #661

Merged
using-system merged 3 commits into
mainfrom
fix/observe-run-environment-stop
Sep 26, 2026
Merged

using-system merged 3 commits into
mainfrom
fix/observe-run-environment-stop

Conversation

@using-system

Copy link
Copy Markdown
Owner

Summary

The mission side of environment targeting:

  • An environment hard stop persists nothing (observe-run Setup step 5 covers its own divergence and step 4's split; the self-check's persistence item is exempted; the observe-run-report reference says so; /odd-observe and /odd-verify relay the stop — both values, the queries, the remedies — instead of running show when no stored path comes back).
  • On a replay, the handoff's environment= is the report's, and the configured one plays no part (one clause in Setup step 5).
  • grafana-discover.py prints each service's deployment environment, read from the log stream and metric series labels it already fetches (no extra query); several values print as a list. Verified live on Grafana Cloud.
  • Recall matches the same service set, the lineage get-status keys by, so a run's baseline sits in the same status row as the run.

Choices amended against the issue text are recorded on the issue: #658 (comment)

Tests

  • tests/skills/odd-memory/test_odd_recall.py: equal-set matching; tests/skills/observability-cli-guides/test_grafana_scripts.py: environment read off streams and series (both, logs only when series fail, several values, none) on masked fake-gcx replays.
  • pytest tests/skills 878 passed (re-run after rebasing on fix(mcp): refuse the environment writes that resolve to a wrong or hidden entry #660), ruff 0.16.4 check/format clean, check_stack_reference.py green, apm install --target claude + apm audit green.

Review

A separate reviewer sub-agent checked the whole branch against the issue under the bound-review rule (recall callers, /odd-verify, benchmark recall, stored reports, a live re-run of the discover script): no finding, ready to merge on the first round. The rebase onto #660 was conflict-free.

Harness measurement

test-plugin-harnessing, copilot / openai/gpt-5.6-luna / medium, /odd-observe drive mission on the local llms-benchmark stack, ABBA (base1, after1, after2, base2); every changed file was in the deployed path.

Side (mean of 2) Turns Commands Tokens in Tokens out nanoAIU Wall Preflight Observation
main 39.5 38.5 2.81 M 15.7 k 11.53 G 341 s 43 s 169 s
branch 38 47.5 2.37 M 14.8 k 10.35 G 365 s 23.5 s 202 s

Main's spread today over eight samples: turns 35–45, commands 30–51, tokens in 2.66–3.41 M, wall 310–431 s, preflight 43–75 s. No degradation on the four axes: turns, output and wall inside main's spread; input tokens and AIU slightly below it on both branch samples; the shorter preflight is paid back in the observation, so wall is level — not claimed as a gain on two samples. One branch sample made two --help calls, all on scripts this branch does not touch (replay_benchmark.py, probe_services.py, grafana-context.py, gcx_local.py). No sample had a stored baseline for these services, so the recall change was not exercised by the mission (it is covered by the unit tests). Every sample wrote a report (4–5 anomalies, 2–5 gaps; main's range today 4–7 / 4–6).

Closes #658

🤖 Generated with Claude Code

using-system and others added 3 commits September 26, 2026 21:24
…e status keys by

A recall matched a report sharing any one service, so a single-service
run could take a report of a larger service set as its baseline, which
get-status renders as another lineage than the run's own.

Closes #658

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ment

Read off the deployment_environment_name label (deployment_environment
as fallback) of the log streams and metric series the probe already
fetches, so a Grafana mission detects the environment without a
hand-composed query; the log streams answer alone when the series
query fails.

Closes #658

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ompares the report's environment

Setup step 5 says an environment hard stop (a divergence or step 4's
split) writes no report and what the reply carries instead, and that on
a replay the handoff's environment is the report's; the self-check's
persistence item exempts the stop; the observation-report reference
mirrors it; /odd-observe and /odd-verify relay the stop instead of
running show when no stored path comes back.

Closes #658

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

1 participant