fix(agent): an environment hard stop persists nothing, and recall follows the status lineage - #661
Merged
Merged
Conversation
…e status keys by A recall matched a report sharing any one service, so a single-service run could take a report of a larger service set as its baseline, which get-status renders as another lineage than the run's own. Closes #658 Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ment Read off the deployment_environment_name label (deployment_environment as fallback) of the log streams and metric series the probe already fetches, so a Grafana mission detects the environment without a hand-composed query; the log streams answer alone when the series query fails. Closes #658 Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ompares the report's environment Setup step 5 says an environment hard stop (a divergence or step 4's split) writes no report and what the reply carries instead, and that on a replay the handoff's environment is the report's; the self-check's persistence item exempts the stop; the observation-report reference mirrors it; /odd-observe and /odd-verify relay the stop instead of running show when no stored path comes back. Closes #658 Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
This was referenced Sep 26, 2026
Merged
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
The mission side of environment targeting:
observe-run-reportreference says so;/odd-observeand/odd-verifyrelay the stop — both values, the queries, the remedies — instead of runningshowwhen no stored path comes back).environment=is the report's, and the configured one plays no part (one clause in Setup step 5).grafana-discover.pyprints each service's deployment environment, read from the log stream and metric series labels it already fetches (no extra query); several values print as a list. Verified live on Grafana Cloud.get-statuskeys by, so a run's baseline sits in the same status row as the run.Choices amended against the issue text are recorded on the issue: #658 (comment)
Tests
tests/skills/odd-memory/test_odd_recall.py: equal-set matching;tests/skills/observability-cli-guides/test_grafana_scripts.py: environment read off streams and series (both, logs only when series fail, several values, none) on masked fake-gcx replays.pytest tests/skills878 passed (re-run after rebasing on fix(mcp): refuse the environment writes that resolve to a wrong or hidden entry #660), ruff 0.16.4 check/format clean,check_stack_reference.pygreen,apm install --target claude+apm auditgreen.Review
A separate reviewer sub-agent checked the whole branch against the issue under the bound-review rule (recall callers,
/odd-verify, benchmark recall, stored reports, a live re-run of the discover script): no finding, ready to merge on the first round. The rebase onto #660 was conflict-free.Harness measurement
test-plugin-harnessing, copilot /openai/gpt-5.6-luna/ medium,/odd-observedrive mission on the local llms-benchmark stack, ABBA (base1, after1, after2, base2); every changed file was in the deployed path.Main's spread today over eight samples: turns 35–45, commands 30–51, tokens in 2.66–3.41 M, wall 310–431 s, preflight 43–75 s. No degradation on the four axes: turns, output and wall inside main's spread; input tokens and AIU slightly below it on both branch samples; the shorter preflight is paid back in the observation, so wall is level — not claimed as a gain on two samples. One branch sample made two
--helpcalls, all on scripts this branch does not touch (replay_benchmark.py,probe_services.py,grafana-context.py,gcx_local.py). No sample had a stored baseline for these services, so the recall change was not exercised by the mission (it is covered by the unit tests). Every sample wrote a report (4–5 anomalies, 2–5 gaps; main's range today 4–7 / 4–6).Closes #658
🤖 Generated with Claude Code