fix(skill): the report synthesis reads its baseline from the frontmatter, and three query scripts match their surface - #662
Merged
Conversation
…t a replay's rulings `new` records the recall's first line as a `baseline` field outside a replay (`none` without one), and `show`/`persist` read it instead of the first report filename on section 1's prose line; a report predating the field falls back to that line, where a value opening with none yields no baseline. A re-measure with no check table counts its baseline rulings. The body contract names section 7's Verdict column on a replay. Closes #659 Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
… one `azure-monitor-traces.py exemplars` took the slowest requests over the union of the operations, so one slow operation filled every slot, and had no p50 pick. It now partitions the slowest by operation (`--slow` per operation, default 1) and adds, per operation, the request nearest its p50. Fixtures re-captured live, masked. Closes #659 Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…gation `grafana-metrics.py labels --label` exited 1 on Grafana Cloud as soon as one series under the selector was aggregated by an Adaptive Metrics rule. It now retries once without those series (`__aggregation__=""`) and says so in its output. Closes #659 Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
`grafana-traces.py ops` withheld every operation's calls when the span metrics counter reset inside the window. It now reads the counter's increase() over the window for the reset operations, as `histogram` and `counter` already do, and marks the row. Closes #659 Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…anges Every observe-run agent of the field campaign tripped on `echo ====` as a separator between reads. The warning is replaced by the form to use: `sed -n 'A,Bp;C,Dp' <file>`, no separator line. Closes #659 Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
`new` took the `baseline` field from the recall's first line even when the mission named another report, which the agent then diffed against, so `show` and `persist` named the wrong baseline. `new --baseline <file>` records the named report as given (refused on a replay, whose baseline is `--verifies`, and when nothing stored carries that name); the recall fills the field only without it. Closes #659 Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ce-set rule Closes #659 Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Merged
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
odd_report.py newwrites abaseline:frontmatter field — the mission's named report (--baseline), else the recall's first match (same shared matcher asodd_recall.py, same-service-set rule), elsenone— andpersist/showread it (orverifieson a replay). Older reports keep the prose fallback, which no longer takes a.mdname from a line that opens withnone/no previous.| Check | Before | After | Verdict |table.azure-monitor-traces.py exemplarsranks per operation and returns a p50 exemplar per operation (--slow Nis now per operation, default 1).grafana-metrics.py labelsretries once without the series an Adaptive Metrics rule aggregates, with a note.grafana-traces.py opsfalls back toincrease()of the span-metrics counter after a reset instead of withholding the calls.sed -n 'A,Bp;C,Dp'form.Choices amended against the issue text are recorded on the issue: #659 (comment)
Tests
test_odd_report.py(baseline field from the recall /none/ named / refused cases / same service set only, older-report prose guard, re-measure headline, Verdict column read bysynthesis), rewritten exemplars test on re-captured masked fixtures, new Adaptive Metrics and reset-fallback tests for the grafana scripts.pytest tests/skills tests/hooks1387 passed after the rebase on fix(agent): an environment hard stop persists nothing, and recall follows the status lineage #661; ruff 0.16.4 check/format clean;check_stack_reference.pygreen;apm install --target claude+apm auditgreen.show,synthesisandcheckover all 18 stored reports are byte-identical to main's;get-status --fullidentical.Review
A separate reviewer sub-agent checked the whole branch under the bound-review rule:
new --baseline(0 commands on the normal path, one contract line) and a test.Harness measurement
test-plugin-harnessing, copilot /openai/gpt-5.6-luna/ medium,/odd-observedrive mission on the local llms-benchmark stack, ABBA, both lab branches rebuilt right before the chain; every changed file in the deployed path.Main's spread today over ten samples: turns 35–45, commands 29–54, tokens in 2.57–3.55 M, wall 310–431 s. No branch figure is above main's spread; several are below it (turns on both samples) — not claimed as a gain on two samples, this is a fix owing no degradation. 1 premium request per sample. Every sample persisted a report on the first try (branch:
baseline: noneand "no previous report" headline, consistent); no sample had a stored baseline, so the "vs baseline" path is covered by the unit tests only. An earlier chain was discarded (one main sample exited 1 at copilot shutdown, its replacement recalled a leftover report).Closes #659
🤖 Generated with Claude Code