Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
6 changes: 3 additions & 3 deletions .apm/agents/observe-run.agent.md
Original file line number Diff line number Diff line change
Expand Up @@ -570,9 +570,9 @@ subscript: `CD=customDimensions; echo
math expression: operand expected`, exit 1, so the CLI never runs, and
`"tostring($CD[1])"` prints `tostring(c)`, one character of the scalar,
where bash prints both as written — write `${CD}[...]`, or the literal
name; no word starting with `=` — zsh looks up a command named `===` for
`echo ====` and fails with `=== not found` where bash prints it — write
`echo "----- $f"`.
name; no separator line between reads — zsh runs `echo ====` as a
command lookup, `=== not found`, and aborts the chain — read several
ranges of one file as `sed -n 'A,Bp;C,Dp' <file>`.

Then query per signal from what came back — keyed by **operation**,
the smallest unit the service serves distinctly: on an HTTP server the
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -152,7 +152,7 @@ these scripts do.
```bash
python3 <Skills>/observability-cli-guides/scripts/azure-monitor-traces.py operations --app <app_insights_app> --service <svc> --from <start> --to <end>
python3 <Skills>/observability-cli-guides/scripts/azure-monitor-traces.py dependencies --app <app_insights_app> --service <svc> --since 30m
python3 <Skills>/observability-cli-guides/scripts/azure-monitor-traces.py exemplars --app <app_insights_app> --service <svc> --slow 3 --failed 3 --since 30m
python3 <Skills>/observability-cli-guides/scripts/azure-monitor-traces.py exemplars --app <app_insights_app> --service <svc> --since 30m
python3 <Skills>/observability-cli-guides/scripts/azure-monitor-traces.py trace <operation_Id> --app <app_insights_app> --from <start> --to <end>
python3 <Skills>/observability-cli-guides/scripts/azure-monitor-traces.py watch --app <app_insights_app> --identity <the run's User-Agent prefix> --from <dispatch instant> --state <scratch>/<slug>-watch.json --length <the manifest's scheduled length> --expect <its scheduled request count> [--to <deadline>] [--bin 30s] [--ended-after 4] [--settle auto] [--every 5s] [--max 8m] [--service <svc>]... [--dimension <customDimensions key>]... [--json]
```
Expand All @@ -163,7 +163,7 @@ optional `--to`, the deadline), `--json`; `operations`,
`dependencies`, `exemplars` and `watch` take `--service`; `operations` adds `--top`
(rows, default 20) and `--bin <duration>` (adds the request count,
failures and p95 per time bucket); `exemplars` adds `--operation <request
name>` (repeatable), `--slow N` (the slowest requests, default 3),
name>` (repeatable), `--slow N` (the slowest requests per operation, default 1),
`--failed N` (the newest failed requests, default 3); `trace` takes the
`operation_Id`; `watch` takes `--identity` (the prefix, matched with
`startswith`), `--state` (its state file), `--length` (the manifest's
Expand Down Expand Up @@ -193,13 +193,16 @@ non-empty one wins; default `user_agent.original` then
join leaves out - a client span whose request is outside the window).
Verified 2026-09-11: seven joined rows (the load generator's client
spans and the service's own outgoing calls), five in `all`.
- `exemplars` - Output: `slow` and `failed_requests`, each request with
- `exemplars` - Output: `p50` (per operation, the request nearest its
p50, with that `p50`), `slow` (ranked per operation) and
`failed_requests`, each request with
`timestamp`, `name`, `duration`, `resultCode`, `operation_Id`, `id`
and, from one union over the picked ids, its `dependencies`,
`exceptions` and `logs` (the `traces` rows of warning level and above).
Verified 2026-09-11: the slowest carried its `payment.authorize`
dependency, a failed one its failed `storage.delete` and the warning
line explaining it.
line explaining it. Verified 2026-09-26: five operations, a `p50`
and two `slow` each.
- `trace` - Output: `summary` (`root`, `duration_ms`, `spans`,
`failed_spans`, `logs`, `exceptions`, `services`) and `nodes`
(depth-first: `request`/`dependency` spans with `duration`, `success`,
Expand Down
8 changes: 5 additions & 3 deletions .apm/skills/observability-cli-guides/references/grafana.md
Original file line number Diff line number Diff line change
Expand Up @@ -192,7 +192,9 @@ lists the distinct metric names behind one or more `--match` selectors
not evidence of an absent metric). `labels` takes one `--match` selector
and a window: with `--label <name>` it lists that label's values across
the series the selector matches in the window (each with its series
count), without it the label *names* those series carry. `instant` and
count; on Grafana Cloud, the series an Adaptive Metrics rule aggregated
left out and said so, verified 2026-09-26), without it the label *names*
those series carry. `instant` and
`range` take a raw PromQL expression for anything the first four do not
shape — always with a selector or an aggregation, since a bare metric
name lists every series it has; `instant` takes `--at` (default now),
Expand Down Expand Up @@ -230,8 +232,8 @@ Six subcommands, the whole surface above (`--since <duration>` replaces

- `ops` — per operation of each service: rooted and containing trace
counts, span-level p50/p95/p99 and calls (span metrics, settled,
bucket-interpolated; `RESET` and calls withheld when the counter fell
inside the window; absence said), trace-level p50/p95/max over the
bucket-interpolated; `RESET` and calls from `increase()` when the
counter fell inside the window, verified 2026-09-26; absence said), trace-level p50/p95/max over the
rooted traces (integer ms), the worst containing trace. `--fetch` adds
each operation's p50, worst-rooted and worst-containing exemplar with
its summary. `--name` adds an operation the roots do not show.
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -3,7 +3,7 @@

azure-monitor-traces.py operations --app <app_insights_app> --service orders-api --from ... --to ...
azure-monitor-traces.py dependencies --app <app_insights_app> --service orders-api --since 30m
azure-monitor-traces.py exemplars --app <app_insights_app> --service orders-api --slow 3 --failed 3 --since 30m
azure-monitor-traces.py exemplars --app <app_insights_app> --service orders-api --since 30m
azure-monitor-traces.py trace <operation_Id> --app <app_insights_app> [--since 24h]
azure-monitor-traces.py watch --app <app_insights_app> --identity odd-bench/<name> --from <dispatch instant> --state <scratch>/<slug>-watch.json [--to <deadline>]

Expand All @@ -13,9 +13,10 @@
dependencies, exemplars and watch take --service (repeatable, a
cloud_RoleName; none = every service). operations adds --top (rows, default
20) and --bin (a duration: adds the request count, failures and p95 per
time bucket). exemplars adds --operation (the request name, repeatable),
--slow N (the slowest requests, default 3), --failed N (the newest failed
requests, default 3). trace takes the operation_Id. watch takes --identity
time bucket). exemplars picks, per operation, the request nearest its p50
and the slowest; it adds --operation (the request name, repeatable),
--slow N (the slowest requests per operation, default 1), --failed N (the
newest failed requests, default 3). trace takes the operation_Id. watch takes --identity
(the run's User-Agent prefix, matched with startswith), --state (its state
file; the same invocation again resumes it), --bin (default 30s),
--ended-after (empty closed bins that end a started run with no schedule
Expand Down Expand Up @@ -334,7 +335,19 @@ def cmd_exemplars(ns) -> tuple[int, dict]:
calls = [
ai_call(
ns.app,
f"requests {svc}{op}| top {ns.slow} by duration desc | project {EX_COLS}",
f"requests {svc}{op}| summarize p50=percentile(duration, 50) by cloud_RoleName, name"
f" | join kind=inner (requests {svc}{op}) on cloud_RoleName, name"
" | extend odd_gap = abs(duration - p50)"
" | summarize arg_min(odd_gap, timestamp, duration, success, resultCode, operation_Id, id, p50) by cloud_RoleName, name"
f" | order by p50 desc | project {EX_COLS}, p50",
frm,
to,
),
ai_call(
ns.app,
f"requests {svc}{op}| extend odd_key = strcat(cloud_RoleName, '/', name)"
f" | partition hint.strategy=native by odd_key (top {ns.slow} by duration desc)"
f" | order by duration desc | project {EX_COLS}",
frm,
to,
),
Expand All @@ -346,9 +359,10 @@ def cmd_exemplars(ns) -> tuple[int, dict]:
),
]
res = run_many(calls)
slow = ai_rows(res[0].data) if res[0].ok else []
failed = ai_rows(res[1].data) if res[1].ok else []
ids = list(dict.fromkeys([r["operation_Id"] for r in slow + failed]))
p50 = ai_rows(res[0].data) if res[0].ok else []
slow = ai_rows(res[1].data) if res[1].ok else []
failed = ai_rows(res[2].data) if res[2].ok else []
ids = list(dict.fromkeys([r["operation_Id"] for r in p50 + slow + failed]))
detail: dict[str, dict] = {
i: {"dependencies": [], "exceptions": [], "logs": []} for i in ids
}
Expand Down Expand Up @@ -392,10 +406,11 @@ def cmd_exemplars(ns) -> tuple[int, dict]:
"message": row["message"],
}
)
for r in slow + failed:
for r in p50 + slow + failed:
r.update(detail.get(r["operation_Id"], {}))
out = {
"window": [frm, to],
"p50": p50,
"slow": slow,
"failed_requests": failed,
"failed": failures(res),
Expand All @@ -420,7 +435,14 @@ def _render_ex(r: dict) -> list[str]:


def render_exemplars(o: dict) -> str:
out = [f"slowest requests, {o['window'][0]}..{o['window'][1]}:"]
out = [f"p50 request per operation, {o['window'][0]}..{o['window'][1]}:"]
for r in o["p50"]:
lines = _render_ex(r)
lines[0] += f" (operation p50 {r['p50']:.4g} ms)"
out += lines
if not o["p50"]:
out.append(" (none)")
out.append("slowest requests per operation:")
for r in o["slow"]:
out += _render_ex(r)
if not o["slow"]:
Expand Down Expand Up @@ -1039,7 +1061,7 @@ def main() -> int:
c = sub.add_parser("exemplars")
c.add_argument("--service", action="append")
c.add_argument("--operation", action="append")
c.add_argument("--slow", type=int, default=3)
c.add_argument("--slow", type=int, default=1)
c.add_argument("--failed", type=int, default=3)
d = sub.add_parser("trace")
d.add_argument("operation_id")
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -142,6 +142,19 @@ def cmd_names(ns) -> tuple[int, dict]:
}


AGGREGATED = "Can't query aggregated metric"


def _unaggregated(match: str) -> str:
"""The selector without the series an Adaptive Metrics rule aggregated."""
if match.rstrip().endswith("}"):
head = match.rstrip()[:-1]
return (
head + (", " if head.rstrip()[-1:] != "{" else "") + '__aggregation__=""}'
)
return match + '{__aggregation__=""}'


def cmd_labels(ns) -> tuple[int, dict]:
"""A label's values (--label) or the label names behind a selector."""
frm, to = resolve_window(ns)
Expand All @@ -152,6 +165,14 @@ def cmd_labels(ns) -> tuple[int, dict]:
at, win = _settled(frm, to, "0s")
expr = f"count by ({ns.label}) (last_over_time({ns.match}{win}))"
r = run_gcx(["metrics", "query", expr, "--time", at])
cmds, note = [r.command], None
if not r.ok and AGGREGATED in (r.error or ""):
# Grafana Cloud refuses the whole selector when one series under
# it is aggregated: count the others, and say so
expr = f"count by ({ns.label}) (last_over_time({_unaggregated(ns.match)}{win}))"
r = run_gcx(["metrics", "query", expr, "--time", at])
cmds.append(r.command)
note = "an Adaptive Metrics rule aggregates series under the selector: counted without them"
values: dict[str, int] = {}
for x in prom_result(r.data) if r.ok else []:
try:
Expand All @@ -163,7 +184,8 @@ def cmd_labels(ns) -> tuple[int, dict]:
"error": r.error,
"label": ns.label,
"values": dict(sorted(values.items(), key=lambda kv: -kv[1])),
"commands": [r.command],
"note": note,
"commands": cmds,
}
r = run_gcx(["metrics", "series", ns.match, "--from", frm, "--to", to])
names: dict[str, int] = {}
Expand Down Expand Up @@ -311,6 +333,8 @@ def render(o: dict) -> str:
out += [f"{v or '(unset)'} ({c} {unit})" for v, c in o["values"].items()] or [
"(no series)"
]
if o.get("note"):
out.append(o["note"])
elif isinstance(o.get("rows"), dict):
cols = [
c
Expand Down
31 changes: 27 additions & 4 deletions .apm/skills/observability-cli-guides/scripts/grafana-traces.py
Original file line number Diff line number Diff line change
Expand Up @@ -247,7 +247,30 @@ def _span_quantiles(keys, frm: str, to: str, settle: str) -> tuple[dict, list, b
else:
e["span_calls"] = None
out[key] = e
return out, results, present
# a reset withholds the subtraction: the counter's increase() over the
# window, which counts across resets, stands in (as histogram and counter do)
reset = [k for k in keys if out[k].get("span_calls_reset")]
extra = run_many(
[
[
"metrics",
"query",
f'sum(increase(traces_spanmetrics_calls_total{{service="{s}", span_name="{n}"}}{win}))',
"--time",
at,
]
for s, n in reset
]
)
for key, r in zip(reset, extra):
res = prom_result(r.data) if r.ok else []
try:
out[key]["span_calls_increase"] = (
round(float(res[0]["value"][1])) if res else None
)
except (KeyError, IndexError, TypeError, ValueError):
out[key]["span_calls_increase"] = None
return out, results + list(extra), present


def cmd_ops(ns) -> tuple[int, dict]:
Expand Down Expand Up @@ -1256,10 +1279,10 @@ def render(o: dict) -> str:
)
for k, e in o["operations"].items():
out.append(
f"{k:44s} {_f(e['rooted_traces']):>6} {_f(e['containing_traces']):>7} {_f(e.get('span_p50_ms')):>8} {_f(e.get('span_p95_ms')):>7} {_f(e.get('span_p99_ms')):>7} {_f(e.get('span_calls')):>6} | "
f"{k:44s} {_f(e['rooted_traces']):>6} {_f(e['containing_traces']):>7} {_f(e.get('span_p50_ms')):>8} {_f(e.get('span_p95_ms')):>7} {_f(e.get('span_p99_ms')):>7} {_f(e.get('span_calls') if e.get('span_calls') is not None else e.get('span_calls_increase')):>6} | "
f"{_f(e['trace_p50_ms']):>9} {_f(e['trace_p95_ms']):>7} {_f(e['trace_max_ms']):>7} | {e['worst_containing_trace']} ({_f(e['worst_containing_ms'])} ms)"
f"{' TRUNCATED' if e['truncated'] else ''}"
f"{' RESET inside the window (calls withheld)' if e.get('span_calls_reset') else ''}"
f"{' RESET inside the window (calls = increase())' if e.get('span_calls_reset') else ''}"
)
for s, nr in (o.get("never_rooted") or {}).items():
out.append(
Expand All @@ -1271,7 +1294,7 @@ def render(o: dict) -> str:
)
)
out.append(
" span p50/p95/p99 and calls: span metrics, settled (bucket-interpolated latency; calls = raw settled - raw start, withheld on a reset)"
" span p50/p95/p99 and calls: span metrics, settled (bucket-interpolated latency; calls = raw settled - raw start, increase() on a reset)"
+ (
""
if o.get("span_metrics_present")
Expand Down
9 changes: 6 additions & 3 deletions .apm/skills/odd-memory/references/observe-run-report.md
Original file line number Diff line number Diff line change
Expand Up @@ -15,7 +15,7 @@ python3 <this skill's directory>/scripts/odd_report.py new [--repo <observed rep
--service <name> [--service <name> ...] --stack <stack> --env <detected environment> \
--mode <drive|observe|post-hoc|verify|re-measure> \
--window <start>/<end> | --from <start> --to <end> --run-name <slug> \
[--verifies <baseline>] [--workload <text>] [--instance <service>=<identity> ...] \
[--verifies <baseline>] [--baseline <named report>] [--workload <text>] [--instance <service>=<identity> ...] \
[--process-restarted <true|false|service=true|false> ...] [--repository <value>] \
[--at <UTC instant>] [--no-revision] [--custom-stack]
python3 <this skill's directory>/scripts/odd_report.py check <path>
Expand All @@ -41,7 +41,9 @@ instrumentation` is the other kind's, stated in its own reference.
`-observe-<stack>` suffix in observe mode, the `verify-` and
`remeasure-` prefixes, the next free ordinal when the path is taken),
fills `date`, `revision`, `tree_anchor` and `repository` from the
repository itself, writes the frontmatter and the seven-section
repository itself, `baseline` outside a replay — the mission's named
baseline when it names one (`--baseline`), else the recall's first
line, or `none`, writes the frontmatter and the seven-section
skeleton — eight with `--custom-stack`, the flag a mission passes
when the handoff names a custom stack: the frontmatter then carries
`stack_friction: 0` and the skeleton the `## 8. Stack friction`
Expand Down Expand Up @@ -278,7 +280,8 @@ heading:
way: a check keyed more coarsely than the operations it rules can
never be re-read per operation later. In a verify or re-measure, this
table rules the baseline's **checks**, each under the key the baseline
gave it; a check key is never a finding id, and a check ruled here
gave it, as `| Check | Before | After | Verdict |` — a ruling
outside a Verdict column is one no script reads; a check key is never a finding id, and a check ruled here
never stands in for section 3's ruling on a baseline finding — the two
tables answer to different keys. A baseline check grouped more
coarsely than the operations it rules — by the route alone, its verbs
Expand Down
23 changes: 3 additions & 20 deletions .apm/skills/odd-memory/scripts/odd_recall.py
Original file line number Diff line number Diff line change
Expand Up @@ -58,6 +58,7 @@
check_report,
parse_value,
read_report,
recall_matches,
)
from odd_report import (
git_root as _git_root,
Expand Down Expand Up @@ -109,25 +110,7 @@ def check(report: dict, stored_names: set[str], root: Path) -> list[str]:
# --- the contract's checks, as the memory invariant applies them ---------------------


# --- the matching rules, as the references state them --------------------------------


def matches(report: dict, scope: dict) -> bool:
if "unreadable" in report:
return False
fm = report["frontmatter"]
if scope["stack"] and str(fm.get("stack")) != scope["stack"]:
return False
if report["kind"] == "instrumentation":
project = str(fm.get("project") or "")
target = scope["project"]
return not target or target == project or target.startswith(project + "/")
# the same service set: the lineage get-status keys a report by
if scope["services"] and set(scope["services"]) != set(as_list(fm.get("services"))):
return False
if scope["environment"] and str(fm.get("environment")) != scope["environment"]:
return False
return not scope["modes"] or str(fm.get("mode")) in scope["modes"]
# --- the matching rules, as the references state them: recall_matches, in odd_report --


def cell(value: Any) -> str:
Expand Down Expand Up @@ -185,7 +168,7 @@ def recall(root: Path, kind: str, scope: dict) -> tuple[list[str], list[str]]:
# every stored report is checked, matched or not: a flaw in the very
# field the scope matches on must never hide the report silently
problems = {r["name"]: check(r, stored, root) for r in reports}
matched = [r for r in reports if matches(r, scope)]
matched = [r for r in reports if recall_matches(r, scope)]
matched_names = {r["name"] for r in matched}
for r in matched:
out.append(line_of(r))
Expand Down
Loading