You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Adds agentic-advisor to the linearb-ai marketplace: before writing code it grades how fragile the target is from LinearB health signals + local git history and holds the agent to a LOW/MEDIUM/HIGH effort level. Telemetry to the user's own LinearB org is on by default when LINEARB_API_TOKEN is set and turned off with LINEARB_TELEMETRY=0; LINEARB_API_URL supports on-prem/regional hosts.
Purpose: Implement the agentic-advisor plugin to size AI coding effort based on LinearB health signals and local git history.
Main changes:
Added a skill to grade repo fragility and mandate LOW/MEDIUM/HIGH coding postures
Implemented auto-trigger hooks for prompt analysis and first-edit blocking via agentic-advisor-trigger.sh
Added agentic-advisor-report.sh for telemetry reporting of effort decisions and token usage
Generated by LinearB AI and added by gitStream. AI-generated content may contain inaccuracies. Please verify before using.
💡 Tip: You can customize your AI Description using GuidelinesLearn how
Public release of LinearB's effort-grading skill: before writing code it reads
LinearB repo health (rework, incidents, unreviewed merges) + local git history
and holds the agent to a LOW/MEDIUM/HIGH effort level.
- install: /plugin install agentic-advisor@linearb-ai
- LINEARB_API_URL overrides the API host for on-prem/regional deployments
- telemetry is opt-in (LINEARB_API_TOKEN), reported as
agentic_advisor.effort_decision only to the user's own org
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The reason will be displayed to describe this comment to others. Learn more.
✨ PR Review
Solid, well-documented plugin addition with thoughtful fire-and-forget telemetry design. Found a security concern around how the API token is passed to curl (exposing it via process listings) and a potential false-positive bug in the dedup logic that uses unanchored substring matching against state files.
3 issues detected:
🔒 Security - Secret value is embedded in the curl command line, exposing it via process listings on shared systems. 🛠️
Details:The post() function passes the token directly as a literal value inside the curl command line (-H "x-api-key: $token"). On most systems this means the token is visible to any other local user via ps aux//proc/<pid>/cmdline for the lifetime of the subprocess, which contradicts the script's own stated design goal ("Secret-safe: the token is only ever passed as an env-var reference to curl").
🛠️ A suggested code correction is included in the review comments.
🐞 Bug - Substring matching on dedup keys can cause false positives and silently drop legitimate telemetry events. 🛠️
Details:Dedup checks (grep -q "$key" "$state") use the raw 16-char sha1-derived key as a grep pattern without anchoring (-x/^$) or using -F. Since the key is appended one-per-line, a key that happens to be a substring of another previously-stored key (or vice versa) could cause a false match, silently suppressing a legitimate new event from being reported.
🛠️ A suggested code correction is included in the review comments.
🔒 Security - Predictable temp file/directory paths without permission hardening can be exploited via symlink attacks on shared systems.
Details:Marker/state files are created at predictable paths under the shared ${TMPDIR:-/tmp}/agentic-advisor directory using plain mkdir -p/touch, without exclusive creation or permission hardening. On a multi-user system this is susceptible to a symlink/race attack where another local user pre-creates these paths pointing elsewhere, causing subsequent writes (mv -f, touch, append) to follow the symlink.
Generated by LinearB AI and added by gitStream. AI-generated content may contain inaccuracies. Please verify before using.
💡 Tip: You can customize your AI Review using GuidelinesLearn how
… accuracy
- telemetry stays on by default; LINEARB_TELEMETRY=0 turns reporting off,
and the README now says so plainly
- reporter and skill pass the API token to curl on stdin (-H @-), so it no
longer appears in curl's argv / process listings
- skill skips the familiarity signal when there is no git identity, and the
phase-2 combine rule keeps all three axes
- README: current phase-2 signal and 24h repo-health cache
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The reason will be displayed to describe this comment to others. Learn more.
✨ PR Review
Well-documented, defensive shell hooks for the new agentic-advisor plugin. The previously flagged token-exposure issue appears fixed (token is now passed to curl via stdin, not argv). The unanchored grep dedup and the predictable-path marker/state files without hardening (previously raised) remain present but unchanged, so they are not re-reported. One new concern found around session-id fallback causing shared state across unrelated sessions.
1 issues detected:
🐞 Bug - Falling back to a fixed "default" filename when session_id is missing can cause state to be shared/corrupted across unrelated concurrent sessions. 🛠️
Details:When session_id is empty, the dedup/state file falls back to a fixed name (${marker_dir}/default.reported, and similarly default.vseen). If multiple hook invocations across genuinely different sessions all lack a session_id (e.g. older Claude Code versions, or non-interactive invocations), they would share the same state file. This could cause one session's reported decisions to suppress another session's events as duplicates, or cause the baseline/graded bucket determination and verdict-count tracking to behave inconsistently across unrelated sessions.
🛠️ A suggested code correction is included in the review comments.
Generated by LinearB AI and added by gitStream. AI-generated content may contain inaccuracies. Please verify before using.
💡 Tip: You can customize your AI Review using GuidelinesLearn how
- README: task/file checks are never cached and run whenever the skill runs
- SKILL: the combine-rule note now says 'the axes', not 'the two axes'
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The reason will be displayed to describe this comment to others. Learn more.
✨ PR Review
Overall the hooks are carefully engineered with explicit fail-open/fail-safe behavior and detailed documentation. The previously flagged token-exposure issue appears fixed (token is now passed to curl via stdin instead of as a command-line argument). The dedup-pattern-matching and predictable-marker-path issues from the previous review still exist unchanged. One new maintainability concern is the unbounded growth of per-session marker/state files with no cleanup mechanism.
1 issues detected:
🧹 Maintainability - No cleanup/expiry mechanism exists for the per-session marker and state files, causing them to accumulate indefinitely.
Details:The marker directory accumulates .reported, .vseen, and (in the trigger script) .done/.held files per session indefinitely, with no cleanup, expiry, or rotation logic anywhere in the hook. Over many sessions on a shared machine or CI runner, ${TMPDIR:-/tmp}/agentic-advisor will grow without bound, consuming disk/inode resources and never being reclaimed.
Generated by LinearB AI and added by gitStream. AI-generated content may contain inaccuracies. Please verify before using.
💡 Tip: You can customize your AI Review using GuidelinesLearn how
… contact email
- skill: the cache write uses a per-process temp file ($f.$$.tmp), so two
concurrent sessions grading the same repo can't clobber each other's temp file
- reporter: holdout first_edit ignores docs/text files (same list as the trigger),
so a docs-only session is no baseline beacon and coding_tokens start at the first code edit
- contact email: support@linearb.io (LinearB's public support address)
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The External Custom Metrics API requires timestamp to be an ISO-8601 string, but this produces a numeric Unix-millisecond value and every payload inserts it with --argjson. Those requests will fail validation, while the detached curl suppresses the error, so the advertised telemetry can silently report nothing. Generate a UTC ISO-8601 timestamp and pass it with --arg ts at all three payload builders.
This issue also appears in the following locations of the same file:
Using the shared $f.tmp name does not make concurrent writers safe: two sessions can truncate/write or move the same temporary file, allowing one writer to alter the file another has already renamed. Use a unique temporary file in the destination directory (for example $f.$$ or mktemp "${f}.XXXXXX") before the atomic mv.
This issue also appears on line 112 of the same file.
This pull request has used all 3 of its automatic AI Reviews.
✨ Comment /gs review to run one now. No arguments needed. It uses the same settings as your automatic reviews: gitStream Settings in Managed Mode, your CM files in Self-Managed.
Re: Copilot's "Use ISO-8601 timestamps for custom metrics payloads" (agentic-advisor-report.sh:161): not an issue — no change needed.
The custom-metrics endpoint accepts both formats: PostCustomMetricModel declares timestamp: Union[str, int] and converts either to UTC in its validator, and the model's own example titled "Example with timestamp that is EPOCH" uses "timestamp": 1706622966150 — epoch milliseconds, exactly what the reporter sends (date +%s × 1000). This reporter has been sending epoch-ms since July and its events land in reported_metrics with correct timestamps, so requests are not failing validation.
Stop runs when the agent finishes the response, not when this intermediate verdict first appears. If the response continues through exploration and edits, process termination before completion still loses the decision, so this does not provide the claimed early, crash-resistant reporting. Emit from a lifecycle point after the verdict but before coding (for example, the first relevant PreToolUse) and retain Stop as a fallback.
Preliminary and corrected verdicts are both counted
This loop reports every verdict in the transcript, so the documented phase-2 correction emits both the preliminary grade and the corrected grade. Since consumers are told to count phase=decision rows, one task is double-counted and the stale grade remains in aggregates. Report only the latest decision per repo, or include a stable session/sequence identifier so consumers can select the corrected decision.
A Write into a newly created directory bypasses the hold because git -C "$dir" fails when the target directory does not exist. This is common when adding a new component/module, and it lets the first source write proceed without grading. Walk up to the nearest existing parent for both the repository check and marker key.
HTTP failures are misclassified as dormant repositories
curl treats HTTP 4xx/5xx responses as successful here, and the pipeline does not use pipefail. A failed measurements request can therefore collapse to no selected row and be classified by line 90 as a dormant/LOW repository instead of the required unavailable/MEDIUM fallback. Make HTTP failures produce a failing pipeline before interpreting an empty successful response as dormancy.
The reason will be displayed to describe this comment to others. Learn more.
✨ PR Review
The two new hook scripts and plugin manifest are well-documented and mostly defensive against failure modes already called out in prior review rounds (token now goes to curl via stdin). One new concern stands out: the JSON payload itself is still passed as a literal curl argument, which can leak contributor PII through the process table on shared machines.
3 issues detected:
🔒 Security - Sensitive user/repo data is passed as a visible process argument instead of via stdin/file. 🛠️
Details:The post() function passes the full JSON payload via -d "$1" as a literal curl command-line argument. While the token itself is now protected (passed on stdin via -H @-), the payload contains contributor_email, repo_url, branch, ticket, and session_name — all visible to any local user via ps aux or /proc/<pid>/cmdline for the duration of the backgrounded curl process.
🛠️ A suggested code correction is included in the review comments.
🐞 Bug - The `||` fallback pattern shares a single stdin read between two commands, so a failing-but-present primary tool silently starves the fallback of input. 🛠️
Details:sha1() { { shasum 2>/dev/null || sha1sum 2>/dev/null; } | cut -c1-16; } pipes stdin into whichever tool runs first. If shasum exists on PATH but fails for a reason other than "not found" (e.g. unsupported flag, permission issue), the || triggers sha1sum as a fallback, but the piped stdin has already been consumed by the failed shasum invocation, so sha1sum hashes empty input — silently producing a wrong/empty key used for dedup and bucket assignment.
🛠️ A suggested code correction is included in the review comments.
🐞 Bug - The pipeline does not exclude blank model values before selecting the most recent entry. 🛠️
Details:model is derived by taking the .message.model of the last matching assistant entry in the transcript and filtering out <synthetic>, but it does not filter out blank/empty values. If the most recent assistant turn lacks a model field (e.g. a tool-only turn), tail -1 can return an empty string even though earlier turns carried a valid model, silently dropping the model tag from the reported event.
🛠️ A suggested code correction is included in the review comments.
Generated by LinearB AI and added by gitStream. AI-generated content may contain inaccuracies. Please verify before using.
💡 Tip: You can customize your AI Review using GuidelinesLearn how
…, payload off argv
- skill: curl -f + pipefail; a failed call (HTTP 4xx/5xx, timeout) is 'unavailable',
never 'dormant'/'not found'; 401/403 = invalid/expired token -> MEDIUM + tell the user
- trigger: a Write into a not-yet-existing folder walks up to the nearest existing
parent, so the first file of a new module is still held for grading
- cache + hook state moved from shared /tmp to ~/.cache/<plugin> (umask 077); the skill
only trusts cache files you own; reporter skips if it has no private state dir
- reporter: payload sent via a private temp file (--data-binary @file), so neither the
token nor contributor data appears in curl's argv
- sha1 helper picks shasum/sha1sum up front; model tag skips blank values
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Token accounting silently requires python3, although the documented requirements only list jq, git, and curl. Without Python, parsed remains {}: decision events omit grading_tokens, graded SessionEnd exits without a tokens event, and baseline sessions emit nothing. Add a non-Python fallback or make Python an explicit requirement.
Any Jira-style key unconditionally marks the prompt as a code task. A question such as “summarize ABC-123” therefore triggers the mandatory skill/API sweep, despite the plugin README stating that questions are skipped. Require code-writing intent (or explicitly exclude question-only prompts) before taking the ticket branch.
Native edit tools bypass the editing gate
plugins/agentic-advisor/hooks/hooks.json:15
The backstop only intercepts Edit and Write. Native NotebookEdit/MultiEdit calls can therefore modify source before any effort grade, and the trigger/report parsers also ignore those tool names. Extend the matcher and path parsing (notebook_path for notebooks) so every native editing tool follows the same gate.
Bug PR threshold cannot trigger without bug metrics
These thresholds include bug-PR counts, but signal gathering only requests rework and unreviewed-merge metrics plus incidents; no step obtains bug PRs. Consequently the 5+ bug PRs HIGH condition can never fire and may under-grade the repository. Either gather this signal or remove it from the grading criteria.
The example gave '9 fix/revert commits' as the reason for HIGH, contradicting
the rule that fix/revert wording is a weak hint that never raises effort alone.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This \b expression is GNU-specific; the default BSD grep -E on macOS does not treat it as a word boundary, so the keyword gate never auto-fires there. Use explicit alphanumeric boundaries so prompt-time grading works on both supported platforms.
This records the verdict count before confirming that a repository URL and reportable verdict are available. For example, a session launched from a parent directory can print a verdict before its first edit; this Stop writes vseen, then exits because repo_url is empty. After the later edit makes repository resolution possible, the unchanged count triggers the fast exit and the decision is never reported. Persist vseen only after the Stop path has successfully processed the current verdicts.
Emit ISO-8601 timestamp instead of epoch milliseconds
The reported-metrics API requires timestamp to be an ISO-8601 string, but this produces a JSON number in epoch milliseconds. Every telemetry payload can therefore be rejected with a validation error, which is hidden by the detached curl. Generate a JSON-encoded ISO timestamp so the existing --argjson ts bindings emit a string.
Recognize all supported code-writing verbs in task matcher
This matcher omits common code-writing verbs explicitly advertised by the skill, such as “add”, “create”, “write”, and “modify”. For prompts like “add a component”, no start-of-task nudge is emitted; the backstop runs only at the first edit, after the agent may already have read and explored files, contrary to the mandatory-first-step behavior. Include those supported task verbs in the prompt gate.
Handle zero matching commits without treating probe as failed
grep -c prints 0 but exits with status 1 when this developer has no matching commits—the exact unfamiliar-code case this probe is meant to detect. The Bash tool will report the probe as failed (and pipefail preserves that failure), so the agent may discard the zero-commit signal instead of raising change-area effort. Count with awk, which exits successfully for zero matches while still letting a git log failure propagate.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Adds
agentic-advisorto thelinearb-aimarketplace: before writing code it grades how fragile the target is from LinearB health signals + local git history and holds the agent to a LOW/MEDIUM/HIGH effort level. Telemetry to the user's own LinearB org is on by default whenLINEARB_API_TOKENis set and turned off withLINEARB_TELEMETRY=0;LINEARB_API_URLsupports on-prem/regional hosts.🤖 Generated with Claude Code
✨ PR Description
Purpose: Implement the
agentic-advisorplugin to size AI coding effort based on LinearB health signals and local git history.Main changes:
agentic-advisor-trigger.shagentic-advisor-report.shfor telemetry reporting of effort decisions and token usageGenerated by LinearB AI and added by gitStream.
AI-generated content may contain inaccuracies. Please verify before using.
💡 Tip: You can customize your AI Description using Guidelines Learn how