You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
LINBEE-29986 | fix: hold once until the effort verdict is recorded (v1.1.0) - #3
At medium effort the model kept the verdict in thinking, so nothing was shown or reported (0/5 live runs); a one-time hold now asks for a printf of the verdict line, which the reporter also reads (5/5 live runs).
Purpose: Enhance the agentic-advisor plugin to enforce recording of effort verdicts before allowing code or file operations.
Main changes:
Updated PreToolUse matcher to include Read, Grep, Glob, Task, and Agent tools.
Implemented verdict_state to block tool use until a valid verdict line is recorded.
Expanded verdict detection to support both assistant text and printf Bash commands.
Generated by LinearB AI and added by gitStream. AI-generated content may contain inaccuracies. Please verify before using.
💡 Tip: You can customize your AI Description using GuidelinesLearn how
At medium effort the model keeps the verdict in thinking, so nothing was shown
or reported (0/5 live). The hold asks for a printf of the verdict line; the
reporter now reads it from that Bash call too (5/5 live).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The reason will be displayed to describe this comment to others. Learn more.
✨ PR Review
The changes implement a sensible one-time hold to force the model to externalize the effort verdict (either as text or as a printf'd Bash call) so it's not lost in "thinking". The dual-source regex matching (assistant text or Bash command) is applied consistently across the reporter and trigger scripts, but widening detection to arbitrary Bash tool_use commands introduces a false-positive risk, and the PreToolUse hook now fires on nearly every tool invocation.
3 issues detected:
🐞 Bug - Matching the verdict pattern against any Bash command text is too permissive and can misattribute unrelated commands as the recorded verdict. 🛠️
Details:The verdict detection now also scans the command field of any assistant-authored Bash tool_use call for the pattern "LinearB: ... (LOW|MEDIUM|HIGH) effort". If the model runs any unrelated Bash command whose text happens to contain this substring (e.g. grepping for the pattern, echoing documentation, or printing an example), it will be incorrectly treated as the actual verdict, corrupting the grading_tokens/duration telemetry.
🛠️ A suggested code correction is included in the review comments.
🐞 Bug - The broad pattern match can be satisfied by unrelated Bash commands that merely contain matching text, bypassing the intended one-time hold. 🛠️
Details:verdict_printed() greps the raw transcript for a line containing both "role":"assistant" and the verdict pattern, now including the possibility that the pattern appears inside an arbitrary Bash tool_use command string rather than the dedicated printf call. This mirrors the same false-positive risk as in agentic-advisor-report.sh and could cause the hold to be skipped even though no genuine verdict was recorded.
🛠️ A suggested code correction is included in the review comments.
🚀 Performance - Expanding the matcher increases the frequency of relatively expensive transcript scans on every tool call.
Details:The PreToolUse matcher was widened from Edit|Write to Edit|Write|Read|Grep|Glob|Task|Agent, meaning the trigger script (including the skill_ran/verdict_printed grep over the whole transcript) now runs on nearly every tool invocation during a session, not just file edits. This adds repeated transcript scanning overhead throughout the session.
Generated by LinearB AI and added by gitStream. AI-generated content may contain inaccuracies. Please verify before using.
💡 Tip: You can customize your AI Review using GuidelinesLearn how
The reason will be displayed to describe this comment to others. Learn more.
✨ PR Review
This revision extends the verdict regex to also match printf-ed Bash commands (addressing the earlier medium-effort false-negative issue) and introduces a one-time PreToolUse hold until a verdict is detected. The core mechanism is reasonable, but the new verdict-hold marker is not scoped per repository the way the existing hold/done markers are, which can cause inconsistent behavior across multi-repo sessions.
1 issues detected:
🐞 Bug - The verdict-hold marker only tracks session_id, not per-repo state, unlike equivalent markers elsewhere in the same script. 🛠️
Details:The new verdict-hold marker ($marker_dir/${session_id:-default}.vheld) is keyed only by session_id, unlike the existing .held/.done markers which also incorporate a per-repo key. If a single Claude session touches multiple repositories, the one-time verdict hold will only ever fire for the first repo encountered; subsequent repos where the skill ran but no verdict was printed will never trigger the hold, silently skipping the telemetry-recording nudge.
🛠️ A suggested code correction is included in the review comments.
Generated by LinearB AI and added by gitStream. AI-generated content may contain inaccuracies. Please verify before using.
💡 Tip: You can customize your AI Review using GuidelinesLearn how
…s a Bash verdict
Gate and reporter now ignore Bash commands that merely contain the pattern
(grep/echo), accept a printf on any line of a multi-line command, and key the
one-time hold per skill run so a second repo is gated too. Skill note allows
the printf fallback; README documents the task-complexity axis.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The reason will be displayed to describe this comment to others. Learn more.
✨ PR Review
The PR extends the verdict-detection pattern to also match Bash printf calls and adds a one-time PreToolUse hold gate; this fixes the medium-effort case where the verdict stayed in "thinking". The implementation is mostly consistent with the stated goal, but the new per-run gating counter has a fragility concern, and several previously raised issues (false-positive substring matching, widened PreToolUse matcher, verdict_printed grep) remain unaddressed in this diff.
1 issues detected:
🐞 Bug - Two independent, differently-implemented counting mechanisms (grep line count vs. jq index) are used to track the same logical "current skill run", risking marker/key misalignment. 🛠️
Details:The vdone marker's uniqueness key ("runs") is derived from a raw grep -cE count of lines in the transcript matching the "skill":"...agentic-advisor..." substring, rather than from an actual count of distinct Skill tool_use invocations. If this substring also appears in tool_result lines that echo back the input, or in any other transcript line unrelated to a genuine new invocation, the count can be inflated or inconsistent across calls. Since verdict_state() independently determines the "latest" Skill call via a different (index-based) mechanism, misalignment between the two counting approaches could cause the per-run vdone marker to be reused or skipped incorrectly, making the one-time hold fire for the wrong run or not fire at all for a genuine new skill invocation.
🛠️ A suggested code correction is included in the review comments.
Generated by LinearB AI and added by gitStream. AI-generated content may contain inaccuracies. Please verify before using.
💡 Tip: You can customize your AI Review using GuidelinesLearn how
This pull request has used all 3 of its automatic AI Reviews.
✨ Comment /gs review to run one now. No arguments needed. It uses the same settings as your automatic reviews: gitStream Settings in Managed Mode, your CM files in Self-Managed.
The new Bash path records any printf command containing the loose effort substring, not the exact verdict contract. For example, printf 'LinearB: demo HIGH effort' passes both checks and is emitted as a valid telemetry decision despite omitting the required delimiters, evidence, and plan. Use the same full-line validator as the trigger before assigning verdict.
Verdict gate accepts prose instead of the complete verdict format
The verdict gate does not enforce the documented verdict shape. This regex accepts any prose substring such as Prose only: LinearB: demo HIGH effort, so the code marks this run as done and permits tools even though no blockquote, evidence, or plan was recorded. Validate the complete > LinearB: <repo> — <LEVEL> effort (<evidence>) — <plan> line here, and keep the reporter's matching rules aligned.
Settled marker fails to short-circuit repeated transcript scans
The .vdone marker does not short-circuit the transcript scans: every subsequent Read/Grep/Glob/Task/Agent call still runs skill_ran over the full transcript and then scans it again to calculate runs before checking the marker. Because these newly matched tools are frequent and the transcript grows throughout a session, this adds increasing latency to every operation. Persist the processed transcript offset/run count, or otherwise inspect only content added since the settled marker.
Re: Copilot's "Previously missed" — Bash path accepts loose effort substrings (agentic-advisor-report.sh:231) and Verdict gate accepts prose instead of the complete verdict format (agentic-advisor-trigger.sh:49): intentionally not changing.
Strict validation would drop real grades. Real output varies: LOW verdicts often omit the — <plan> clause (optional for LOW per SKILL.md), some use - instead of —, some drop the leading > . A full-shape check would lose those decisions from telemetry, and the hold fires only once, so the session would continue without a recorded grade.
The false positive doesn't occur in practice. A match must be assistant-authored, after the latest skill call, as text or a printf. The SKILL.md examples are loaded content (not assistant) and the templates use <LOW|MEDIUM|HIGH>, which doesn't match; grep/echo commands containing the pattern are already rejected (tested).
Same matcher the reporter has used for text verdicts since July, with no bad rows observed in the data.
Gate and reporter stay aligned on the same lenient rule.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
At medium effort the model kept the verdict in thinking, so nothing was shown or reported (0/5 live runs); a one-time hold now asks for a
printfof the verdict line, which the reporter also reads (5/5 live runs).🤖 Generated with Claude Code
✨ PR Description
Purpose: Enhance the agentic-advisor plugin to enforce recording of effort verdicts before allowing code or file operations.
Main changes:
PreToolUsematcher to include Read, Grep, Glob, Task, and Agent tools.verdict_stateto block tool use until a valid verdict line is recorded.printfBash commands.Generated by LinearB AI and added by gitStream.
AI-generated content may contain inaccuracies. Please verify before using.
💡 Tip: You can customize your AI Description using Guidelines Learn how