Repository navigation
bumped v0.70 - #7
Conversation
PRAgent ReviewVerdict: Reviewed 25 changed file(s); no blocking issues found. Final Pre-Merge SummaryChanged files: CHANGELOG.md (modified, +43/-0); README.md (modified, +133/-6); docs/docs/index.html (modified, +146/-19); docs/index.html (modified, +11/-6); pyproject.toml (modified, +1/-1); src/agentic_evals/init.py (modified, +7/-1); src/agentic_evals/evaluators.py (modified, +3/-1); src/agentic_evals/gate.py (modified, +35/-11); src/agentic_evals/models.py (modified, +69/-0); src/agentic_evals/runner.py (modified, +122/-3); 15 more file(s). Impact On Existing CodeNo merge conflicts or blocking findings were detected. The owner should still confirm tests and product expectations before merging. Owner Permission RequiredOwner approval required: approve merge if the summarized changes match intent. Quality GateQuality gate retained all 0 actionable finding(s).
Code ReviewThe PR correctly adds new features for reporting breakdowns by tags and metrics, introduces tool argument and order checks, adds a rubric builder, and improves gate thresholds. Tests and docs are comprehensive and aligned. No correctness bugs or API contract breaks are found in the diff. Security ReviewNo direct security issues found; the PR adds metrics-based gating, skipped checks, tool argument and order validation, and rubric builders without introducing dangerous patterns. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: b03b4b5341
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| values_match = any( | ||
| isinstance(tool_args, dict) | ||
| and all( | ||
| arg in tool_args and tool_args[arg] == value for arg, value in expected_args.items() |
There was a problem hiding this comment.
Enforce type equality for expected tool arguments
When arguments contain JSON values of different types, Python's == can accept them as equal—for example, True == 1, False == 0, and nested {"enabled": true} compares equal to {"enabled": 1}. This lets expected_tool_arguments pass despite the documented no-type-coercion contract, potentially allowing an invalid tool invocation through an evaluation gate; compare value types recursively as well as values.
Useful? React with 👍 / 👎.
Summary
Adds four general-purpose features and bumps the version to 0.7.0. Everything is additive: no existing field, default or score name changes, reports saved by 0.6.0 still load, and
schema_versionstays1.0.Intent
A report could only say how the suite did overall. This PR lets it say where the failures are (by tag and by metric), lets a gate act on that, and closes two gaps that forced users to write custom evaluators: checking tool-argument values and tool order, and writing an LLM rubric from scratch. None of the features is tied to a domain; tags, metrics, tools and criteria are all supplied by the user.
What's added
Report breakdowns
EvaluationSummary.tagsandEvaluationSummary.metrics: per-tag and per-metric counts, pass rate and average score (TagSummary,MetricSummary).Score.metric: a stable grouping key. It defaults toScore.name; implicit checks such ascontains:<text>andrequired_tool:<tool>share the metric named by their prefix.Score.nameis unchanged.Score.skip(name, explanation)andScore.skipped: records a check that did not apply. A skipped score never fails the case and is left out of averages and pass rates.CaseEvaluation.tagsandCaseEvaluation.metadata, copied from theTestCase.Slice gates
GateConfig.min_tag_pass_rateandGateConfig.min_metric_pass_rate: pass-rate floors for one tag or one metric, applied on top ofmin_pass_rate.Tool expectations
TestCase.expected_tool_arguments: expected argument values per tool, scored astool_arg_values:<tool>.TestCase.required_tool_order: tools whose first calls must come in the listed order, scored astool_order.Rubric builder
make_rubric(name, criteria, levels=..., with_reference=...): builds aRubricTemplatefrom criteria in plain language, pass/fail by default or with a caller-supplied graded scale.Deliberate behaviour, for review
min_tag_pass_rate/min_metric_pass_ratebut absent from the report produces a failure reason, so a renamed tag cannot switch a check off.5does not match"5".required_tool_order=["a", "b"]fails ifbis called before the firsta, even when a latera,bpair is in order. A listed tool that is never called also fails.CallableEvaluatorleavespassedandrequireduntouched on a skipped score and only setsevaluator_type.Eval()ignores a skippedScorereturned by a scorer.Fixed
evaluate_suiteraisedZeroDivisionErrorwhen no case produced a score (for example, every evaluator returned an empty list).average_scoreis now reported as0.0.Voice-agent project
run_evals.pyreads the per-metric table fromreport.summary.metricsinstead of aggregating by hand, and gains a per-category table fromreport.summary.tags. Metric rows are now grouped (containsrather thancontains:1234).expected_tool_arguments.scenario_runner.pymaps scores to rules byScore.metricinstead of parsing name prefixes. The scenario lab's owntool_argumentsevaluator is unchanged.Docs and website
make_rubric.## 0.7.0 - 2026-10-03.taskargument, the eval-pack YAML shape,load_packbeing passed a string instead of aPath, and a broken link to the skills directory.Version
Bumped to 0.7.0 in
pyproject.toml,__init__.py,uv.lockand the website badge.Testing
ruff check,ruff format --checkandmypypass.pytest: 170 passed. New test files:test_report_breakdowns.py,test_tool_expectations.py; new cases intest_gate.pyandtest_scorers_rubric.py. No existing test was changed.examples/run clean.twine check.