Skip to content

bumped v0.70 - #7

Merged
jothsnapraveena merged 1 commit into
mainfrom
feat/report-breakdowns
Oct 4, 2026
Merged

jothsnapraveena merged 1 commit into
mainfrom
feat/report-breakdowns

Conversation

@jothsnapraveena

Copy link
Copy Markdown
Contributor

Summary

Adds four general-purpose features and bumps the version to 0.7.0. Everything is additive: no existing field, default or score name changes, reports saved by 0.6.0 still load, and schema_version stays 1.0.

Intent

A report could only say how the suite did overall. This PR lets it say where the failures are (by tag and by metric), lets a gate act on that, and closes two gaps that forced users to write custom evaluators: checking tool-argument values and tool order, and writing an LLM rubric from scratch. None of the features is tied to a domain; tags, metrics, tools and criteria are all supplied by the user.

What's added

Report breakdowns

  • EvaluationSummary.tags and EvaluationSummary.metrics: per-tag and per-metric counts, pass rate and average score (TagSummary, MetricSummary).
  • Score.metric: a stable grouping key. It defaults to Score.name; implicit checks such as contains:<text> and required_tool:<tool> share the metric named by their prefix. Score.name is unchanged.
  • Score.skip(name, explanation) and Score.skipped: records a check that did not apply. A skipped score never fails the case and is left out of averages and pass rates.
  • CaseEvaluation.tags and CaseEvaluation.metadata, copied from the TestCase.

Slice gates

  • GateConfig.min_tag_pass_rate and GateConfig.min_metric_pass_rate: pass-rate floors for one tag or one metric, applied on top of min_pass_rate.

Tool expectations

  • TestCase.expected_tool_arguments: expected argument values per tool, scored as tool_arg_values:<tool>.
  • TestCase.required_tool_order: tools whose first calls must come in the listed order, scored as tool_order.

Rubric builder

  • make_rubric(name, criteria, levels=..., with_reference=...): builds a RubricTemplate from criteria in plain language, pass/fail by default or with a caller-supplied graded scale.

Deliberate behaviour, for review

  • A missing slice fails the gate. A tag or metric named in min_tag_pass_rate / min_metric_pass_rate but absent from the report produces a failure reason, so a renamed tag cannot switch a check off.
  • Argument values are not coerced. 5 does not match "5".
  • Tool order is judged on first calls. required_tool_order=["a", "b"] fails if b is called before the first a, even when a later a, b pair is in order. A listed tool that is never called also fails.
  • Skipped scores bypass the threshold override. CallableEvaluator leaves passed and required untouched on a skipped score and only sets evaluator_type.
  • Eval() ignores a skipped Score returned by a scorer.

Fixed

  • evaluate_suite raised ZeroDivisionError when no case produced a score (for example, every evaluator returned an empty list). average_score is now reported as 0.0.

Voice-agent project

  • run_evals.py reads the per-metric table from report.summary.metrics instead of aggregating by hand, and gains a per-category table from report.summary.tags. Metric rows are now grouped (contains rather than contains:1234).
  • The baseline suite also sets expected_tool_arguments.
  • scenario_runner.py maps scores to rules by Score.metric instead of parsing name prefixes. The scenario lab's own tool_arguments evaluator is unchanged.

Docs and website

  • README: new sections for breakdowns, skipped checks, tool call expectations, slice gates and make_rubric.
  • Changelog: ## 0.7.0 - 2026-10-03.
  • Skills: five cards updated to point at the new fields.
  • Website: both pages document the new features.
  • Four existing examples on the docs page were corrected because they did not run against the real API: the first-eval task argument, the eval-pack YAML shape, load_pack being passed a string instead of a Path, and a broken link to the skills directory.

Version

Bumped to 0.7.0 in pyproject.toml, __init__.py, uv.lock and the website badge.

Testing

  • ruff check, ruff format --check and mypy pass.
  • pytest: 170 passed. New test files: test_report_breakdowns.py, test_tool_expectations.py; new cases in test_gate.py and test_scorers_rubric.py. No existing test was changed.
  • Every example added to the README and website was run as real code.
  • The voice-agent suite, all ten preset scenarios and the five scripts in examples/ run clean.
  • The package builds as 0.7.0 and passes twine check.
  • Not checked: the website pages were not opened in a browser, so the layout of the new sections is unverified.

@deepagentlabs-pragent

Copy link
Copy Markdown

PRAgent Review

Verdict: PASS
Risk: LOW
Confidence: 75%

Reviewed 25 changed file(s); no blocking issues found.

Final Pre-Merge Summary

Changed files: CHANGELOG.md (modified, +43/-0); README.md (modified, +133/-6); docs/docs/index.html (modified, +146/-19); docs/index.html (modified, +11/-6); pyproject.toml (modified, +1/-1); src/agentic_evals/init.py (modified, +7/-1); src/agentic_evals/evaluators.py (modified, +3/-1); src/agentic_evals/gate.py (modified, +35/-11); src/agentic_evals/models.py (modified, +69/-0); src/agentic_evals/runner.py (modified, +122/-3); 15 more file(s).

Impact On Existing Code

No merge conflicts or blocking findings were detected. The owner should still confirm tests and product expectations before merging.

Owner Permission Required

Owner approval required: approve merge if the summarized changes match intent.

Quality Gate

Quality gate retained all 0 actionable finding(s).

Severity Count
Critical 0
High 0
Medium 0
Low 0
Info 0

Code Review

The PR correctly adds new features for reporting breakdowns by tags and metrics, introduces tool argument and order checks, adds a rubric builder, and improves gate thresholds. Tests and docs are comprehensive and aligned. No correctness bugs or API contract breaks are found in the diff.

Security Review

No direct security issues found; the PR adds metrics-based gating, skipped checks, tool argument and order validation, and rubric builders without introducing dangerous patterns.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: b03b4b5341

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

values_match = any(
isinstance(tool_args, dict)
and all(
arg in tool_args and tool_args[arg] == value for arg, value in expected_args.items()

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Enforce type equality for expected tool arguments

When arguments contain JSON values of different types, Python's == can accept them as equal—for example, True == 1, False == 0, and nested {"enabled": true} compares equal to {"enabled": 1}. This lets expected_tool_arguments pass despite the documented no-type-coercion contract, potentially allowing an invalid tool invocation through an evaluation gate; compare value types recursively as well as values.

Useful? React with 👍 / 👎.

@manemsai
manemsai self-requested a review October 4, 2026 05:44
@jothsnapraveena
jothsnapraveena merged commit 98e6cdf into main Oct 4, 2026
6 of 7 checks passed

This branch was successfully deployed

1 active deployment
pypi — b03b4b53 Deployed Oct 4, 2026 by jothsnapraveena via publish-pypi #7
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants