Skip to content

Latest commit

 

History

201 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

AgenticLens

AgenticLens logo

Open-source observability, evaluation, and operational intelligence for production AI systems.

CI Docs PyPI Python License: MIT GitHub stars GitHub forks PyPI downloads

Public asset Link
Website and docs GitHub Pages
Product roadmap agenticlens-roadmap.md
Workflow specification docs/workflow-schema-spec.md
Research roadmap AgenticLens_Research_and_Development_Roadmap.md

AgenticLens is an open-source Python operational toolkit for LLM applications and agentic workflows. It helps developers instrument the AI runtime they are actually building: workflows, agents, LLM calls, prompts, context, retrieval, memory, tools, MCP actions, evaluations, safety signals, and reliability events.

It then turns that runtime into inspectable local artifacts, telemetry, and actionable recommendations.

Think of it as a lightweight, local cProfile for AI workflows: no hosted dashboard, no required backend, no account, and no data egress just to inspect a run.

The product idea is simple:

instrument the AI runtime once, export everywhere

Development continues with judge calibration and dataset management, followed by experiments and statistical comparison. The open-source readiness plan tracks reproducible examples, artifact interoperability, independent validation, governance, and onboarding alongside those features. These are planned goals, not a claim of AAIF acceptance or completed independent validation.

The code-backed roadmap audit records implemented, partial, and missing features with source/test evidence and open acceptance gates.

Contents

Why AgenticLens?

LLM applications rarely spend money in one place. Cost often leaks across planners, retrievers, memory, tool calls, repeated system prompts, and final response steps.

Most observability tools can show token usage. AgenticLens focuses on the next questions:

What ran, why did it behave that way, what should I change, and can I export that evidence anywhere I need?

AgenticLens currently detects token waste patterns such as:

  • repeated system prompts that may be cached or deduplicated
  • excessive retrieved chunks in RAG workflows
  • low-utility retrieved chunks that appear unlikely to affect the final answer
  • long conversation history that should be summarized or truncated
  • duplicate tool calls that should be cached
  • model-tier mismatches where a lower-cost model may handle the recorded workload
  • projected token, dollar-per-run, and monthly savings

It can also capture hierarchical agent traces, measure memory and retry overhead, compare repeated baseline and candidate runs, and fail CI when a candidate exceeds configured regression limits.

AgenticLens can now evaluate versioned test suites against recorded outputs and traces. Deterministic checks cover answer content, required or forbidden tool use, required tool arguments, structured JSON output, required output fields, turn counts, latency, and cost. The resulting evidence can be exported as JSON, rendered as a standalone HTML report, and enforced as a release gate in CI. Trusted live Python and HTTP targets can also be executed directly against the same suite. Versioned evaluation datasets, split management, sample export, and judge-calibration reports are available for local evaluation workflows.

Architecture

AgenticLens provides two compatible instrumentation paths:

Application code
├── Workflow profiler
│   ├── profile() and step()
│   ├── provider usage extraction
│   ├── automatic cost calculation
│   └── optimization recommendations
│
└── Research trace API
    ├── trace() and nested spans
    ├── raw execution metrics
    ├── deterministic findings
    └── repeated-run comparison

Local artifacts
├── workflow JSON
├── run trace JSON
├── comparison JSON and CSV
├── Markdown
└── Jira-oriented output

Use the workflow profiler for step-level token and cost optimization with provider response extraction. Use the trace API for hierarchical execution evidence, memory/retry analysis, and repeated-run experiments. Applications may use both while the research API evolves.

AgenticLens remains package-first, local-first, framework-neutral, CI-friendly, and advisory-first. It does not require a hosted backend and does not automatically change production prompts, models, tools, or routing.

That boundary matters in the broader DeepAgentLabs ecosystem: AgenticLens observes and evaluates, while agenticops-control-tower is the future operator-facing control plane above the package layer rather than a hosted requirement of AgenticLens itself.

Step-Level Token Optimization

AgenticLens is designed to make token waste visible at the level where engineers can actually fix it:

Workflow area What AgenticLens flags Typical fix
Prompting Repeated system prompt prefixes Cache or deduplicate stable prompt blocks
RAG Too many retrieved chunks Lower top-k or tighten retrieval filters
RAG Low-utility chunks unlikely to affect the final answer Rerank, prune, or improve retrieval scoring
Memory Long conversation history Summarize or truncate older turns
Tools Duplicate tool calls with the same arguments Cache tool results
Multi-agent handoffs Large context passed between agents Pass structured summaries or key facts
Model selection A lower-cost candidate can process the recorded token volume Evaluate the candidate against quality requirements before switching

The analyze command reports reducible tokens by step, so teams can see whether the biggest opportunity is in retrieval, memory, planning, tool use, or final response generation.

For multi-agent workflows, pass agent_name and optional handoff metadata:

with step(
    "Research answer",
    type="llm_call",
    agent_name="research_agent",
    agent_role="researcher",
    handoff_from="planner_agent",
    handoff_to="answer_agent",
    handoff_tokens=5200,
):
    ...

Status

AgenticLens is early-stage software. Workflow profiling, structured execution tracing, cost calculation, deterministic diagnostics, repeated-run comparison, export, CLI, and the rule-based recommendation engine are implemented. The research trace API is experimental and may evolve before a stable 1.0 release.

Capability Status
Workflow and step profiler Implemented
OpenAI and Anthropic usage extraction Implemented
Live, cached, bundled, and overridden pricing Implemented
Rule-based token optimization Implemented
RAG chunk-utility analysis Implemented
Hierarchical run/span tracing Implemented, experimental
Memory and retry overhead findings Implemented, experimental
Repeated-run regression comparison Implemented, experimental
Unified evaluator SDK and versioned test suites Implemented, experimental
Provider-neutral custom and LLM-as-a-Judge adapters Implemented, experimental
Live Python and HTTP evaluation targets Implemented, experimental
Quality, tool-use, latency, and cost release gates Implemented, experimental
Standalone evaluation HTML report Implemented, experimental
Versioned evaluation dataset management Implemented, experimental
Judge calibration with confidence intervals Implemented, experimental
Multi-variant repeated experiment runner Implemented, experimental
AIOS draft validation and conformance CLI Implemented, experimental
OTLP/HTTP JSON trace export Implemented, experimental
OTLP/OpenTelemetry ingestion adapter (gen_ai.* semconv, import-otlp) Implemented, experimental
Live OTLP receiver + real-time dashboard (optional [api] extra, serve-otlp) Implemented, experimental
Persistent local trace history (--db, /history, agenticlens history) Implemented, experimental
Local HTML dashboard (timeline, cost, findings, gate, comparison) Implemented, experimental
Statistical significance testing Planned
Framework trace adapters (native LangChain/LangGraph instrumentation, not already-OTel data) Planned
ModelFit and governance Planned

Evaluation and Release Gates

Versioned YAML or JSON suites turn expected agent behavior into executable acceptance criteria. AgenticLens checks response content, required and forbidden tools, required tool arguments, structured JSON output, required output fields, turn counts, end-to-end latency, and estimated cost against recorded run traces.

agenticlens evaluate suite.yaml samples.json \
  --save evaluation.json \
  --html evaluation.html

agenticlens gate evaluation.json \
  --min-pass-rate 0.95 \
  --min-average-score 0.98 \
  --max-failed-cases 1

The evaluation command produces machine-readable JSON and an optional standalone HTML report. The gate command returns exit status 2 when a configured release threshold fails, making it suitable for CI.

AgenticLens can also manage local evaluation datasets:

agenticlens dataset summary dataset.json
agenticlens dataset split dataset.json --save dataset-split.json --seed 7
agenticlens dataset export-samples dataset-split.json --split test --save samples-test.json
agenticlens experiment run experiment.yaml suite.yaml --save experiment-report.json

... and calibrate an llm_judge evaluator's verdicts against a versioned, human-labeled reference set (labels.json: a CalibrationDataset of bare {case_id, passed} reference labels — a different, simpler shape than the dataset.json above):

agenticlens calibrate evaluation.json labels.json --evaluator answer_quality --save calibration.json

Reports agreement rate with a 95% Wilson confidence interval plus a true/false accept/reject confusion breakdown; requires exact case-id and suite-name/version matching between the report and the reference set.

Dashboard Report

analyze, inspect, and compare can each render their own standalone HTML view (--html), and dashboard composes several already-saved artifacts (workflow, run trace, evaluation report, comparison report) into one page — an agent timeline, cost-by-workflow-area breakdown, waste findings, a release gate, and a baseline-vs-candidate comparison, whichever inputs are present:

agenticlens analyze workflow.json --html analyze.html
agenticlens inspect run.json --html inspect.html
agenticlens compare baseline/ candidate/ --html compare.html

agenticlens dashboard \
  --workflow workflow.json \
  --run run.json \
  --evaluation agenticlens-evaluation.json \
  --comparison comparison.json \
  --save agenticlens-dashboard.html

The dashboard is a single self-contained HTML file with no external requests (fonts, CDNs, or otherwise) and no data leaves the machine — the same local-first posture as the rest of AgenticLens.

AIOS Validation and Conformance

AgenticLens can validate AI Operations Specification draft workflow and run artifacts against the sibling ai-operations-spec schemas and semantic rules.

agenticlens validate workflow.json --version 0.4
agenticlens conformance run.json --version 0.4 --spec-root ../ai-operations-spec

validate performs schema checks. conformance adds semantic graph and reference checks and reports draft alignment rather than stable conformance, because AIOS v0.4 remains a draft.

OpenTelemetry Export

Structured trace() runs can now emit OTLP/HTTP JSON spans when configured:

from agenticlens import SpanType, trace

with trace(
    "support-agent",
    otlp_endpoint="http://localhost:4318/v1/traces",
) as recording:
    with recording.span("planner", SpanType.PLANNING) as planner:
        planner.record_tokens(input_tokens=120, output_tokens=30)

Or configure export through environment variables:

export AGENTICLENS_OTLP_TRACES_ENDPOINT=http://localhost:4318/v1/traces
export AGENTICLENS_OTLP_HEADERS='Authorization=Bearer local-dev-token'
export AGENTICLENS_OTLP_TIMEOUT_SECONDS=10

See examples/operational_intelligence_demo.py for a runnable local example that writes both an AgenticLens run artifact and an OTLP payload.

OpenTelemetry Ingestion

The inverse of the export above: convert OTLP/HTTP JSON trace exports — from an OTel Collector's file exporter, another vendor's trace dump, or AgenticLens's own OTLP export — into AgenticLens run files that flow straight into inspect, dashboard, compare, and evaluate:

agenticlens import-otlp otlp-export.json --save-dir imported-runs/
agenticlens dashboard --run imported-runs/<trace-id>.json --save report.html

Every field prefers AgenticLens's own agenticlens.* attributes (a perfect round-trip with the exporter above), falls back to the OpenTelemetry GenAI semantic convention (gen_ai.*, including legacy names such as gen_ai.usage.prompt_tokens and gen_ai.system), and never fabricates a value it cannot find — an unpriced span stays unpriced rather than showing $0.00. Any attribute the adapter doesn't recognize is preserved verbatim on the span rather than discarded, so pointing this at a real production collector export is safe: nothing is silently dropped, and nothing is silently invented. --save-dir is optional — without it, import-otlp prints a summary table as a dry run and writes nothing.

import-otlp is for offline/CI use — export a file, convert it, move on. For watching traces arrive live, see the receiver below.

What survives ingestion, and what doesn't

Pure OpenTelemetry GenAI-semconv data (no agenticlens.* attributes layered on top) round-trips some fields perfectly and structurally cannot carry others — this is a ceiling in what the OTel spec currently standardizes, not a shortcoming of the adapter:

Survives cleanly Structurally invisible to pure OTel
trace_id/span_id/parent_span_id, timestamps, latency Cost — no standardized cost attribute exists; every pure-OTel import shows cost as unavailable, never $0.00
input_tokens/output_tokens (gen_ai.usage.*) Retry attempt number — no OTel concept, so retry-attribution features can't use pure OTel data
model_name, provider, tool_name, agent_name 7 of 11 span types: retrieval, planning, memory_read, memory_write, validation, retry, final_response — only model_call, tool_call, and delegation map from gen_ai.operation.name; everything else becomes custom
error_type/error_message (via the standard OTel exception event) Run-level task_success, experiment_id, variant_id, task_id/task_type, framework — AgenticLens-specific product concepts with no OTel equivalent
application_name (via service.name)

The fix, when you control the source: emit the matching agenticlens.* attribute alongside the standard gen_ai.* ones (e.g. an explicit agenticlens.span_type on a retrieval span, or agenticlens.estimated_cost_usd on a priced call) — the adapter always prefers those first. That gets you full native fidelity without giving up your existing OTel instrumentation.

Live OTLP Receiver

Behind an optional extra (pip install agenticlens[api], adding FastAPI and uvicorn — nothing in the base package requires them). Runs a real OTLP/HTTP endpoint an OTel Collector's otlphttp exporter can be pointed at directly, plus a real-time HTML dashboard that shows traces as they land:

pip install 'agenticlens[api]'
agenticlens serve-otlp --port 4318 --save-dir live-runs/
  • POST http://localhost:4318/v1/traces — the endpoint to give your Collector or SDK's OTLP/HTTP exporter. Reuses the exact same conversion as import-otlp (same field-mapping rules, same never-fabricate guarantee).
  • http://localhost:4318/ — the live dashboard: the most recently received trace's timeline and cost breakdown, auto-refreshing, with a nav strip to switch between recent traces.
  • --save-dir (optional) also writes every received trace to disk as <trace_id>.json, so a live session still leaves a permanent artifact trail rather than being purely ephemeral.
  • The in-memory store is capped (--max-traces, default 200) and evicts the oldest trace first — bounded by construction, not by discipline.

Security note: this endpoint has no built-in authentication and binds to 127.0.0.1 by default on purpose. It's meant for local/dev visibility or inside a network you already trust. Anything reachable beyond localhost needs your own auth/network controls in front of it — this command does not provide any.

The live refresh is plain polling (the page reloads every few seconds), not a websocket/SSE push — a deliberate simplicity choice for this first version, and the natural next upgrade if that latency ever matters.

Persistent history (cross-trace view)

By default the receiver's store is in-memory only — restart it and every trace is gone. --db turns that into a durable local store, and /history gives you the cross-trace view a single-trace dashboard can't: aggregate tokens/cost, error rate, and p95 latency across recent traces, not just one at a time.

agenticlens serve-otlp --port 4318 --db traces.db --save-dir live-runs/
  • --db PATH — persists every received trace to a local SQLite file (Python's stdlib sqlite3, no new dependency, no change to the [api] extra) instead of memory-only. Traces survive a restart.

  • http://localhost:4318/history — aggregate stats (total traces, total tokens, total cost — or "N of M traces priced" when only some are, never a fabricated $0.00) plus a row per recent trace, linking back to its own live dashboard.

  • --max-persisted N caps the persistent store (unbounded by default — the whole point of --db is not throwing history away); kept separate from --max-traces, which still governs the in-memory path.

  • View history offline, without running the receiver at all:

    agenticlens history traces.db --save history.html
    agenticlens history live-runs/ --save history.html   # or a --save-dir of run JSON

Alerting was deliberately left out — this is local visibility for a human watching a dashboard, not a paging system.

Run a trusted live target directly:

agenticlens evaluate-live suite.yaml \
  --target-kind python \
  --target examples/live_evaluation_demo.py:run_case \
  --save evaluation-live.json

evaluate-live is intentionally powerful. Python targets execute local code and HTTP targets can reach arbitrary URLs, so suite files and live targets should be treated as trusted developer-controlled inputs.

The offline LangGraph pitch demonstration exercises the complete workflow:

uv sync --extra langgraph
uv run python -m examples.pitch_demo.run_pitch_demo

It performs a real supervisor graph execution, records structured multi-agent spans, evaluates output and tool behavior, applies the release gate, and writes the presentation-ready report to examples/pitch_demo/artifacts/evaluation.html.

Installation

For local development from this repository:

git clone https://github.com/DeepAgentLabs/agenticlens.git
cd agenticlens
uv sync --extra dev

If you do not use uv, install in editable mode with development extras:

python -m venv .venv
. .venv/bin/activate
pip install -e ".[dev]"

On Windows PowerShell:

python -m venv .venv
.\.venv\Scripts\Activate.ps1
pip install -e ".[dev]"

Quickstart

Instrument your workflow with explicit profile() and step() blocks:

from agenticlens import profile, step

with profile("Customer Support Agent"):
    with step(
        "Planner",
        type="planner",
        provider="openai",
        model="gpt-4o-mini",
        prompt=planner_prompt,
    ) as s:
        response = planner_llm.invoke(planner_prompt)
        s.record(response)

    with step(
        "Retriever",
        type="retriever",
        chunk_count=12,
        avg_tokens_per_chunk=80,
    ):
        chunks = retriever.search(user_question)

    with step(
        "Final Answer",
        type="final_response",
        provider="openai",
        model="gpt-4o-mini",
        final_answer="Refunds are processed to the original payment method.",
    ) as s:
        response = answer_llm.invoke(final_prompt)
        s.record(response)

Then profile and analyze a script:

uv run agenticlens profile examples/recommendations_demo.py --save workflow.json
uv run agenticlens analyze workflow.json

Example output:

Budget Optimization Run cost: $0.0068; reducible: ~$0.0024/run (35%), ~$2.38/month.

Optimization Suggestions
  * Long conversation history
  * Excessive retrieved chunks
  * Repeated system prompt
  * Low-utility retrieved chunks
  * Duplicate tool call

Estimated Savings: 35%

Structured Agent Tracing

The research trace API represents one agent execution as a Run containing nested Span objects. It is framework- and provider-neutral and is additive to the existing profile() and step() API.

from agenticlens import SpanType, trace

with trace(
    "customer-support-agent",
    environment="staging",
    prompt_version="support-v4",
) as recording:
    with recording.span("create-plan", SpanType.PLANNING) as planner:
        plan = create_plan()
        planner.record_tokens(input_tokens=300, output_tokens=75)

    with recording.span("load-history", SpanType.MEMORY_READ) as memory:
        history = load_customer_history()
        memory.record_tokens(input_tokens=800)

    with recording.span("search-account", SpanType.TOOL_CALL) as tool:
        account = search_account()
        tool.record_tokens(input_tokens=50, output_tokens=120)

    with recording.span("generate-answer", SpanType.MODEL_CALL) as model:
        answer = generate_answer(account, history)
        model.record_tokens(input_tokens=1200, output_tokens=250)
        model.record_cost(0.021)

recording.save("run.json")

Supported span types include planning, model calls, memory reads and writes, retrieval, tool calls, validation, retries, delegation, final responses, and custom operations.

Each run can report:

  • total input, output, and combined tokens
  • end-to-end and per-span latency
  • estimated cost
  • tokens and latency by span type
  • tool-call and retry counts
  • execution status and captured exceptions
  • memory and retry overhead

Parent-child span relationships preserve execution structure, such as a retry that occurred inside a failed tool call. Saved traces are validated for duplicate IDs, missing parents, self-parent relationships, and cyclic parent graphs.

Inspect a saved trace:

agenticlens inspect run.json

The terminal report includes a run summary, nested span tree, token and latency distributions, errors, retries, deterministic findings, and next-best-analysis guidance when findings indicate a likely follow-up investigation path.

Save a Markdown trace report:

agenticlens inspect run.json --save trace-report.md

Trace lifecycle

trace() entered
    ↓
Run created with status "running"
    ↓
Nested spans record operations
    ↓
Each span records timing, usage, status, and optional evidence
    ↓
Exceptions mark the active span and run as "failed"
    ↓
Run receives its completion time and final status
    ↓
Run is saved as portable JSON

Exceptions are recorded but not swallowed. The original exception continues to propagate so application behavior is unchanged:

with trace("tool-agent") as recording:
    with recording.span("lookup", SpanType.TOOL_CALL):
        raise TimeoutError("Customer database timed out")

Runs carry identity, application, framework, task, experiment, timing, status, success, error, and metadata fields. Spans carry parent relationships, type, agent, provider, model, tool, retry, usage, timing, cost, status, error, references, optional redacted payloads, and extensible attributes.

Run totals are reproducible from the recorded spans:

total_input_tokens  = sum(span.input_tokens)
total_output_tokens = sum(span.output_tokens)
total_tokens        = input + output
estimated_cost      = sum(known span costs)
end-to-end latency  = completed_at - started_at

Privacy-Preserving Capture

Prompts, responses, and tool arguments are not captured by default. Applications must explicitly opt in:

with recording.span("model", SpanType.MODEL_CALL) as span:
    response = call_model(request)
    span.record_io(input_data=request, output_data=response)

Explicitly captured values pass through a recursive redactor. The default redactor covers common secret fields, authorization values, bearer tokens, cookies, passwords, API keys, and email addresses. A custom redactor= function can be supplied for application-specific requirements.

The built-in redactor is a defense-in-depth control, not a complete compliance or data-loss-prevention system. Teams should still minimize payload capture and apply their own retention and access policies.

Memory and Retry Diagnostics

AgenticLens calculates:

memory_share = memory_tokens / total_tokens
retry_token_share = retry_tokens / total_tokens
retry_latency_share = retry_latency / total_latency

It also reports retry count, retry latency, and retry cost. When memory or retry consumption exceeds a configured threshold, AgenticLens produces a deterministic finding containing:

  • the measured values
  • the threshold that was exceeded
  • severity and confidence
  • exact span IDs that contributed to the finding

These findings identify measurable overhead. Retry findings also retain evidence about likely triggering failures and whether a retry appears to have recovered, failed, or remained unresolved.

Repeated-Run Comparison

Agent systems are nondeterministic, so one run is rarely sufficient. Store baseline and candidate traces in separate directories:

results/
  baseline/
    run-001.json
    run-002.json
  candidate/
    run-001.json
    run-002.json

Then compare them:

agenticlens compare results/baseline results/candidate

For each group, AgenticLens calculates:

  • run count and task-success rate
  • mean, median, and P95 tokens
  • mean, median, and P95 latency
  • standard deviation and coefficient of variation
  • mean cost and cost per successful task

The comparison detects relative regressions in success rate, mean tokens, latency, and cost. The threshold is configurable:

agenticlens compare results/baseline results/candidate \
  --regression-threshold 0.05 \
  --save comparison.json

Use CSV for tabular analysis:

agenticlens compare results/baseline results/candidate \
  --save comparison.csv \
  --format csv

Use Markdown for a review-friendly report:

agenticlens compare results/baseline results/candidate \
  --save comparison.md \
  --format md

Use --fail-on-regression to return a nonzero exit status in CI:

agenticlens compare results/baseline results/candidate \
  --regression-threshold 0.05 \
  --fail-on-regression

Require a minimum cohort size before trusting a comparison:

agenticlens compare results/baseline results/candidate --min-samples 5

Current comparisons are descriptive. They do not claim statistical significance or causal attribution, especially for small or uncontrolled samples.

Designing a useful comparison

For credible results:

  1. Use the same test cases for baseline and candidate conditions.
  2. Keep unrelated settings fixed.
  3. Record prompt, model, tool, and dataset versions in run metadata.
  4. Run multiple trials per test case.
  5. Preserve failed runs instead of deleting them.
  6. Compare success and quality alongside cost and latency.
  7. Review trace-level evidence before accepting an aggregate conclusion.

A 5% regression flag means the configured relative threshold was exceeded. It does not mean the difference is statistically significant.

Interpreting cost per successful task

Average request cost can favor a cheap but unreliable configuration:

cost_per_successful_task = total_recorded_cost / successful_runs
Variant Mean run cost Success rate Cost per success
Small model $0.04 50% $0.08
Larger model $0.06 100% $0.06

Here, the larger model costs more per attempt but less per successful task.

Evaluation correctness and compatibility (unreleased)

  • Structured-output checks use JSON Schema Draft 2020-12, including enum, numeric limits, composition, and embedded references. Invalid schemas and unsupported dialects raise configuration errors. External references are not fetched. The format keyword retains its standard annotation-only behavior.
  • Duplicate or unknown sample IDs raise errors before scoring; missing samples still produce failed cases. NaN and Infinity are rejected as invalid JSON.
  • Cost totals are unavailable if any span/case is unpriced. Set estimated_cost_usd=0.0 explicitly for free spans. Cost gates reject incomplete reports; comparison cost metrics are unavailable for incompletely priced groups.
  • Explicit task_success overrides execution status in comparisons. Status is used only when task_success is absent.

For example, an output of 0 now fails a schema with type=integer and minimum=1. A successful execution with task_success=False counts as a failed task. An evaluation containing costs 0.01 and null has total_cost_usd=null and fails a configured cost gate. Existing reports can retain old totals; regenerate them for corrected per-span costs and comparison results.

These fixes preserve the report field shapes but intentionally tighten behavior. Schema-dependent applications may now fail checks that were previously ignored.

Judge calibration against human labels

Compare one saved llm_judge score with a versioned reference dataset, without making new model calls:

agenticlens calibrate evaluation.json labels.json --evaluator quality --save calibration.json

labels.json must match the evaluation report's suite name/version and contain exactly one boolean label for every case:

{
  "name": "support-human-review",
  "version": "1",
  "suite_name": "support",
  "suite_version": "1",
  "labels": [
    {"case_id": "case-1", "passed": true},
    {"case_id": "case-2", "passed": false}
  ]
}

The selected score name must occur exactly once per case and have type llm_judge. Calibration uses its saved passed verdict, preserving the evaluation threshold. Python users can call calibrate_judge(report, CalibrationDataset.model_validate_json(labels_text), evaluator="quality") from agenticlens.evaluation.

The JSON report includes agreement, a two-sided 95% Wilson interval, true/false accepts and rejects, and per-case verdicts with trace IDs. It retains dataset and suite versions. Fewer than 30 cases produces an exploratory-evidence warning; one-class references also produce a warning. Neither warning is an automatic quality gate.

This is verdict agreement reporting, not probability calibration, automatic threshold tuning, or proof that human labels are correct. Use independently reviewed, representative cases; correlated examples and judge-assisted labels can overstate reliability. Do not mix judge models/prompts or threshold settings within a calibration run; retain that configuration with the source evaluation report. The interval assumes independent cases. See the NIST Wilson interval reference. Dataset lifecycle management and broader statistical calibration remain planned.

Using Regression Checks in CI

Store or download a reviewed baseline, generate candidate traces in the build, and compare them:

name: Agent regression check

on:
  pull_request:

jobs:
  agent-regression:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - uses: astral-sh/setup-uv@v6
      - run: uv sync --extra dev
      - name: Generate candidate traces
        run: uv run python benchmarks/run_candidate.py
      - name: Compare with baseline
        run: |
          uv run agenticlens compare \
            benchmarks/baseline \
            benchmarks/candidate \
            --regression-threshold 0.05 \
            --save comparison.json \
            --fail-on-regression
      - uses: actions/upload-artifact@v4
        if: always()
        with:
          name: agenticlens-comparison
          path: comparison.json

--fail-on-regression returns exit code 2 when the comparison is valid but regressions are detected. Invalid inputs or unreadable traces return exit code 1.

Portable Schemas

Versioned JSON Schemas are provided for:

  • run traces: schemas/trace.schema.json
  • deterministic findings: schemas/finding.schema.json and schemas/v2/finding.schema.json
  • comparison reports: schemas/report.schema.json

The schemas are included in wheel distributions under agenticlens/schemas. They allow external systems to validate and consume artifacts without depending on AgenticLens internal Python classes.

For compatibility-sensitive integrations, prefer the versioned schema URLs published in each schema's $id instead of the unversioned convenience alias.

Artifact Purpose
Workflow JSON Existing profiler output and recommendation input
Run trace JSON Hierarchical research execution record
Finding JSON Deterministic diagnostic evidence
Comparison JSON Complete machine-readable baseline/candidate report
Comparison CSV Flat metric deltas for analysis and charts

The run trace and workflow artifact are related but currently distinct. Consumers should inspect the artifact schema rather than assuming they are interchangeable.

Core Concepts

Workflow

A workflow is one complete execution of an LLM application, such as answering a customer support question or running a multi-agent task.

with profile("Refund Support"):
    ...

Step

A step is a meaningful unit inside that workflow: planner, retriever, memory, tool call, LLM call, or final response.

with step("Retrieve Policy Chunks", type="retriever", chunk_count=10):
    ...

Recommendation

A recommendation is a rule-based optimization suggestion. Recommendations carry token savings, estimated percentage savings, dollar impact when pricing is known, confidence when relevant, and quality-risk notes for heuristics such as RAG chunk utility.

AI Runtime Objects

AgenticLens is moving toward an object-based model aligned with the AI Operations Specification. At a high level, the runtime includes:

  • Workflow
  • Request
  • Agent
  • LLM
  • Prompt
  • Context
  • RAG
  • Memory
  • Tool
  • MCP
  • Evaluation
  • Safety
  • Reliability
  • Incident

These runtime objects emit AI-native events such as workflow.run, agent.step, llm.call, prompt.render, rag.retrieve, memory.read, and tool.call.

Features

Area Capability
Profiling Explicit profile() and step() context managers
Tracing Framework-neutral Run and nested Span execution traces
Metrics Prompt tokens, completion tokens, total tokens, latency, TPS, cost
Diagnostics Memory-share and retry-overhead findings with span-level evidence
Comparison Repeated runs, P95, variability, cost per success, regression detection
Evaluation Structured-output, tool-argument, and turn-count checks
Privacy Opt-in payload capture with recursive redaction
Providers OpenAI and Anthropic response usage extraction
Costing User overrides, cached live LiteLLM pricing, bundled fallback pricing
Recommendations Repeated prompts, excessive chunks, low-utility chunks, long history, duplicate tool calls
Budget impact Dollar-per-run and monthly savings projections
CLI profile, report, analyze, inspect, compare, evaluate, evaluate-live, experiment run, and gate
Export Workflow reports, run traces, JSON, CSV, Markdown, and Jira
Schemas Versioned trace, finding, and comparison-report JSON Schemas
Tooling pytest, Ruff, mypy, GitHub Actions

Cost Calculation

AgenticLens calculates per-step cost from the provider, model, prompt tokens, and completion tokens recorded by the profiler:

input_cost = (prompt_tokens / 1000) * input_price_per_1k
output_cost = (completion_tokens / 1000) * output_price_per_1k
total_cost = input_cost + output_cost

Pricing resolution order:

  1. User-supplied pricing override
  2. Live LiteLLM community pricing feed, when enabled
  3. Bundled src/agenticlens/config/pricing.yaml
  4. Unknown model: cost is reported as None, not $0.00

Live pricing is enabled by default. AgenticLens downloads LiteLLM's community-maintained model pricing table and stores it in:

~/.cache/agenticlens/live_pricing_cache.json

The default cache lifetime is 24 hours and the default network timeout is five seconds. A fresh cache avoids another network request. If refresh fails, AgenticLens uses the stale cache when one exists; otherwise it falls back to the bundled table.

Live entries are converted from cost per token into AgenticLens's internal USD per 1,000-token representation. Model lookup supports direct model names, provider/model names, and explicit aliases for provider feeds whose versioned keys differ from AgenticLens model names.

Configure pricing with an AgenticLens YAML file:

pricing_overrides:
  "openai:internal-fine-tune":
    input_per_1k: 0.002
    output_per_1k: 0.008

live_pricing:
  enabled: true
  ttl_seconds: 86400
  timeout_seconds: 5
  cache_path: ".agenticlens/live_pricing_cache.json"

Point AgenticLens at the file with:

export AGENTICLENS_CONFIG=agenticlens.yaml

On Windows PowerShell:

$env:AGENTICLENS_CONFIG = "agenticlens.yaml"

For hermetic builds, offline execution, or tests, disable remote pricing:

export AGENTICLENS_DISABLE_LIVE_PRICING=1

User overrides always win, including when live pricing is enabled. This is useful for negotiated provider rates, private deployments, fine-tuned models, or internal chargeback prices.

When pricing cannot be resolved, AgenticLens emits an UnknownModelPricingWarning and preserves the cost as None. Reports render that value as unavailable rather than incorrectly treating an unknown model as free.

Model-Swap Cost Analysis

The model-swap recommender recalculates the current step cost using the active pricing configuration and compares it with lower-cost candidates. It uses the live LiteLLM table when available and the bundled table as a fallback.

Candidate discovery is restricted by default to a curated list of direct model providers so gateway and reseller aliases do not overwhelm the comparison. Recommendations include:

  • current provider and model
  • candidate provider and model
  • measured token volume used in the estimate
  • current and projected candidate cost
  • projected dollar and percentage savings
  • a quality-risk warning

A cheaper model is a candidate for evaluation, not an automatic replacement. AgenticLens does not claim equivalent quality and does not change production routing.

Cost-Aware Reports and Comparisons

Resolved step costs flow into:

  • workflow total cost
  • per-step and per-agent CLI summaries
  • JSON, CSV, Markdown, and Jira exports
  • projected recommendation savings
  • repeated-run mean cost
  • cost per successful task
  • baseline-versus-candidate cost regression detection

Trace spans also accept explicitly recorded estimated costs through span.record_cost(). Trace cost is currently caller-supplied; automatic pricing resolution is implemented for the existing profile() and step() workflow profiler.

Configuration Reference

AgenticLens loads YAML configuration from an explicit path passed to load_config(), from AGENTICLENS_CONFIG, or from defaults.

pricing_overrides:
  "openai:internal-fine-tune":
    input_per_1k: 0.002
    output_per_1k: 0.008

live_pricing:
  enabled: true
  url: "https://raw.githubusercontent.com/BerriAI/litellm/main/model_prices_and_context_window.json"
  cache_path: ".agenticlens/live_pricing_cache.json"
  ttl_seconds: 86400
  timeout_seconds: 5

recommender:
  system_prompt_prefix_tokens: 50
  max_chunks: 8
  history_token_limit: 4000
  monthly_runs: 1000
  warning_savings_pct: 5
  critical_savings_pct: 20
  warning_savings_usd: 0.005
  critical_savings_usd: 0.05
  rag_min_chunk_utility_score: 0.08
  rag_min_low_utility_chunks: 2
  handoff_token_limit: 3000
  model_swap_min_savings_pct: 15
  model_swap_providers:
    - openai
    - anthropic
    - gemini
Environment variable Purpose
AGENTICLENS_CONFIG Path to an AgenticLens YAML configuration file
AGENTICLENS_DISABLE_LIVE_PRICING Disable remote pricing and use cache/static fallback

Configuration through [tool.agenticlens] in pyproject.toml is planned but is not implemented yet.

RAG Chunk Utility

The RAG utility rule identifies retrieved chunks that are unlikely to influence the final answer. It supports multiple signal types (in priority order):

Signal Type Supported Fields Source
Citation cited, used, referenced (boolean) Your app logic
Reranker reranker_score, rerank_score, cross_encoder_score (0–1) Cross-encoder models
Embedding embedding_similarity, cosine_similarity, semantic_score (0–1) Vector search
Generic utility_score, relevance_score (0–1) Custom scoring
Fallback Word-overlap against final answer Automatic

Example chunk metadata:

{"text": "...", "reranker_score": 0.92}
{"text": "...", "cosine_similarity": 0.85}
{"text": "...", "cited": True}
{"text": "...", "utility_score": 0.12}

When rich signals (reranker, embedding, citation) are available, confidence is higher and quality risk is lower. If no explicit signals are present, it falls back to lightweight word-overlap against the final answer.

For a complete guide, see docs/rag-chunk-utility.md.

Examples

Run the recommendation demo:

uv run agenticlens profile examples/recommendations_demo.py --save workflow.json
uv run agenticlens analyze workflow.json

Other examples:

  • examples/basic_usage.py
  • examples/rag_customer_support_demo.py
  • examples/multiagent_support_demo.py
  • examples/multiagent_token_optimization_demo.py
  • examples/reference_workflows/langgraph_supervisor.py — offline LangGraph supervisor
  • examples/export_demo.py — export to Markdown and Jira
  • examples/live_evaluation_demo.py — trusted live Python target for evaluate-live
  • examples/dataset_and_calibration_demo.py — dataset splitting, judge labels, evaluation, and calibration together
  • examples/experiment_runner_demo.py — repeated multi-variant experiment manifest and comparison flow
  • examples/rag_scoring_demo.py — RAG chunk utility with reranker/embedding/citation signals
  • examples/custom_llm_judge.py — registering a custom LLMJudgeEvaluator for the shared evaluator contract
  • examples/operational_intelligence_demo.py — structured trace, OTLP export, and AIOS conformance together
  • examples/pitch_demo/ — offline LangGraph pitch demo tying tracing, evaluation, and release gates together (see docs/evaluation-and-release-gates.md)

Some examples call real provider APIs and require provider API keys.

The reference workflows are based on orchestration patterns published by the official framework repositories. See docs/multi-agent-reference-workflows.md for setup, source links, dependency isolation, and instrumentation boundaries.

Exporting Reports

Markdown

from agenticlens.exporters import MarkdownExporter

MarkdownExporter().export(workflow, "report.md")

With Recommendations

All exporters accept an optional recommendations parameter (Jira currently ignores it):

from agenticlens.exporters import MarkdownExporter, JSONExporter, CSVExporter
from agenticlens.recommenders import RecommendationEngine

engine = RecommendationEngine()
recs = engine.run(workflow)

MarkdownExporter().export(workflow, "report.md", recommendations=recs)
JSONExporter().export(workflow, "report.json", recommendations=recs)
CSVExporter().export(workflow, "steps.csv", recommendations=recs)
# CSV also writes steps_recommendations.csv alongside

Jira Integration

Post profiling results directly as a comment on a Jira issue:

from agenticlens.exporters import JiraExporter

JiraExporter(
    base_url="https://yourteam.atlassian.net",
    user_email="you@example.com",
    api_token="your-api-token",
    issue_key="PROJ-123",
).export(workflow)

Set credentials via environment variables for safety — see examples/export_demo.py for a complete example.

For sample output previews of all formats, see docs/export-formats.md.

CLI Reference

Profile a Python script:

uv run agenticlens profile app.py

Save a workflow report:

uv run agenticlens profile app.py --save workflow.json

Display a saved workflow:

uv run agenticlens report workflow.json

Analyze a saved workflow:

uv run agenticlens analyze workflow.json

Inspect a saved run trace:

uv run agenticlens inspect run.json

Save a Markdown trace report:

uv run agenticlens inspect run.json --save trace.md

Compare baseline and candidate traces:

uv run agenticlens compare results/baseline results/candidate

Save a comparison and fail CI on detected regressions:

uv run agenticlens compare results/baseline results/candidate \
  --save comparison.json \
  --fail-on-regression

Save a Markdown comparison report and enforce sample size:

uv run agenticlens compare results/baseline results/candidate \
  --save comparison.md \
  --format md \
  --min-samples 5

Evaluate recorded samples:

uv run agenticlens evaluate suite.yaml samples.json --html evaluation.html

Run a trusted live Python target:

uv run agenticlens evaluate-live suite.yaml \
  --target-kind python \
  --target examples/live_evaluation_demo.py:run_case

Apply a release gate:

uv run agenticlens gate evaluation.json --min-pass-rate 0.95

Command summary

Command Purpose
profile Run an instrumented Python script and optionally save its workflow
report Render an existing workflow JSON artifact
analyze Run optimization recommenders against a workflow
inspect Render a run trace, span tree, distributions, and findings
compare Compare baseline and candidate trace files or directories
validate Run AIOS draft schema validation on a workflow or run artifact
conformance Run AIOS draft schema and semantic checks with draft-alignment reporting
evaluate Score recorded outputs and traces against a test suite
evaluate-live Run a trusted live Python or HTTP target against a suite
experiment run Run repeated live trials for 3+ variants against one suite
gate Enforce release thresholds from an evaluation report

The compare command accepts either one JSON trace file or a directory of *.json traces for each condition.

Current Limitations

  • The research trace API and workflow profiler use separate artifact types.
  • Trace-span cost must currently be recorded by the caller.
  • Memory findings measure consumption, not semantic relevance or contribution.
  • Comparisons do not calculate confidence intervals or significance tests yet.
  • Built-in evaluators are deterministic; semantic, safety, and RAG-quality checks rely on application-supplied CallableEvaluator/LLMJudgeEvaluator logic rather than a bundled LLM-as-a-Judge provider.
  • Model-swap recommendations estimate cost and do not guarantee quality.
  • Live pricing uses a community-maintained feed that may lag provider changes.
  • Default redaction cannot guarantee removal of every domain-specific secret or personal identifier.
  • evaluate-live assumes trusted targets and suite definitions.
  • Framework adapters and a local dashboard remain planned.

Development

A Makefile provides shorthand for common tasks:

make install     # install dev dependencies
make check       # run all quality gates (lint + format + typecheck + test)
make test-cov    # tests with coverage report
make docs        # build documentation
make help        # list all available targets

Or run individual steps:

uv sync --extra dev --extra docs
uv run pytest
uv run ruff check .
uv run ruff format .
uv run mypy

Useful targeted checks while working:

uv run ruff check src tests
uv run ruff format --check src tests

Project Structure

src/agenticlens/
  instrumentation/ structured run and span tracing, payload redaction
  analysis/        memory and retry diagnostics
  comparison/      repeated-run statistics, regression reports, export
  reports/         trace inspection rendering
  profiler/       workflow and step profiling
  metrics/        cost and performance calculation
  providers/      provider response usage extraction
  recommenders/   rule-based optimization suggestions
  exporters/      JSON, CSV, Markdown, and Jira exports
  cli/            Typer CLI and Rich rendering
  config/         pricing and settings
  models/         Pydantic data models
schemas/           versioned trace, finding, and report JSON Schemas

Roadmap

Near-term priorities:

  • experiment manifests and confidence intervals
  • richer dataset curation workflows beyond versioned local dataset artifacts
  • automatic pricing resolution for research trace spans
  • prompt caching opportunity detection
  • integrations for LangChain, LangGraph, LiteLLM, and OpenAI Agents SDK
  • OpenTelemetry and OpenInference trace import
  • optional prompt compression handoff

See the product roadmap for committed product direction and the research roadmap for experimental research plans.

Contributing

Contributions are welcome. Good first areas include:

  • provider integrations
  • recommender rules
  • example workflows
  • docs and tutorials
  • export formats
  • test coverage

Please read CONTRIBUTING.md before opening a pull request.

Security

Please report vulnerabilities privately. See SECURITY.md.

Code of Conduct

This project follows CODE_OF_CONDUCT.md.

License

AgenticLens is released under the MIT License. See LICENSE.

About

An open-source profiler for AI agents that analyzes token usage, cost, latency, and optimization opportunities across LLM workflows.

Resources

Code of conduct

Contributing

Security policy

Stars

58 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages