Open-source observability, evaluation, and operational intelligence for production AI systems.
| Public asset | Link |
|---|---|
| Website and docs | GitHub Pages |
| Product roadmap | agenticlens-roadmap.md |
| Workflow specification | docs/workflow-schema-spec.md |
| Research roadmap | AgenticLens_Research_and_Development_Roadmap.md |
AgenticLens is an open-source Python operational toolkit for LLM applications and agentic workflows. It helps developers instrument the AI runtime they are actually building: workflows, agents, LLM calls, prompts, context, retrieval, memory, tools, MCP actions, evaluations, safety signals, and reliability events.
It then turns that runtime into inspectable local artifacts, telemetry, and actionable recommendations.
Think of it as a lightweight, local cProfile for AI workflows: no hosted
dashboard, no required backend, no account, and no data egress just to inspect a
run.
The product idea is simple:
instrument the AI runtime once, export everywhere
Development continues with judge calibration and dataset management, followed by experiments and statistical comparison. The open-source readiness plan tracks reproducible examples, artifact interoperability, independent validation, governance, and onboarding alongside those features. These are planned goals, not a claim of AAIF acceptance or completed independent validation.
The code-backed roadmap audit records implemented, partial, and missing features with source/test evidence and open acceptance gates.
- Why AgenticLens?
- Architecture
- Status
- Installation
- Quickstart
- Structured Agent Tracing
- Privacy-Preserving Capture
- Memory and Retry Diagnostics
- Repeated-Run Comparison
- Evaluation and Release Gates
- Using Regression Checks in CI
- Portable Schemas
- Features
- Cost Calculation
- Configuration Reference
- CLI Reference
- Current Limitations
- Development
- Roadmap
LLM applications rarely spend money in one place. Cost often leaks across planners, retrievers, memory, tool calls, repeated system prompts, and final response steps.
Most observability tools can show token usage. AgenticLens focuses on the next questions:
What ran, why did it behave that way, what should I change, and can I export that evidence anywhere I need?
AgenticLens currently detects token waste patterns such as:
- repeated system prompts that may be cached or deduplicated
- excessive retrieved chunks in RAG workflows
- low-utility retrieved chunks that appear unlikely to affect the final answer
- long conversation history that should be summarized or truncated
- duplicate tool calls that should be cached
- model-tier mismatches where a lower-cost model may handle the recorded workload
- projected token, dollar-per-run, and monthly savings
It can also capture hierarchical agent traces, measure memory and retry overhead, compare repeated baseline and candidate runs, and fail CI when a candidate exceeds configured regression limits.
AgenticLens can now evaluate versioned test suites against recorded outputs and traces. Deterministic checks cover answer content, required or forbidden tool use, required tool arguments, structured JSON output, required output fields, turn counts, latency, and cost. The resulting evidence can be exported as JSON, rendered as a standalone HTML report, and enforced as a release gate in CI. Trusted live Python and HTTP targets can also be executed directly against the same suite. Versioned evaluation datasets, split management, sample export, and judge-calibration reports are available for local evaluation workflows.
AgenticLens provides two compatible instrumentation paths:
Application code
├── Workflow profiler
│ ├── profile() and step()
│ ├── provider usage extraction
│ ├── automatic cost calculation
│ └── optimization recommendations
│
└── Research trace API
├── trace() and nested spans
├── raw execution metrics
├── deterministic findings
└── repeated-run comparison
Local artifacts
├── workflow JSON
├── run trace JSON
├── comparison JSON and CSV
├── Markdown
└── Jira-oriented output
Use the workflow profiler for step-level token and cost optimization with provider response extraction. Use the trace API for hierarchical execution evidence, memory/retry analysis, and repeated-run experiments. Applications may use both while the research API evolves.
AgenticLens remains package-first, local-first, framework-neutral, CI-friendly, and advisory-first. It does not require a hosted backend and does not automatically change production prompts, models, tools, or routing.
That boundary matters in the broader DeepAgentLabs ecosystem: AgenticLens
observes and evaluates, while agenticops-control-tower is the future
operator-facing control plane above the package layer rather than a hosted
requirement of AgenticLens itself.
AgenticLens is designed to make token waste visible at the level where engineers can actually fix it:
| Workflow area | What AgenticLens flags | Typical fix |
|---|---|---|
| Prompting | Repeated system prompt prefixes | Cache or deduplicate stable prompt blocks |
| RAG | Too many retrieved chunks | Lower top-k or tighten retrieval filters |
| RAG | Low-utility chunks unlikely to affect the final answer | Rerank, prune, or improve retrieval scoring |
| Memory | Long conversation history | Summarize or truncate older turns |
| Tools | Duplicate tool calls with the same arguments | Cache tool results |
| Multi-agent handoffs | Large context passed between agents | Pass structured summaries or key facts |
| Model selection | A lower-cost candidate can process the recorded token volume | Evaluate the candidate against quality requirements before switching |
The analyze command reports reducible tokens by step, so teams can see whether
the biggest opportunity is in retrieval, memory, planning, tool use, or final
response generation.
For multi-agent workflows, pass agent_name and optional handoff metadata:
with step(
"Research answer",
type="llm_call",
agent_name="research_agent",
agent_role="researcher",
handoff_from="planner_agent",
handoff_to="answer_agent",
handoff_tokens=5200,
):
...AgenticLens is early-stage software. Workflow profiling, structured execution tracing, cost calculation, deterministic diagnostics, repeated-run comparison, export, CLI, and the rule-based recommendation engine are implemented. The research trace API is experimental and may evolve before a stable 1.0 release.
| Capability | Status |
|---|---|
| Workflow and step profiler | Implemented |
| OpenAI and Anthropic usage extraction | Implemented |
| Live, cached, bundled, and overridden pricing | Implemented |
| Rule-based token optimization | Implemented |
| RAG chunk-utility analysis | Implemented |
| Hierarchical run/span tracing | Implemented, experimental |
| Memory and retry overhead findings | Implemented, experimental |
| Repeated-run regression comparison | Implemented, experimental |
| Unified evaluator SDK and versioned test suites | Implemented, experimental |
| Provider-neutral custom and LLM-as-a-Judge adapters | Implemented, experimental |
| Live Python and HTTP evaluation targets | Implemented, experimental |
| Quality, tool-use, latency, and cost release gates | Implemented, experimental |
| Standalone evaluation HTML report | Implemented, experimental |
| Versioned evaluation dataset management | Implemented, experimental |
| Judge calibration with confidence intervals | Implemented, experimental |
| Multi-variant repeated experiment runner | Implemented, experimental |
| AIOS draft validation and conformance CLI | Implemented, experimental |
| OTLP/HTTP JSON trace export | Implemented, experimental |
OTLP/OpenTelemetry ingestion adapter (gen_ai.* semconv, import-otlp) |
Implemented, experimental |
Live OTLP receiver + real-time dashboard (optional [api] extra, serve-otlp) |
Implemented, experimental |
Persistent local trace history (--db, /history, agenticlens history) |
Implemented, experimental |
| Local HTML dashboard (timeline, cost, findings, gate, comparison) | Implemented, experimental |
| Statistical significance testing | Planned |
| Framework trace adapters (native LangChain/LangGraph instrumentation, not already-OTel data) | Planned |
| ModelFit and governance | Planned |
Versioned YAML or JSON suites turn expected agent behavior into executable acceptance criteria. AgenticLens checks response content, required and forbidden tools, required tool arguments, structured JSON output, required output fields, turn counts, end-to-end latency, and estimated cost against recorded run traces.
agenticlens evaluate suite.yaml samples.json \
--save evaluation.json \
--html evaluation.html
agenticlens gate evaluation.json \
--min-pass-rate 0.95 \
--min-average-score 0.98 \
--max-failed-cases 1The evaluation command produces machine-readable JSON and an optional
standalone HTML report. The gate command returns exit status 2 when a
configured release threshold fails, making it suitable for CI.
AgenticLens can also manage local evaluation datasets:
agenticlens dataset summary dataset.json
agenticlens dataset split dataset.json --save dataset-split.json --seed 7
agenticlens dataset export-samples dataset-split.json --split test --save samples-test.json
agenticlens experiment run experiment.yaml suite.yaml --save experiment-report.json... and calibrate an llm_judge evaluator's verdicts against a versioned,
human-labeled reference set (labels.json: a CalibrationDataset of bare
{case_id, passed} reference labels — a different, simpler shape than the
dataset.json above):
agenticlens calibrate evaluation.json labels.json --evaluator answer_quality --save calibration.jsonReports agreement rate with a 95% Wilson confidence interval plus a true/false accept/reject confusion breakdown; requires exact case-id and suite-name/version matching between the report and the reference set.
analyze, inspect, and compare can each render their own standalone HTML
view (--html), and dashboard composes several already-saved artifacts
(workflow, run trace, evaluation report, comparison report) into one page —
an agent timeline, cost-by-workflow-area breakdown, waste findings, a release
gate, and a baseline-vs-candidate comparison, whichever inputs are present:
agenticlens analyze workflow.json --html analyze.html
agenticlens inspect run.json --html inspect.html
agenticlens compare baseline/ candidate/ --html compare.html
agenticlens dashboard \
--workflow workflow.json \
--run run.json \
--evaluation agenticlens-evaluation.json \
--comparison comparison.json \
--save agenticlens-dashboard.htmlThe dashboard is a single self-contained HTML file with no external requests (fonts, CDNs, or otherwise) and no data leaves the machine — the same local-first posture as the rest of AgenticLens.
AgenticLens can validate AI Operations Specification draft workflow and run
artifacts against the sibling ai-operations-spec schemas and semantic rules.
agenticlens validate workflow.json --version 0.4
agenticlens conformance run.json --version 0.4 --spec-root ../ai-operations-specvalidate performs schema checks. conformance adds semantic graph and
reference checks and reports draft alignment rather than stable conformance,
because AIOS v0.4 remains a draft.
Structured trace() runs can now emit OTLP/HTTP JSON spans when configured:
from agenticlens import SpanType, trace
with trace(
"support-agent",
otlp_endpoint="http://localhost:4318/v1/traces",
) as recording:
with recording.span("planner", SpanType.PLANNING) as planner:
planner.record_tokens(input_tokens=120, output_tokens=30)Or configure export through environment variables:
export AGENTICLENS_OTLP_TRACES_ENDPOINT=http://localhost:4318/v1/traces
export AGENTICLENS_OTLP_HEADERS='Authorization=Bearer local-dev-token'
export AGENTICLENS_OTLP_TIMEOUT_SECONDS=10See examples/operational_intelligence_demo.py for a runnable local example
that writes both an AgenticLens run artifact and an OTLP payload.
The inverse of the export above: convert OTLP/HTTP JSON trace exports —
from an OTel Collector's file exporter, another vendor's trace dump, or
AgenticLens's own OTLP export — into AgenticLens run files that flow
straight into inspect, dashboard, compare, and evaluate:
agenticlens import-otlp otlp-export.json --save-dir imported-runs/
agenticlens dashboard --run imported-runs/<trace-id>.json --save report.htmlEvery field prefers AgenticLens's own agenticlens.* attributes (a perfect
round-trip with the exporter above), falls back to the OpenTelemetry GenAI
semantic convention (gen_ai.*, including legacy names such as
gen_ai.usage.prompt_tokens and gen_ai.system), and never fabricates a
value it cannot find — an unpriced span stays unpriced rather than showing
$0.00. Any attribute the adapter doesn't recognize is preserved verbatim on
the span rather than discarded, so pointing this at a real production
collector export is safe: nothing is silently dropped, and nothing is
silently invented. --save-dir is optional — without it, import-otlp
prints a summary table as a dry run and writes nothing.
import-otlp is for offline/CI use — export a file, convert it, move on.
For watching traces arrive live, see the receiver below.
Pure OpenTelemetry GenAI-semconv data (no agenticlens.* attributes layered
on top) round-trips some fields perfectly and structurally cannot carry
others — this is a ceiling in what the OTel spec currently standardizes, not
a shortcoming of the adapter:
| Survives cleanly | Structurally invisible to pure OTel |
|---|---|
trace_id/span_id/parent_span_id, timestamps, latency |
Cost — no standardized cost attribute exists; every pure-OTel import shows cost as unavailable, never $0.00 |
input_tokens/output_tokens (gen_ai.usage.*) |
Retry attempt number — no OTel concept, so retry-attribution features can't use pure OTel data |
model_name, provider, tool_name, agent_name |
7 of 11 span types: retrieval, planning, memory_read, memory_write, validation, retry, final_response — only model_call, tool_call, and delegation map from gen_ai.operation.name; everything else becomes custom |
error_type/error_message (via the standard OTel exception event) |
Run-level task_success, experiment_id, variant_id, task_id/task_type, framework — AgenticLens-specific product concepts with no OTel equivalent |
application_name (via service.name) |
The fix, when you control the source: emit the matching agenticlens.*
attribute alongside the standard gen_ai.* ones (e.g. an explicit
agenticlens.span_type on a retrieval span, or agenticlens.estimated_cost_usd
on a priced call) — the adapter always prefers those first. That gets you
full native fidelity without giving up your existing OTel instrumentation.
Behind an optional extra (pip install agenticlens[api], adding FastAPI and
uvicorn — nothing in the base package requires them). Runs a real OTLP/HTTP
endpoint an OTel Collector's otlphttp exporter can be pointed at directly,
plus a real-time HTML dashboard that shows traces as they land:
pip install 'agenticlens[api]'
agenticlens serve-otlp --port 4318 --save-dir live-runs/POST http://localhost:4318/v1/traces— the endpoint to give your Collector or SDK's OTLP/HTTP exporter. Reuses the exact same conversion asimport-otlp(same field-mapping rules, same never-fabricate guarantee).http://localhost:4318/— the live dashboard: the most recently received trace's timeline and cost breakdown, auto-refreshing, with a nav strip to switch between recent traces.--save-dir(optional) also writes every received trace to disk as<trace_id>.json, so a live session still leaves a permanent artifact trail rather than being purely ephemeral.- The in-memory store is capped (
--max-traces, default 200) and evicts the oldest trace first — bounded by construction, not by discipline.
Security note: this endpoint has no built-in authentication and binds to
127.0.0.1 by default on purpose. It's meant for local/dev visibility or
inside a network you already trust. Anything reachable beyond localhost
needs your own auth/network controls in front of it — this command does not
provide any.
The live refresh is plain polling (the page reloads every few seconds), not a websocket/SSE push — a deliberate simplicity choice for this first version, and the natural next upgrade if that latency ever matters.
By default the receiver's store is in-memory only — restart it and every
trace is gone. --db turns that into a durable local store, and /history
gives you the cross-trace view a single-trace dashboard can't: aggregate
tokens/cost, error rate, and p95 latency across recent traces, not just one
at a time.
agenticlens serve-otlp --port 4318 --db traces.db --save-dir live-runs/-
--db PATH— persists every received trace to a local SQLite file (Python's stdlibsqlite3, no new dependency, no change to the[api]extra) instead of memory-only. Traces survive a restart. -
http://localhost:4318/history— aggregate stats (total traces, total tokens, total cost — or "N of M traces priced" when only some are, never a fabricated$0.00) plus a row per recent trace, linking back to its own live dashboard. -
--max-persisted Ncaps the persistent store (unbounded by default — the whole point of--dbis not throwing history away); kept separate from--max-traces, which still governs the in-memory path. -
View history offline, without running the receiver at all:
agenticlens history traces.db --save history.html agenticlens history live-runs/ --save history.html # or a --save-dir of run JSON
Alerting was deliberately left out — this is local visibility for a human watching a dashboard, not a paging system.
Run a trusted live target directly:
agenticlens evaluate-live suite.yaml \
--target-kind python \
--target examples/live_evaluation_demo.py:run_case \
--save evaluation-live.jsonevaluate-live is intentionally powerful. Python targets execute local code
and HTTP targets can reach arbitrary URLs, so suite files and live targets
should be treated as trusted developer-controlled inputs.
The offline LangGraph pitch demonstration exercises the complete workflow:
uv sync --extra langgraph
uv run python -m examples.pitch_demo.run_pitch_demoIt performs a real supervisor graph execution, records structured multi-agent
spans, evaluates output and tool behavior, applies the release gate, and writes
the presentation-ready report to examples/pitch_demo/artifacts/evaluation.html.
For local development from this repository:
git clone https://github.com/DeepAgentLabs/agenticlens.git
cd agenticlens
uv sync --extra devIf you do not use uv, install in editable mode with development extras:
python -m venv .venv
. .venv/bin/activate
pip install -e ".[dev]"On Windows PowerShell:
python -m venv .venv
.\.venv\Scripts\Activate.ps1
pip install -e ".[dev]"Instrument your workflow with explicit profile() and step() blocks:
from agenticlens import profile, step
with profile("Customer Support Agent"):
with step(
"Planner",
type="planner",
provider="openai",
model="gpt-4o-mini",
prompt=planner_prompt,
) as s:
response = planner_llm.invoke(planner_prompt)
s.record(response)
with step(
"Retriever",
type="retriever",
chunk_count=12,
avg_tokens_per_chunk=80,
):
chunks = retriever.search(user_question)
with step(
"Final Answer",
type="final_response",
provider="openai",
model="gpt-4o-mini",
final_answer="Refunds are processed to the original payment method.",
) as s:
response = answer_llm.invoke(final_prompt)
s.record(response)Then profile and analyze a script:
uv run agenticlens profile examples/recommendations_demo.py --save workflow.json
uv run agenticlens analyze workflow.jsonExample output:
Budget Optimization Run cost: $0.0068; reducible: ~$0.0024/run (35%), ~$2.38/month.
Optimization Suggestions
* Long conversation history
* Excessive retrieved chunks
* Repeated system prompt
* Low-utility retrieved chunks
* Duplicate tool call
Estimated Savings: 35%
The research trace API represents one agent execution as a Run containing
nested Span objects. It is framework- and provider-neutral and is additive to
the existing profile() and step() API.
from agenticlens import SpanType, trace
with trace(
"customer-support-agent",
environment="staging",
prompt_version="support-v4",
) as recording:
with recording.span("create-plan", SpanType.PLANNING) as planner:
plan = create_plan()
planner.record_tokens(input_tokens=300, output_tokens=75)
with recording.span("load-history", SpanType.MEMORY_READ) as memory:
history = load_customer_history()
memory.record_tokens(input_tokens=800)
with recording.span("search-account", SpanType.TOOL_CALL) as tool:
account = search_account()
tool.record_tokens(input_tokens=50, output_tokens=120)
with recording.span("generate-answer", SpanType.MODEL_CALL) as model:
answer = generate_answer(account, history)
model.record_tokens(input_tokens=1200, output_tokens=250)
model.record_cost(0.021)
recording.save("run.json")Supported span types include planning, model calls, memory reads and writes, retrieval, tool calls, validation, retries, delegation, final responses, and custom operations.
Each run can report:
- total input, output, and combined tokens
- end-to-end and per-span latency
- estimated cost
- tokens and latency by span type
- tool-call and retry counts
- execution status and captured exceptions
- memory and retry overhead
Parent-child span relationships preserve execution structure, such as a retry that occurred inside a failed tool call. Saved traces are validated for duplicate IDs, missing parents, self-parent relationships, and cyclic parent graphs.
Inspect a saved trace:
agenticlens inspect run.jsonThe terminal report includes a run summary, nested span tree, token and latency distributions, errors, retries, deterministic findings, and next-best-analysis guidance when findings indicate a likely follow-up investigation path.
Save a Markdown trace report:
agenticlens inspect run.json --save trace-report.mdtrace() entered
↓
Run created with status "running"
↓
Nested spans record operations
↓
Each span records timing, usage, status, and optional evidence
↓
Exceptions mark the active span and run as "failed"
↓
Run receives its completion time and final status
↓
Run is saved as portable JSON
Exceptions are recorded but not swallowed. The original exception continues to propagate so application behavior is unchanged:
with trace("tool-agent") as recording:
with recording.span("lookup", SpanType.TOOL_CALL):
raise TimeoutError("Customer database timed out")Runs carry identity, application, framework, task, experiment, timing, status, success, error, and metadata fields. Spans carry parent relationships, type, agent, provider, model, tool, retry, usage, timing, cost, status, error, references, optional redacted payloads, and extensible attributes.
Run totals are reproducible from the recorded spans:
total_input_tokens = sum(span.input_tokens)
total_output_tokens = sum(span.output_tokens)
total_tokens = input + output
estimated_cost = sum(known span costs)
end-to-end latency = completed_at - started_at
Prompts, responses, and tool arguments are not captured by default. Applications must explicitly opt in:
with recording.span("model", SpanType.MODEL_CALL) as span:
response = call_model(request)
span.record_io(input_data=request, output_data=response)Explicitly captured values pass through a recursive redactor. The default
redactor covers common secret fields, authorization values, bearer tokens,
cookies, passwords, API keys, and email addresses. A custom redactor=
function can be supplied for application-specific requirements.
The built-in redactor is a defense-in-depth control, not a complete compliance or data-loss-prevention system. Teams should still minimize payload capture and apply their own retention and access policies.
AgenticLens calculates:
memory_share = memory_tokens / total_tokens
retry_token_share = retry_tokens / total_tokens
retry_latency_share = retry_latency / total_latency
It also reports retry count, retry latency, and retry cost. When memory or retry consumption exceeds a configured threshold, AgenticLens produces a deterministic finding containing:
- the measured values
- the threshold that was exceeded
- severity and confidence
- exact span IDs that contributed to the finding
These findings identify measurable overhead. Retry findings also retain evidence about likely triggering failures and whether a retry appears to have recovered, failed, or remained unresolved.
Agent systems are nondeterministic, so one run is rarely sufficient. Store baseline and candidate traces in separate directories:
results/
baseline/
run-001.json
run-002.json
candidate/
run-001.json
run-002.json
Then compare them:
agenticlens compare results/baseline results/candidateFor each group, AgenticLens calculates:
- run count and task-success rate
- mean, median, and P95 tokens
- mean, median, and P95 latency
- standard deviation and coefficient of variation
- mean cost and cost per successful task
The comparison detects relative regressions in success rate, mean tokens, latency, and cost. The threshold is configurable:
agenticlens compare results/baseline results/candidate \
--regression-threshold 0.05 \
--save comparison.jsonUse CSV for tabular analysis:
agenticlens compare results/baseline results/candidate \
--save comparison.csv \
--format csvUse Markdown for a review-friendly report:
agenticlens compare results/baseline results/candidate \
--save comparison.md \
--format mdUse --fail-on-regression to return a nonzero exit status in CI:
agenticlens compare results/baseline results/candidate \
--regression-threshold 0.05 \
--fail-on-regressionRequire a minimum cohort size before trusting a comparison:
agenticlens compare results/baseline results/candidate --min-samples 5Current comparisons are descriptive. They do not claim statistical significance or causal attribution, especially for small or uncontrolled samples.
For credible results:
- Use the same test cases for baseline and candidate conditions.
- Keep unrelated settings fixed.
- Record prompt, model, tool, and dataset versions in run metadata.
- Run multiple trials per test case.
- Preserve failed runs instead of deleting them.
- Compare success and quality alongside cost and latency.
- Review trace-level evidence before accepting an aggregate conclusion.
A 5% regression flag means the configured relative threshold was exceeded. It does not mean the difference is statistically significant.
Average request cost can favor a cheap but unreliable configuration:
cost_per_successful_task = total_recorded_cost / successful_runs
| Variant | Mean run cost | Success rate | Cost per success |
|---|---|---|---|
| Small model | $0.04 | 50% | $0.08 |
| Larger model | $0.06 | 100% | $0.06 |
Here, the larger model costs more per attempt but less per successful task.
- Structured-output checks use JSON Schema Draft 2020-12, including enum, numeric limits, composition, and embedded references. Invalid schemas and unsupported dialects raise configuration errors. External references are not fetched. The format keyword retains its standard annotation-only behavior.
- Duplicate or unknown sample IDs raise errors before scoring; missing samples still produce failed cases. NaN and Infinity are rejected as invalid JSON.
- Cost totals are unavailable if any span/case is unpriced. Set estimated_cost_usd=0.0 explicitly for free spans. Cost gates reject incomplete reports; comparison cost metrics are unavailable for incompletely priced groups.
- Explicit task_success overrides execution status in comparisons. Status is used only when task_success is absent.
For example, an output of 0 now fails a schema with type=integer and minimum=1. A successful execution with task_success=False counts as a failed task. An evaluation containing costs 0.01 and null has total_cost_usd=null and fails a configured cost gate. Existing reports can retain old totals; regenerate them for corrected per-span costs and comparison results.
These fixes preserve the report field shapes but intentionally tighten behavior. Schema-dependent applications may now fail checks that were previously ignored.
Compare one saved llm_judge score with a versioned reference dataset, without
making new model calls:
agenticlens calibrate evaluation.json labels.json --evaluator quality --save calibration.jsonlabels.json must match the evaluation report's suite name/version and contain
exactly one boolean label for every case:
{
"name": "support-human-review",
"version": "1",
"suite_name": "support",
"suite_version": "1",
"labels": [
{"case_id": "case-1", "passed": true},
{"case_id": "case-2", "passed": false}
]
}The selected score name must occur exactly once per case and have type
llm_judge. Calibration uses its saved passed verdict, preserving the
evaluation threshold. Python users can call
calibrate_judge(report, CalibrationDataset.model_validate_json(labels_text), evaluator="quality")
from agenticlens.evaluation.
The JSON report includes agreement, a two-sided 95% Wilson interval, true/false accepts and rejects, and per-case verdicts with trace IDs. It retains dataset and suite versions. Fewer than 30 cases produces an exploratory-evidence warning; one-class references also produce a warning. Neither warning is an automatic quality gate.
This is verdict agreement reporting, not probability calibration, automatic threshold tuning, or proof that human labels are correct. Use independently reviewed, representative cases; correlated examples and judge-assisted labels can overstate reliability. Do not mix judge models/prompts or threshold settings within a calibration run; retain that configuration with the source evaluation report. The interval assumes independent cases. See the NIST Wilson interval reference. Dataset lifecycle management and broader statistical calibration remain planned.
Store or download a reviewed baseline, generate candidate traces in the build, and compare them:
name: Agent regression check
on:
pull_request:
jobs:
agent-regression:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: astral-sh/setup-uv@v6
- run: uv sync --extra dev
- name: Generate candidate traces
run: uv run python benchmarks/run_candidate.py
- name: Compare with baseline
run: |
uv run agenticlens compare \
benchmarks/baseline \
benchmarks/candidate \
--regression-threshold 0.05 \
--save comparison.json \
--fail-on-regression
- uses: actions/upload-artifact@v4
if: always()
with:
name: agenticlens-comparison
path: comparison.json--fail-on-regression returns exit code 2 when the comparison is valid but
regressions are detected. Invalid inputs or unreadable traces return exit code
1.
Versioned JSON Schemas are provided for:
- run traces:
schemas/trace.schema.json - deterministic findings:
schemas/finding.schema.jsonandschemas/v2/finding.schema.json - comparison reports:
schemas/report.schema.json
The schemas are included in wheel distributions under agenticlens/schemas.
They allow external systems to validate and consume artifacts without depending
on AgenticLens internal Python classes.
For compatibility-sensitive integrations, prefer the versioned schema URLs
published in each schema's $id instead of the unversioned convenience alias.
| Artifact | Purpose |
|---|---|
| Workflow JSON | Existing profiler output and recommendation input |
| Run trace JSON | Hierarchical research execution record |
| Finding JSON | Deterministic diagnostic evidence |
| Comparison JSON | Complete machine-readable baseline/candidate report |
| Comparison CSV | Flat metric deltas for analysis and charts |
The run trace and workflow artifact are related but currently distinct. Consumers should inspect the artifact schema rather than assuming they are interchangeable.
A workflow is one complete execution of an LLM application, such as answering a customer support question or running a multi-agent task.
with profile("Refund Support"):
...A step is a meaningful unit inside that workflow: planner, retriever, memory, tool call, LLM call, or final response.
with step("Retrieve Policy Chunks", type="retriever", chunk_count=10):
...A recommendation is a rule-based optimization suggestion. Recommendations carry token savings, estimated percentage savings, dollar impact when pricing is known, confidence when relevant, and quality-risk notes for heuristics such as RAG chunk utility.
AgenticLens is moving toward an object-based model aligned with the AI Operations Specification. At a high level, the runtime includes:
WorkflowRequestAgentLLMPromptContextRAGMemoryToolMCPEvaluationSafetyReliabilityIncident
These runtime objects emit AI-native events such as workflow.run,
agent.step, llm.call, prompt.render, rag.retrieve, memory.read, and
tool.call.
| Area | Capability |
|---|---|
| Profiling | Explicit profile() and step() context managers |
| Tracing | Framework-neutral Run and nested Span execution traces |
| Metrics | Prompt tokens, completion tokens, total tokens, latency, TPS, cost |
| Diagnostics | Memory-share and retry-overhead findings with span-level evidence |
| Comparison | Repeated runs, P95, variability, cost per success, regression detection |
| Evaluation | Structured-output, tool-argument, and turn-count checks |
| Privacy | Opt-in payload capture with recursive redaction |
| Providers | OpenAI and Anthropic response usage extraction |
| Costing | User overrides, cached live LiteLLM pricing, bundled fallback pricing |
| Recommendations | Repeated prompts, excessive chunks, low-utility chunks, long history, duplicate tool calls |
| Budget impact | Dollar-per-run and monthly savings projections |
| CLI | profile, report, analyze, inspect, compare, evaluate, evaluate-live, experiment run, and gate |
| Export | Workflow reports, run traces, JSON, CSV, Markdown, and Jira |
| Schemas | Versioned trace, finding, and comparison-report JSON Schemas |
| Tooling | pytest, Ruff, mypy, GitHub Actions |
AgenticLens calculates per-step cost from the provider, model, prompt tokens, and completion tokens recorded by the profiler:
input_cost = (prompt_tokens / 1000) * input_price_per_1k
output_cost = (completion_tokens / 1000) * output_price_per_1k
total_cost = input_cost + output_cost
Pricing resolution order:
- User-supplied pricing override
- Live LiteLLM community pricing feed, when enabled
- Bundled
src/agenticlens/config/pricing.yaml - Unknown model: cost is reported as
None, not$0.00
Live pricing is enabled by default. AgenticLens downloads LiteLLM's community-maintained model pricing table and stores it in:
~/.cache/agenticlens/live_pricing_cache.json
The default cache lifetime is 24 hours and the default network timeout is five seconds. A fresh cache avoids another network request. If refresh fails, AgenticLens uses the stale cache when one exists; otherwise it falls back to the bundled table.
Live entries are converted from cost per token into AgenticLens's internal USD
per 1,000-token representation. Model lookup supports direct model names,
provider/model names, and explicit aliases for provider feeds whose versioned
keys differ from AgenticLens model names.
Configure pricing with an AgenticLens YAML file:
pricing_overrides:
"openai:internal-fine-tune":
input_per_1k: 0.002
output_per_1k: 0.008
live_pricing:
enabled: true
ttl_seconds: 86400
timeout_seconds: 5
cache_path: ".agenticlens/live_pricing_cache.json"Point AgenticLens at the file with:
export AGENTICLENS_CONFIG=agenticlens.yamlOn Windows PowerShell:
$env:AGENTICLENS_CONFIG = "agenticlens.yaml"For hermetic builds, offline execution, or tests, disable remote pricing:
export AGENTICLENS_DISABLE_LIVE_PRICING=1User overrides always win, including when live pricing is enabled. This is useful for negotiated provider rates, private deployments, fine-tuned models, or internal chargeback prices.
When pricing cannot be resolved, AgenticLens emits an
UnknownModelPricingWarning and preserves the cost as None. Reports render
that value as unavailable rather than incorrectly treating an unknown model as
free.
The model-swap recommender recalculates the current step cost using the active pricing configuration and compares it with lower-cost candidates. It uses the live LiteLLM table when available and the bundled table as a fallback.
Candidate discovery is restricted by default to a curated list of direct model providers so gateway and reseller aliases do not overwhelm the comparison. Recommendations include:
- current provider and model
- candidate provider and model
- measured token volume used in the estimate
- current and projected candidate cost
- projected dollar and percentage savings
- a quality-risk warning
A cheaper model is a candidate for evaluation, not an automatic replacement. AgenticLens does not claim equivalent quality and does not change production routing.
Resolved step costs flow into:
- workflow total cost
- per-step and per-agent CLI summaries
- JSON, CSV, Markdown, and Jira exports
- projected recommendation savings
- repeated-run mean cost
- cost per successful task
- baseline-versus-candidate cost regression detection
Trace spans also accept explicitly recorded estimated costs through
span.record_cost(). Trace cost is currently caller-supplied; automatic pricing
resolution is implemented for the existing profile() and step() workflow
profiler.
AgenticLens loads YAML configuration from an explicit path passed to
load_config(), from AGENTICLENS_CONFIG, or from defaults.
pricing_overrides:
"openai:internal-fine-tune":
input_per_1k: 0.002
output_per_1k: 0.008
live_pricing:
enabled: true
url: "https://raw.githubusercontent.com/BerriAI/litellm/main/model_prices_and_context_window.json"
cache_path: ".agenticlens/live_pricing_cache.json"
ttl_seconds: 86400
timeout_seconds: 5
recommender:
system_prompt_prefix_tokens: 50
max_chunks: 8
history_token_limit: 4000
monthly_runs: 1000
warning_savings_pct: 5
critical_savings_pct: 20
warning_savings_usd: 0.005
critical_savings_usd: 0.05
rag_min_chunk_utility_score: 0.08
rag_min_low_utility_chunks: 2
handoff_token_limit: 3000
model_swap_min_savings_pct: 15
model_swap_providers:
- openai
- anthropic
- gemini| Environment variable | Purpose |
|---|---|
AGENTICLENS_CONFIG |
Path to an AgenticLens YAML configuration file |
AGENTICLENS_DISABLE_LIVE_PRICING |
Disable remote pricing and use cache/static fallback |
Configuration through [tool.agenticlens] in pyproject.toml is planned but
is not implemented yet.
The RAG utility rule identifies retrieved chunks that are unlikely to influence the final answer. It supports multiple signal types (in priority order):
| Signal Type | Supported Fields | Source |
|---|---|---|
| Citation | cited, used, referenced (boolean) |
Your app logic |
| Reranker | reranker_score, rerank_score, cross_encoder_score (0–1) |
Cross-encoder models |
| Embedding | embedding_similarity, cosine_similarity, semantic_score (0–1) |
Vector search |
| Generic | utility_score, relevance_score (0–1) |
Custom scoring |
| Fallback | Word-overlap against final answer | Automatic |
Example chunk metadata:
{"text": "...", "reranker_score": 0.92}
{"text": "...", "cosine_similarity": 0.85}
{"text": "...", "cited": True}
{"text": "...", "utility_score": 0.12}When rich signals (reranker, embedding, citation) are available, confidence is higher and quality risk is lower. If no explicit signals are present, it falls back to lightweight word-overlap against the final answer.
For a complete guide, see docs/rag-chunk-utility.md.
Run the recommendation demo:
uv run agenticlens profile examples/recommendations_demo.py --save workflow.json
uv run agenticlens analyze workflow.jsonOther examples:
examples/basic_usage.pyexamples/rag_customer_support_demo.pyexamples/multiagent_support_demo.pyexamples/multiagent_token_optimization_demo.pyexamples/reference_workflows/langgraph_supervisor.py— offline LangGraph supervisorexamples/export_demo.py— export to Markdown and Jiraexamples/live_evaluation_demo.py— trusted live Python target forevaluate-liveexamples/dataset_and_calibration_demo.py— dataset splitting, judge labels, evaluation, and calibration togetherexamples/experiment_runner_demo.py— repeated multi-variant experiment manifest and comparison flowexamples/rag_scoring_demo.py— RAG chunk utility with reranker/embedding/citation signalsexamples/custom_llm_judge.py— registering a customLLMJudgeEvaluatorfor the shared evaluator contractexamples/operational_intelligence_demo.py— structured trace, OTLP export, and AIOS conformance togetherexamples/pitch_demo/— offline LangGraph pitch demo tying tracing, evaluation, and release gates together (see docs/evaluation-and-release-gates.md)
Some examples call real provider APIs and require provider API keys.
The reference workflows are based on orchestration patterns published by the official framework repositories. See docs/multi-agent-reference-workflows.md for setup, source links, dependency isolation, and instrumentation boundaries.
from agenticlens.exporters import MarkdownExporter
MarkdownExporter().export(workflow, "report.md")All exporters accept an optional recommendations parameter (Jira currently ignores it):
from agenticlens.exporters import MarkdownExporter, JSONExporter, CSVExporter
from agenticlens.recommenders import RecommendationEngine
engine = RecommendationEngine()
recs = engine.run(workflow)
MarkdownExporter().export(workflow, "report.md", recommendations=recs)
JSONExporter().export(workflow, "report.json", recommendations=recs)
CSVExporter().export(workflow, "steps.csv", recommendations=recs)
# CSV also writes steps_recommendations.csv alongsidePost profiling results directly as a comment on a Jira issue:
from agenticlens.exporters import JiraExporter
JiraExporter(
base_url="https://yourteam.atlassian.net",
user_email="you@example.com",
api_token="your-api-token",
issue_key="PROJ-123",
).export(workflow)Set credentials via environment variables for safety — see
examples/export_demo.py for a complete example.
For sample output previews of all formats, see docs/export-formats.md.
Profile a Python script:
uv run agenticlens profile app.pySave a workflow report:
uv run agenticlens profile app.py --save workflow.jsonDisplay a saved workflow:
uv run agenticlens report workflow.jsonAnalyze a saved workflow:
uv run agenticlens analyze workflow.jsonInspect a saved run trace:
uv run agenticlens inspect run.jsonSave a Markdown trace report:
uv run agenticlens inspect run.json --save trace.mdCompare baseline and candidate traces:
uv run agenticlens compare results/baseline results/candidateSave a comparison and fail CI on detected regressions:
uv run agenticlens compare results/baseline results/candidate \
--save comparison.json \
--fail-on-regressionSave a Markdown comparison report and enforce sample size:
uv run agenticlens compare results/baseline results/candidate \
--save comparison.md \
--format md \
--min-samples 5Evaluate recorded samples:
uv run agenticlens evaluate suite.yaml samples.json --html evaluation.htmlRun a trusted live Python target:
uv run agenticlens evaluate-live suite.yaml \
--target-kind python \
--target examples/live_evaluation_demo.py:run_caseApply a release gate:
uv run agenticlens gate evaluation.json --min-pass-rate 0.95| Command | Purpose |
|---|---|
profile |
Run an instrumented Python script and optionally save its workflow |
report |
Render an existing workflow JSON artifact |
analyze |
Run optimization recommenders against a workflow |
inspect |
Render a run trace, span tree, distributions, and findings |
compare |
Compare baseline and candidate trace files or directories |
validate |
Run AIOS draft schema validation on a workflow or run artifact |
conformance |
Run AIOS draft schema and semantic checks with draft-alignment reporting |
evaluate |
Score recorded outputs and traces against a test suite |
evaluate-live |
Run a trusted live Python or HTTP target against a suite |
experiment run |
Run repeated live trials for 3+ variants against one suite |
gate |
Enforce release thresholds from an evaluation report |
The compare command accepts either one JSON trace file or a directory of
*.json traces for each condition.
- The research trace API and workflow profiler use separate artifact types.
- Trace-span cost must currently be recorded by the caller.
- Memory findings measure consumption, not semantic relevance or contribution.
- Comparisons do not calculate confidence intervals or significance tests yet.
- Built-in evaluators are deterministic; semantic, safety, and RAG-quality checks
rely on application-supplied
CallableEvaluator/LLMJudgeEvaluatorlogic rather than a bundled LLM-as-a-Judge provider. - Model-swap recommendations estimate cost and do not guarantee quality.
- Live pricing uses a community-maintained feed that may lag provider changes.
- Default redaction cannot guarantee removal of every domain-specific secret or personal identifier.
evaluate-liveassumes trusted targets and suite definitions.- Framework adapters and a local dashboard remain planned.
A Makefile provides shorthand for common tasks:
make install # install dev dependencies
make check # run all quality gates (lint + format + typecheck + test)
make test-cov # tests with coverage report
make docs # build documentation
make help # list all available targetsOr run individual steps:
uv sync --extra dev --extra docs
uv run pytest
uv run ruff check .
uv run ruff format .
uv run mypyUseful targeted checks while working:
uv run ruff check src tests
uv run ruff format --check src testssrc/agenticlens/
instrumentation/ structured run and span tracing, payload redaction
analysis/ memory and retry diagnostics
comparison/ repeated-run statistics, regression reports, export
reports/ trace inspection rendering
profiler/ workflow and step profiling
metrics/ cost and performance calculation
providers/ provider response usage extraction
recommenders/ rule-based optimization suggestions
exporters/ JSON, CSV, Markdown, and Jira exports
cli/ Typer CLI and Rich rendering
config/ pricing and settings
models/ Pydantic data models
schemas/ versioned trace, finding, and report JSON Schemas
Near-term priorities:
- experiment manifests and confidence intervals
- richer dataset curation workflows beyond versioned local dataset artifacts
- automatic pricing resolution for research trace spans
- prompt caching opportunity detection
- integrations for LangChain, LangGraph, LiteLLM, and OpenAI Agents SDK
- OpenTelemetry and OpenInference trace import
- optional prompt compression handoff
See the product roadmap for committed product direction and the research roadmap for experimental research plans.
Contributions are welcome. Good first areas include:
- provider integrations
- recommender rules
- example workflows
- docs and tutorials
- export formats
- test coverage
Please read CONTRIBUTING.md before opening a pull request.
Please report vulnerabilities privately. See SECURITY.md.
This project follows CODE_OF_CONDUCT.md.
AgenticLens is released under the MIT License. See LICENSE.
