A2A green-agent orchestrator for evaluating agents on the AppWorld benchmark, built on the AgentBeats SDK
-
Updated
Jul 1, 2026 - Python
A2A green-agent orchestrator for evaluating agents on the AppWorld benchmark, built on the AgentBeats SDK
Harness-localized credit assignment and certificate-guided RL for AppWorld agents
Experimental code and selected traces for automatic task discovery and verification
Pi-Bench research workspace for proactive-agent evaluation, memory-backend integration, experiment orchestration, and diagnostics.
Minimal AppWorld agent for auto-agent optimization experiments — target agent that gets autonomously improved
Task-conditioned tail reliability for tool-using agents under equivalent interfaces
To associate your repository with the appworld topic, visit your repo's landing page and select "manage topics."