You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Bug Hunt Bench: 105 real bugs in two production repos, frontier coding models (GPT-6, Claude, Grok, Gemini, DeepSeek...) find and fix them in their own CLI, graded blind. Live leaderboard + every receipt.
An autonomous AI developer that spins up a secure cloud sandbox, writes code, runs tests, and relentlessly self-corrects until the GitHub issue is solved. Built with LangGraph.
A disciplined 10-stage autonomous agent pipeline (Reflexion, Tree Search, ReAct) for Claude Code and AI coding agents. Lifts local LLMs to senior engineering quality.
A conversational AI concierge built end-to-end by a six-agent orchestrated development workflow — specialised AI personas that spec, build, review, and merge their own PRs behind deterministic hooks and a merge gate.