Research engineer working on AI evaluation, red teaming and agent reliability.
I study how to evaluate agentic systems when the evidence itself is unreliable: task outcomes vary across reruns, automatic judges make systematic errors, and a final success label can hide failures in trajectories, tools, providers or the evaluation infrastructure.
- 🎓 MSc Artificial Intelligence for Sustainable Development, UCL Computer Science (2025–26) · BEng Automation, Xi'an Jiaotong University
- 🧪 Research Intern, Holistic AI, London (Jun–Sep 2026): LLM / agent red-teaming and evaluation
- 📄 Routing Is Least Learnable Where It Is Most Valuable: Bounds on Representation Routing for Web Agents — accepted at the EMNLP 2026 Workshop REALM (OpenReview)
Selected work
| Project | What it is |
|---|---|
| Cost-Aware-Routing-for-Web-Usage-Agents | Preregistered study of which page representation (DOM / Set-of-Marks / vision) a web agent should read, across 3 model families × 6 observation modes. Main result: representations carry real, complementary value, but that value is hardest to predict exactly where it matters most. |
| redteam-under-test | Local-first LLM red-team measurement stack: target-conditioned attacks, 103 end-to-end-validated plugins, and the judge itself measured against independent gold labels. |
| agent-redteam-lab | Tool-using agent security: a read-only MCP trace analyzer for boundary crossing, data egress, destructive actions and confused-deputy behaviour, plus a replay harness. |
| coding-agent-guardrails | Rules, hooks and skills that keep AI coding agents (Claude Code, Codex CLI) reliable: git guardrails, evidence before a verdict, session handover, cross-model review. |
| constrained-agent-runner | Control plane for a web chat agent to run constrained, auditable jobs on your own machine, with GitHub as durable transport and hard capability boundaries. |
| FinQA | Controlled post-training attribution on Qwen3 (0.6B→14B): separating protocol alignment, answer-format SFT, explicit reasoning and GRPO before crediting a score change to capability. |
Links · Portfolio · Research CV · Poster & showcase · LinkedIn · jiaming.wei.ai@outlook.com · 中文主页



