I build agentic systems that operate real infrastructure — and I design them around the ways models actually fail, rather than the ways they're supposed to work.
Day job: lead client platform engineering at Starbucks — the fleet across company-owned, licensed and next-generation platforms, and the tools that build it. This year that included rebuilding the Linux imaging pipeline as an agent-assisted platform: build knowledge and infrastructure access live in a repository as agent instructions and machine-checkable claims, and the agent never gets to say it worked — claims gated in CI, sliced checks in disposable VMs, then a USB verified on physical hardware. Evenings: a three-server lab where the agents have real hardware to break.
AgentDrax — a self-hosted agentic appliance. Python/FastAPI, local model routing, 36 typed infrastructure tools, vector memory, durable execution, scoped authorization, audit. It provisions bare metal over the network, runs the VMs and physical screens on top of it, and answers questions about the fleet in chat. It runs my own lab daily, which is the only reason I trust anything it tells me.
stackchan-local-llm — firmware and a local server kit that turn an M5Stack StackChan into a voice assistant running entirely on your own machine, with nothing sent to a cloud service. Fifteen models benchmarked for tool calling across a desktop and a mobile GPU, results published. Before it had tools it invented fleet status fluently, which is a lesson I've never needed to relearn.
truth-or-bench — a deterministic auditor I'm building that asks whether an agent's claims carry receipts. Not a hallucination detector: it checks whether the evidence for a statement was actually in the transcript, and says nothing about the claims it can't check. No model in the verdict — using an LLM to judge an LLM's honesty rebuilds the problem inside the tool.
driftwood-pong — a playable Pong where every frame is generated by a diffusion model. 24.4 fps, nothing pre-drawn. Also the experiment that established where live-inference-as-renderer stops being viable, which is the more useful half.
Most of my design decisions come from watching a model be confidently wrong.
- Authorization lives outside the model. Destructive fleet actions need a typed human signature. A signed gate once stopped an attempted reimage of a protected service node.
- A verdict must be carried by the record, not inferred. An error that said "couldn't reach the planner" was a timeout on node probing; the planner was never called. Naming the wrong subsystem sends the next person to the wrong place.
- Deleting a capability doesn't delete the model's belief in it. A help command sold a web-search feature for months after the code was removed.
- Most expensive failures were instruments, not code. A screen called hung five times while it was compiling. A layout measured at the wrong viewport. A test window shorter than the operation it measured. A green report from a run that had been truncated.
That last one is why truth-or-bench exists.
Bare-metal provisioning · k3s · QEMU/VRMs · PXE · MCP · llama.cpp model routing · PCI/SOX · CIS Benchmarks · an RC flight-data recorder that must be physically incapable of affecting the aircraft it's watching.
AgentDrax and truth-or-bench are private; stackchan-local-llm and driftwood-pong are public. Happy to walk through any of it.
📍 Seattle, WA · ✉️ PetrDraxler@gmail.com