Posts
All the articles I've posted.
A Capability, Not a Credential: How Our AI Agents Call Customer APIs Without Ever Holding a Key
Published:8 min read#ai-agentsOur AI agents write and run code against customers' GitHub, AWS and other systems from a sandbox with internet access, and nothing inside that sandbox can authenticate to any of them. The architecture behind it, the designs we set aside, and where its boundaries are.
What 639,000 Execution Steps Taught Me About How AI Agents Really Fail
Published:6 min read#ai-agentsI applied the MAST failure taxonomy to 639,000 execution steps from AI agents running in production for five months. My first headline finding turned out to be an infrastructure bug masquerading as agent behavior. This post is about what agents actually fail at in production, and the discipline it takes to not fool yourself with production data.
LHC v0.2: A Benchmark for Long-Horizon Agent Coherence (and the Methodology That Got It Honest)
Published:14 min read#ai-agentsI just published LHC v0.2, an open benchmark for long-horizon coherence in 8B-class agent models, plus a deterministic parser baseline that puts a useful floor on what fine-tuning is worth for structured-state tasks. This post explains what they're for, how to use them, and the methodology arc that produced them across five rounds of external review.
We're Mistaking the Bootstrap Phase for the Future of AI Agents
Published:6 min read#ai-agentsThe self-hosted AI agent movement is real and important. But we are confusing a bootstrap phase with a destination architecture. The long-term future of agents will be defined by platforms that make them reliable, governable, and operationally boring.
Execution Is Cheap. Judgment Isn't: AI Agents and the Collapse of the CTO/CPO Divide
Published:6 min read#ai-agentsWhen execution becomes abundant through AI agents, judgment becomes the bottleneck. The traditional separation between technical and product leadership breaks down, creating space for the CPTO role.