Blog
Thoughts on building AI agents, incident response, and developer tools.
When software breaks at 2am, every minute costs thousands.
Software teams spend 30-40% of their engineering time investigating what broke and why. We built the system that does it automatically, with 78% root cause accuracy on 200 real production bugs.
Read more
How we outperformed Claude Code and Codex on root cause accuracy
We benchmarked 200 production bugs across 12 engineering teams (3-12 engineers each). All three systems get identical observability MCP tools and full repo access. Sonarly and Claude Code run Claude Opus 4.6, and we compared against OpenAI Codex (GPT-5.3) as a second baseline. The difference is the Context Graph, self-contradiction, and bug reproduction.
Read more
How I stopped my agent from being lazy
A simple output format trick that forces AI agents to actually investigate before concluding. Based on a real production incident where 4 out of 5 analyses were wrong.
Read more
How my agent built its own tool
A technique I stumbled on while building Sonarly that changed how I think about agent design.
Read more