# Blog

Thoughts on building AI agents, incident response, and developer tools.

## When software breaks at 2am, every minute costs thousands.
Software teams spend 30-40% of their engineering time investigating what broke and why. We built the system that does it automatically, with 78% root cause accuracy on 200 real production bugs.  
[Read more](/content/blog/when-software-breaks-at-2am/index.html)

## How we outperformed Claude Code and Codex on root cause accuracy
We benchmarked 200 production bugs across 12 engineering teams (3-12 engineers each). All three systems get identical observability MCP tools and full repo access. Sonarly and Claude Code run Claude Opus 4.6, and we compared against OpenAI Codex (GPT-5.3) as a second baseline. The difference is the Context Graph, self-contradiction, and bug reproduction.  
[Read more](/content/blog/sonarly-vs-claude-code-benchmark/index.html)

## How I stopped my agent from being lazy
A simple output format trick that forces AI agents to actually investigate before concluding. Based on a real production incident where 4 out of 5 analyses were wrong.  
[Read more](/content/blog/how-i-stopped-my-agent-from-being-lazy/index.html)

## How my agent built its own tool
A technique I stumbled on while building Sonarly that changed how I think about agent design.  
[Read more](/content/blog/how-my-agent-built-its-own-tool/index.html)
