Agent benchmarks

Who builds better, and how we measure it.

Controlled head-to-head runs between coding agents and models. Same prompt, same repo constraints, same success criteria, then an honest write-up of what shipped, what broke, and what it cost in tokens and time.