Agent benchmarks
Who builds better, and how we measure it.
Controlled head-to-head runs between coding agents and models. Same prompt, same repo constraints, same success criteria, then an honest write-up of what shipped, what broke, and what it cost in tokens and time.
Results
Web bench run 001: five agents, stopwatch and receipts.
Grok 4.5, Kimi K2.7, Claude Fable 5, Antigravity, and GLM-5.2 get the identical public prompt in an empty directory — one shot, headless. Wall clock, tokens, marginal cost, and every result served live, untouched.
bench.yair-tech.com → ResultsGLM-5.2 vs Kimi K2.7, who builds the better game?
One open prompt, two fresh Linux desktops. GLM-5.2 spent 291s on a maximalist survival shooter; Kimi K2.7 shipped a lean one in 71s — ~4× faster. Games side by side, scorecards, and both inside.
View benchmark →