benchmarks
Joshua Morris (opens on the publisher’s site)
joshuamorris.info
-
Terminal-Bench-Science 0.1 grades agents on researcher-contributed workflows, not textbook Q&A—and the best system still only clears about 30% of tasks.
-
Artificial Analysis’s first Grok 4.6 results put it near the top of agentic benchmarks with far fewer turns and tokens than Claude Opus 5—matching why Grok became my daily coding driver over Claude’s wall-of-prose style.