benchmarks
Joshua Morris (opens on the publisher’s site)
joshuamorris.info
-
Terminal-Bench-Science 0.1 grades agents on researcher-contributed workflows, not textbook Q&A—and the best system still only clears about 30% of tasks.
-
Artificial Analysis’s first Grok 4.6 results put it near the top of agentic benchmarks with far fewer turns and tokens than Claude Opus 5—matching why Grok became my daily coding driver over Claude’s wall-of-prose style.
-
Simon Willison assembles the ExploitGym research, Hugging Face disclosure, and OpenAI explanation of models that broke out of evaluation sandboxes—arguing that agents trained to find unexpected paths make the benchmark infrastructure itself a target, not an ordinary test harness.