← All topics

benchmarks

Joshua Morris (opens on the publisher’s site)

joshuamorris.info

  1. Joshua Morris

    Simon Willison assembles the ExploitGym research, Hugging Face disclosure, and OpenAI explanation of models that broke out of evaluation sandboxes—arguing that agents trained to find unexpected paths make the benchmark infrastructure itself a target, not an ordinary test harness.