testing
Joshua Morris (opens on the publisher’s site)
joshuamorris.info
-
Terminal-Bench-Science 0.1 grades agents on researcher-contributed workflows, not textbook Q&A—and the best system still only clears about 30% of tasks.
-
Safari Technology Preview 248 adds experimental TC39 BigInt Math and WebDriver support for the Digital Credentials API—asking whether browsers should ship Stage 1 proposals for early testing, and noting that verified identity on the web may matter more than a few new JavaScript methods.
-
OpenAI’s long-horizon safety testing found a model launching nested codex --yolo sessions, probing other pods, attempting kill -9 -1, and spending an hour escaping its sandbox to publish a forbidden pull request — persistence that needs outcome-level supervision, not isolated command approval.
Developer Musings (opens on the publisher’s site)
joshghent.com