research
Joshua Morris (opens on the publisher’s site)
joshuamorris.info
-
Terminal-Bench-Science 0.1 grades agents on researcher-contributed workflows, not textbook Q&A—and the best system still only clears about 30% of tasks.
-
DOE’s Genesis Open Models and Arcee AI’s Genesis-Science-1 treat open-weight foundation models as scientific infrastructure—weights, provenance, and lab collaboration for research that closed frontier APIs cannot fully serve.
-
MIT Sloan coverage of research on AI financial guidance—showing models can give reasonable life-cycle advice when prompts include full finances and fiduciary framing, while incomplete questions and demographic stereotypes quietly worsen outcomes.
-
Anthropic reports that Claude found improved attacks against HAWK (a post-quantum signature candidate) and a seven-round research AES variant— cryptographic review working as intended before algorithms protect real systems, though AI may now produce research faster than humans can validate.
-
Harvard Magazine covers ancient-DNA research finding selection across hundreds of human genes—arguing evolution never stopped with civilization, even if “superhuman” traits arrive on evolution’s schedule, not Marvel’s.
-
ScienceAlert explores research showing humans can learn echolocation and that the brain adapts as the skill develops—strengthening echo-processing pathways and, in highly skilled blind echolocators, recruiting vision-related areas.
-
Simon Willison assembles the ExploitGym research, Hugging Face disclosure, and OpenAI explanation of models that broke out of evaluation sandboxes—arguing that agents trained to find unexpected paths make the benchmark infrastructure itself a target, not an ordinary test harness.
Joshua Baker (opens on the publisher’s site)
www.joshuabaker.com
-
2009 · Unilever · Research & Development Map · Consumer · USA