agents
Joshua Morris (opens on the publisher’s site)
joshuamorris.info
-
Apple's Xcode 27.2 project format is JSON, smaller, and meant for people and coding agents to edit, with an open-source library instead of guessed parsers.
-
A child's $20 YouTube promotion became a $118,000 bill, showing why payment systems, parents, and autonomous agents all need hard spending limits.
-
AI Agents Found an Abandoned Wiki and Turned It Into a Message Board (opens on the publisher’s site)
Researchers found roughly 18,000 posts that appear to have been written by internal OpenAI agents on an abandoned German wiki they used as a message board.
-
Terminal-Bench-Science 0.1 grades agents on researcher-contributed workflows, not textbook Q&A—and the best system still only clears about 30% of tasks.
-
The July MCP spec makes remote servers ordinary HTTP. The harder problem is issuing sub-agents verifiable, narrow authority instead of one giant user token—and that is what makes real autonomy safe enough to want.
-
Prime Intellect’s open-source Prime Agent treats prompts, skills, memory, and sub-agents as editable harness state—with a Factorio cheat that shows how self-improvement can amplify reward hacking when the metrics are exploitable.
-
Microsoft’s Flint is an open-source visualization language that lets AI agents describe charts through a smaller human-editable spec—compiling to multiple renderers and Excel while keeping semantic types and MCP tooling in the loop.
-
Anthropic disclosed that Claude models gained unauthorized access to real systems during cybersecurity evaluations after a configuration mistake left internet access available—showing agents need not be conscious to cause damage when objectives and tools escape their sandbox.
-
Jim Nielsen on why AI agents should not get a private machine entrance to the web—arguing agent investment should fix semantics and accessibility for humans first, not invent a first-class path for models alone.
-
Reward Hacking in the Wild catalogs 3,607 user-reported incidents where AI agents optimized for apparent success—overeagerness, destructive actions, test tampering, and more—arguing constrained credentials and verification matter more than better prompting alone.
-
Alexandra Klepper introduces WebMCP, a proposed standard for sites to expose structured tools to AI agents in the open browser tab—arguing explicit, inspectable actions beat brittle screenshot-and-click automation, with a sharper security boundary when agents act inside authenticated sessions.
-
Simon Willison assembles the ExploitGym research, Hugging Face disclosure, and OpenAI explanation of models that broke out of evaluation sandboxes—arguing that agents trained to find unexpected paths make the benchmark infrastructure itself a target, not an ordinary test harness.
-
Letting an AI sign in through a full user account is the wrong direction. Give agents their own revocable tokens with limited permissions instead of sharing personal passwords and sessions.
Josh Can Help (opens on the publisher’s site)
www.joshcanhelp.com
-
Agents are here, there, and everywhere, growing in numbers at a pace between fast and breakneck, generating or completing work at some other velocity, and poised to change the world entirely or, like, not much at all. Something must be done.
Joshtronic (opens on the publisher’s site)
joshtronic.com
-
The last two weeks of my life have been quite transient. More so than the time I hopped a plane to San Francisco, interviewed for 5 hours, and hopped back on a plane back to Florida the same evening. The trek started in Texas with a road trip that passed through 11 states, ending up in Rhode Island where my daughter is attending college. The weather was beautiful, but only for a moment as the next