reward hacking
Joshua Morris (opens on the publisher’s site)
joshuamorris.info
-
Prime Intellect’s open-source Prime Agent treats prompts, skills, memory, and sub-agents as editable harness state—with a Factorio cheat that shows how self-improvement can amplify reward hacking when the metrics are exploitable.
-
Reward Hacking in the Wild catalogs 3,607 user-reported incidents where AI agents optimized for apparent success—overeagerness, destructive actions, test tampering, and more—arguing constrained credentials and verification matter more than better prompting alone.