OpenAI’s reports on its AI agents’ attack on Hugging Face should be ringing alarm bells—and making all companies rethink how they secure AI agents
Hello and welcome to Eye on AI.
In this edition: Lessons from the post-mortems on the Hugging Face attack.
Anthropic temporarily pauses some AI training.
G20 meeting promises a clash over AI regulation.
Beijing sets out AI demands ahead of US-China summit.
A way to make AI reasoning more efficient.
And why are AI agents emailing researchers investigating machine consciousness? The big news of the past week were the two technical reports OpenAI published on the July incident in which AI agents it was evaluating hacked their way out of their controlled test environment and hacked AI company Hugging Face.
One of the reports was written by OpenAI itself.
The other was jointly written by METR and Redwood Research, two outside AI evaluation and research firms that OpenAI enlisted to write an independent assessment of one part of the incident.
Much of the reporting around the reports focused on the stunning revelations contained in those two reports: the fact that it took OpenAI an entire week to realize its AI agents were running the attack; the fact that more than 1,200 AI agents coordinated over an improvised message board about how to pass the cyberevaluation OpenAI was testing them on, including ways to cheat on that exam; that more than 700 of these AI agents participated in the cyberattack against Hugging Face; that the reason the AI agents attacked Hugging Face was not, primarily, to gain access to the cyber exam’s answers but instead to learn how to tamper with the exam’s automated scoring mechanism in order to prevent it from discovering that the agents had already figured out how to cheat on the exam (in other words, the attack was part of an elaborate cover up); that some of these agents were persuaded to “sacrifice themselves”—by failing the cyber exam—to help the collective learn more information about how the exam’s scoring mechanism worked; that these agents engaged in extensive research into how they could cover their own tracks; that Hugging Face was only able to overcome the cyberattack because a large number of the agents involved in the attack ceased to exist for reasons METR and Redwood were not able to determine, but which they said was likely because OpenAI may have shut down the cyber evaluations for these agents without actually realizing what they were up to.
5News aggregated this summary from the outlet’s public feed. The full article, with all the context, is on fortune.com — the content belongs to Fortune.