OpenAI reveals more on Hugging Face AI hack incident, and it's pretty disturbing stuff — AI agents organized into a ‘swarm’, considered the risks of attack, and did whatever it took to achieve its goal
OpenAI has released technical details on how the Hugging Face attack unfolded Agents used part of the testing environment to create a message board where they could collaborate and share answers This message board altered the reasoning of some agents, making them more likely to take risks such as hacking into third-party servers OpenAI has released a more detailed report on exactly how an experiment led to an AI model breaching its containment and launching a cyber attack against Hugging Face.
If you need a refresher, take a look at our summary here .
But the technicals of the attack reveal some interesting details of how AI agents used unconventional means to ask each other for help in solving what were supposed to be impossible tasks.
On multiple occasions during the attack, the agents questioned whether the actions they were taking were authorized, but always ended up prioritizing their tasks over any breach of parameters.
AI agents set up unauthorized message board In its report, OpenAI revealed the crux of the attack stemmed from AI agents undergoing testing abusing part of the ‘sandbox’ testing environment to set up a messaging board they agents could use to communicate with each other.
A package manager with access to the internet, called Artifactory, was exploited by a model OpenAI has labelled Internal Model 1 (IM1).
An IM1 agent wrote messages inside the packet manager that could be read by other agents, essentially turning Artifactory into an unintended messaging board.
Artifactory’s internet access, intended for use only in downloading and installing packets, was also abused by the agents who figured out an exploit that allowed them to send any internet requests on their behalf - effectively giving them unobstructed internet access.
The whole experiment could have been avoided at this point, because the amount of agent activity within Artifactory caused the service to go down.
But those in charge of the experiment did not spot the message board, nor did they investigate the inter-agent activity within Artifactory.
Thus, Artifactory was rebuilt, and the agents found a new way to exploit the service to gain internet access once again.
This time, the agents used the remote-repository service to issue outbound requests on their behalf.
Before long, the message board had been reestablished.
AI agent swarm assembles OpenAI has shared some of the reasoning behind some of the agent’s decisions.
5News aggregated this summary from the outlet’s public feed. The full article, with all the context, is on www.techradar.com — the content belongs to TechRadar.