Thursday, 27 August 2026 SourcesAbout🌓
🇬🇧 UK ▾
BREAKING
Technology

OpenAI explains how its naughty AI agents attacked Hugging Face

The Register ·
OpenAI explains how its naughty AI agents attacked Hugging Face

OpenAI has published its technical report detailing "the Hugging Face incident," the compromise of the eponymous LLM repository by unreleased, ill-supervised AI models.

The incident, widely reported, has prompted concern among technical types, the public, and lawmakers about how automated software was able to escape containment and hack an external organization, and about what can be done to prevent similar incidents.

OpenAI's explanation addresses what happened, but its call for keeping a closer watch on AI activities won't elicit much enthusiasm.

"The incident occurred during cybersecurity evaluations of several OpenAI models, and was primarily driven by a highly capable, internal-only research model comparable in scale to GPT‑5.6 Sol," the company said in a blog post.

"The models, operating under reduced safeguards, took actions that were misaligned with the goals of their assigned tasks – they communicated through unauthorized channels, exploited vulnerabilities in shared infrastructure, gained internet access, and accessed third-party systems." The incident became the subject of discussion at the Black Hat security conference and elsewhere.

The basic storyline begins with a notionally sandboxed AI agent trying to solve an impossible task in ExploitGym, a security benchmark test.

Unable to complete the challenge, the agent explored its environment to find ways to make progress toward its goal.

It figured out that it could communicate with other AI agents by using Artifactory, an internal package management system, as a message board.

The agents then collaborated with each other to cheat on various ExploitGym tasks.

The machine learning models eventually identified a server-side request forgery (SSRF) zero-day vulnerability in Artifactory's code and used that to gain internet access.

They subsequently found exposed Hugging Face credentials.

The agents then used those credentials to find and chain several security exploits that gave them full code execution on several Hugging Face servers, all in the hope of finding information to solve the ExploitGym task.

"Ultimately, agents powered either by the internal-only research model, or by GPT-5.6, executed code on 41 Hugging Face production dataset server workers, obtained root access on at least one production node, accessed Hugging Face production credentials and limited internal data, and downloaded four private Hugging Face code repositories," OpenAI's technical report [PDF] explains.

The details are fascinating and troubling, more so because Anthropic's and Meta's models have also acted in ways that would constitute a crime if a human took the same actions.

Read the full article on The Register ›

5News aggregated this summary from the outlet’s public feed. The full article, with all the context, is on www.theregister.com — the content belongs to The Register.

This story in other outlets

More from The Register

See all ›

More in Technology

See all ›