Wednesday, August 26, 2026 SourcesAbout🌓
🇺🇸 US ▾
BREAKING
Business

OpenAI, independent firms publish reports on rogue AI attack on Hugging Face. Here are the main takeaways—and what OpenAI still hasn’t disclosed.

Fortune ·
OpenAI, independent firms publish reports on rogue AI attack on Hugging Face. Here are the main takeaways—and what OpenAI still hasn’t disclosed.

OpenAI today published the findings of its internal investigation into the July incident in which several AI models it was testing hacked their way out of their test environment and launched a cyberattack against the AI company Hugging Face.

Although many details of the rogue AI incident have already been made public by OpenAI, there are a few new items disclosed in the 37-page technical post-mortem.

Also today, independent research firms METR and Redwood Research published a 91-page analysis of the event.

OpenAI asked METR and Redwood to perform the analysis, but only to look at the events that occurred between July 7 and July 13, which is the time period during which many key events leading to the incident occurred.

The METR and Redwood report focuses on how the agents collaborated on a secret messaging board to execute the attack, as OpenAI first disclosed in an August 5 presentation at the Black Hat security conference.

OpenAI’s report contains the full account of what happened before the attack through to the days that followed.

OpenAI was not aware its agents were hacking Hugging Face Among the main takeaways from OpenAI’s report is that the company did not know its agents had breached Hugging Face, an online repository for AI models and datasets, until a week after the event happened.

The ability to monitor and identify unwanted behavior by AI agents in real-time or near real-time is critical to preventing another breach like this—perhaps one with even more serious consequences, such as an attack on a bank or hospital.

The report makes it clear that OpenAI’s monitoring systems were inadequate and failed to alert the AI researchers conducting the cybersecurity evaluation that its AI agents were engaging in unintended and potentially dangerous activities.

OpenAI has said that it has since improved its monitoring of what the agents it is training are doing, including much more scrutiny of the AI agents’ so-called “chain of thought.” This is the kind of “thinking aloud” that an AI model performs, reasoning in text about what it is trying to do and how it is approaching a task.

OpenAI also said it is monitoring more closely what tools the AI agents are accessing and what actions they are taking.

The activities leading up to the breach began in May, as OpenAI staffers revealed at Black Hat.

But OpenAI’s investigation found, according to the report, that the heart of the attack began on July 8, and reached its apex between July 11 and July 13.

Hugging Face disclosed the incident on July 16, and OpenAI reached out to the company on July 17 “in its capacity as a Hugging Face customer to assess any impact to OpenAI data.” In a stunning turn of events, OpenAI soon learned it was the culprit.

Read the full article on Fortune ›

5News aggregated this summary from the outlet’s public feed. The full article, with all the context, is on fortune.com — the content belongs to Fortune.

This story in other outlets

More from Fortune

See all ›

More in Business

See all ›