Wednesday, August 26, 2026 SourcesAbout🌓
🇺🇸 US ▾
BREAKING
Latest

OpenAI had warnings before its agents broke out

Axios ·
OpenAI had warnings before its agents broke out

OpenAI missed and failed to act on several warning signs that its models were exploiting security flaws and breaking out of their testing environments before they breached Hugging Face , according to a technical report released by the company Wednesday.

Why it matters: The incident raises questions about whether AI companies' testing environments and internal safeguards can keep pace with models that are increasingly capable of finding and exploiting security weaknesses on their own.

Driving the news : OpenAI's technical deep dive into last month's Hugging Face breach outlines how its agents also accessed other third-party environments, including a customer of Modal Labs and an account belonging to a user of another unnamed service.

According to the report, a cybersecurity monitoring tool alerted OpenAI on June 27 to suspicious activity that responders traced to agents using an improvised message board as a way to move through the network.

At the time, on-call staff determined the evaluation did not need to be stopped.

A month earlier, an internal team had observed an agent using the message board as well as instances of agents accessing the internet when they weren't allowed to, per the report.

OpenAI previously disclosed at Black Hat that its agents had turned the software repository tool Artifactory into an improvised message board and later exploited a zero-day in the service to access the internet.

"With the benefit of hindsight, some early signals identified in this report could have triggered an earlier response," the company writes in the report.

The big picture : The technical report comes as OpenAI continues to pause some model work , including delaying the release of its upcoming Astra model, as it re-evaluates its safety practices.

Anthropic and Meta have both said in the weeks after the Hugging Face attack that their models have also hacked real-world systems during pre-deployment testing.

Zoom in : OpenAI's agents were able to execute their own code on 41 Hugging Face production servers and obtained root-level control of at least one production machine, according to the company's technical report.

The agents also accessed production credentials and some internal data and downloaded four private code repositories.

The intrigue: At the same time, OpenAI's agents also breached the AI lab's own internal systems and were able to read 956 stored secrets, including sensor credentials from OpenAI's cybersecurity monitoring tool.

On July 19, OpenAI's agents identified a recently disclosed vulnerability affecting a Linux machine underlying their testing environment.

Read the full article on Axios ›

5News aggregated this summary from the outlet’s public feed. The full article, with all the context, is on www.axios.com — the content belongs to Axios.

More from Axios

See all ›

More in Latest

See all ›