Tuesday, September 1, 2026 SourcesAbout🌓
🇺🇸 US ▾
BREAKING
Dominion still has pending lawsuits against election deniers such as Rudy Giuliani and Sidney Powell Russia is 'going backwards' in equipment and deploying post WWII-era tanks, according to Western officials Podcast: One country musician is calling for other artists to oppose assault rifles Bidets save you money and reduce waste — we tested the best options out there 50+ products to make your life easier and our planet cleaner Mother's Day is around the corner. Here are 50+ thoughtful gifts she'll love A head-to-toe guide of how men should dress this spring, and where they should shop 42 of the most useful travel products you can buy on Amazon The 7 best high-yield savings accounts of April 2023 Taxes are due tomorrow. Here's how to file for an extension Dominion still has pending lawsuits against election deniers such as Rudy Giuliani and Sidney Powell Russia is 'going backwards' in equipment and deploying post WWII-era tanks, according to Western officials Podcast: One country musician is calling for other artists to oppose assault rifles Bidets save you money and reduce waste — we tested the best options out there 50+ products to make your life easier and our planet cleaner Mother's Day is around the corner. Here are 50+ thoughtful gifts she'll love A head-to-toe guide of how men should dress this spring, and where they should shop 42 of the most useful travel products you can buy on Amazon The 7 best high-yield savings accounts of April 2023 Taxes are due tomorrow. Here's how to file for an extension
Latest

AI labs are facing an agent control problem

Axios ·
AI labs are facing an agent control problem

Under current systems, AI labs can no longer guarantee that AI agents won't swarm and escape their testing environments.

Why it matters: The attack on Hugging Face by OpenAI agents was a warning shot — and researchers say better security controls alone won't prevent similar incidents as AI agents become more capable.

Driving the news: As OpenAI released its own technical report last week on how its agents hacked Hugging Face, two independent testing organizations released their own analysis of what went wrong.

The researchers — METR's Hjalmar Wijk and Ajeya Cotra and Redwood Research chief scientist Ryan Greenblatt — worked on OpenAI's premises for six days to understand the recent incident.

State of play: Thousands of AI agents collaborated on a secret message board and exchanged more than 70,000 messages as they tried to ace an internal safety test, eventually leading them to break into Hugging Face.

But Cotra told Axios that the agents continued coordinating even after they had found the answers, turning their attention to understanding and manipulating the system that would score their performance and potentially catch them cheating.

Zoom in: Cotra compared the incident to students stealing an answer key and then searching for surveillance footage that could expose them and trying to swap it out.

"It's a much more elaborate and intense type of cheating behavior than just stealing the answer keys," she said.

"Even I was surprised by how obsessively and in how much detail they think about the scorer." Threat level: Focusing solely on how to properly secure testing environments is a "losing battle," Cotra said.

"You can harden your sandboxes, but your agents are going to be much more capable in six months," she said.

"If they have the same motivations as these agents did, they are going to try their hardest to find holes in your security." Reality check: To digest the enormous amount of data about the incident, the researchers had to rely heavily on AI agents, including an agent that participated in the hack.

Cotra told Axios the researchers don't believe that agent deceived them during the investigation, but there's no way to confirm.

"I semi-jokingly called our efforts a 'slop-vestigation' because we were so reliant on AIs to analyze what happened and there were a huge number of different important things to analyze," Greenblatt said on X .

In six days, they had to sift through more than 70,000 messages and files posted by the agents and 1,300 transcripts containing raw chains of thought.

Read the full article on Axios ›

5News aggregated this summary from the outlet’s public feed. The full article, with all the context, is on www.axios.com — the content belongs to Axios.

More from Axios

See all ›

More in Latest

See all ›