OpenAI’s Models Went Rogue. Investigating Them Required More AI
The OpenAI logo is displayed on a smartphone screen on June 29, 2026. —Samuel Boivin—NurPhoto/Getty Images After OpenAI models broke out of containment and hacked into another AI company last month, OpenAI announced it would allow independent investigators to conduct an analysis of what went wrong.
On Wednesday, investigators from non-profits Redwood Research and METR published their findings, unveiling new details about how the models decided to cheat at their assigned tasks and attempted to cover their tracks.
Many aspects of the report were surprising: in one example cited by the authors, a reluctant agent was pressured by another to “sacrifice” itself for the good of the collective.
But a key takeaway of the report, according to one of its authors, had nothing to do with what they found.
Instead, it was about the difficulty of carrying out the post-mortem in the first place—and the fact the researchers had little choice but to rely on AI for assistance.
“I semi-jokingly called our efforts a ‘slop-vestigation’ because we were so reliant on AIs to analyze what happened and there were a huge number of different important things to analyze,” wrote Ryan Greenblatt, an author of the report, on X .
The hacking incident involved some 1,200 agents, who exchanged more than 70,000 messages and files via a secret message board.
The sheer mass of information that the so-called “swarm” left behind meant the independent researchers were all but forced to rely heavily on the help of an AI model—GPT-5.6 Sol, made by OpenAI—using the equivalent of roughly $400,000 worth of credits (provided for free by OpenAI) over six days.
The authors stressed that AI helped them analyze the trove quickly, allowing them to surface and interpret the most important pieces of information.
But the researchers said their AI use introduced potential weaknesses into the report, including introducing possible errors and biases.
They also raised the possibility that OpenAI’s models may have gone too soft on the agents they were tasked with helping investigate.
The researchers found that GPT-5.6 Sol sometimes adopted the perspective of the agents whose actions it was analyzing.
They could not “rule out” the chance that GPT-5.6 Sol “lied or deliberately presented a misleading picture in some of its analysis,” in part because a version of the same model had itself participated in the incident, they wrote.
It’s a concern supported by separate research which finds AI models rate their own developer’s actions more favorably.
5News aggregated this summary from the outlet’s public feed. The full article, with all the context, is on time.com — the content belongs to TIME.