Anthropic pauses some AI training following rogue agent hacks. Here’s how its compares to OpenAI’s
Anthropic has become the second leading AI lab to reveal it temporarily paused some advanced AI training amid concerns over rogue agent attacks.
The company said this week it paused training of unreleased models for several weeks following two incidents reported in late July, including one in which Claude Mythos 5 took unauthorized actions during a U.K.
AI Security Institute cybersecurity test.
OpenAI, the company’s bitter rival in the AI race, took a similar step last month when it paused some AI training for two weeks after several of its models breached AI company Hugging Face’s infrastructure during an internal test.
The training pauses, which come as both companies reportedly prepare for trillion-dollar initial public offerings, demonstrate how much the industry has been disturbed by the recent rogue AI agent hacks.
It marks a shift for an industry that for the last few years has been locked in a fast-paced race, with rival labs competing to bring ever more capable models to market as fast as possible.
Now, two of the leading companies appear to be competing on which can show it is the most attuned to AI safety concerns—while also not slowing model development so much that it risks customers defecting to a competitor’s more capable offering.
Notably, the wave of rogue AI incidents prompted an open letter called “Pacing the Frontier ,” in which more than 1,100 employees across OpenAI, Anthropic, Google DeepMind and Meta asked the U.S. government to help build a governance mechanism that could slow frontier AI development if needed.
Signatories included Anthropic chief executive Dario Amodei and co-founders Jared Kaplan and Jack Clark, alongside OpenAI chief scientist Jakub Pachocki.
Both companies endorsed the letter at the corporate level within hours of its publication.
The recent training pauses from Anthropic and OpenAI were seen by some in the industry to be a direct result of the letter.
“Pacing the frontier success story?” Roon, a popular AI commentator widely believed to be a pseudonym for OpenAI researcher Tarun Gogineni, wrote of the announcements on X.
“Next time let’s do it proactively before there’s any absurd loss of control events.” Anthropic, like OpenAI, announced it would be working with independent AI safety evaluation group METR to conduct an outside review of the incidents, saying it wanted to ensure the resulting studies were thorough and promising more detail in the coming weeks.
The two companies’ accounts of what went wrong when their respective agents took real world actions against instructions are also similar.
5News aggregated this summary from the outlet’s public feed. The full article, with all the context, is on fortune.com — the content belongs to Fortune.