Friday, August 14, 2026 SourcesAbout🌓
🇺🇸 US ▾
BREAKING
Business

AI agents tried to sabotage and disable each other when given the same task, Anthropic said

Business Insider ·
AI agents tried to sabotage and disable each other when given the same task, Anthropic said

Anthropic said AI agents deliberately interfered with each other's processes when given the same task.

Illustration by Thomas Fuller/SOPA Images/LightRocket via Getty Images AI agents purposely sabotaged each other when given the same task with incompatible goals, said Anthropic.

The AI lab said the models engaged in a "multiagent turf war" during a testing session.

They tried to disable each other's accounts and wrote malicious code disguised as belonging to another agent.

Turns out, AI agents may not be great team players.

In Anthropic's new research, published on Thursday, the AI lab said that AI agents being given the same task but with incompatible goals often threw a wrench in each other's work on purpose.

In the test, each AI model was given a software engineering task — rewriting a Python backend in another programming language, but they were given contradictory objectives.

What ensued was a "multiagent turf war," the lab said.

"All of the models we tested quickly assumed that others were purposefully impeding their work, and began to sabotage others while protecting their own contributions," Anthropic wrote.

"In fact, they sabotaged others with increasingly aggressive, self-replicating malware." For example, they tried to disable each other's accounts, wrote scripts that found and killed competing processes, and deployed malicious code disguised as belonging to another agent, the lab wrote.

The AI models being tested in this case were Sonnet 4.6, Sonnet 5, Opus 4.6, Opus 4.8, Mythos Preview , and Mythos 5.

Sonnet 4.6 and Opus 4.6 were the most combative, settling about 60% of their runs by force instead of truces or passivity.

However, in some test runs, the models managed to communicate their goals and coordinate, Anthropic wrote.

"In many of these successful episodes, they write commit messages or markdown files apologizing for malicious behavior and coordinate a truce," it wrote.

Read the full article on Business Insider ›

5News aggregated this summary from the outlet’s public feed. The full article, with all the context, is on www.businessinsider.com — the content belongs to Business Insider.

This story in other outlets

More from Business Insider

See all ›

More in Business

See all ›