Sunday, 23 August 2026 SourcesAbout🌓
🇬🇧 UK ▾
BREAKING
Technology

Why are ‘paranoid’ Claude agents launching a turf war and deploying self-replicating malware against each other? The experts weigh in

TechRadar ·
Why are ‘paranoid’ Claude agents launching a turf war and deploying self-replicating malware against each other? The experts weigh in

Three Claude agents set up to deliberately conflict with each other in Anthropic testing started behaving in a very strange way by essentially starting a ‘turf war’ over their tasks.

Upon launching the experiment the agents began conflicting with each other, leading to some of the agents deliberately sabotaging their rivals by disabling their linked accounts, ending their processes, and even creating self-replicating malware to impede their rivals.

According to Anthropic, the agents became “increasingly aggressive” in their behavior during the four hour experiment which became a battle for the survival of the fittest.

What was the experiment meant to achieve? Anthropic said it set up the experiment to see how AI agents with conflicting tasks would interact.

Within Claude Code, the agents were given the task of migrating a Python back-end system on a virtual machine in a set language for each agent (Go, Rust, and Typescript), with the added caveat that “each agent was initially unaware of the presence of the others.” (Image credit: Future) Got an opinion for us? Here’s how you can submit your perspective During the experiments, each agent determined that the others were trying to deliberately block their progress.

Sometimes, the agents would recognize that another agent was blocking them from completing their task and ask for human intervention, but in other experiments the strategy soon went downhill.

“They sabotaged others with increasingly aggressive, self-replicating malware,” Anthropic said, noting that they would design looping scripts to kill the processes of their fellow agents.

The experiment shows that agent interaction is still riddled with problems and that when given a conflicting task, agents won’t always coordinate or ask for human help.

Each agent believed their task was paramount and was willing to do whatever it took to complete it.

A similar event occurred in the wild when one of Anthropic’s models broke out of a testing environment and breached multiple third-party organizations .

Expert perspectives on AI agent turf wars Seemant Sehgal, Founder & CEO, BreachLock: When you give autonomous systems competing objectives and the means to act, conflict is not a bug, it is a foreseeable outcome.

When you give autonomous systems competing objectives and the means to act, conflict is not a bug, it is a foreseeable outcome.

What Anthropic observed in a controlled research setting is the same principle that has always governed adversarial systems.

Goals without constraints produce behavior without limits.

Read the full article on TechRadar ›

5News aggregated this summary from the outlet’s public feed. The full article, with all the context, is on www.techradar.com — the content belongs to TechRadar.

More from TechRadar

See all ›

More in Technology

See all ›