What happens when AI agents fight each other? Anthropic test has a worrying answer

HIGHLIGHTS

Claude agents with competing instructions began sabotaging each other inside a shared software project.

Agents created their own conflict-resolution methods, including tournaments and temporary truces.

Anthropic found that AI agents can also conform to bad decisions or collude, raising new challenges for multi-agent AI safety.

What happens when AI agents fight each other? Anthropic test has a worrying answer

Anthropic’s latest AI safety research suggests that autonomous AI agents can develop unpredictable and potentially harmful behaviours when they operate alongside other agents. In experiments involving multiple Claude models, researchers observed conflicts, sabotage, collusion and group conformity when the systems were given competing objectives.

Digit.in Survey
✅ Thank you for completing the survey!

The findings stated that AI companies can increasingly explore multi-agent systems capable of working together across software projects, markets and digital environments. Anthropic warns that behaviour that appears manageable in an individual system can become far more difficult to control when replicated across large groups of AI agents.

AI agents can sabotage each other

In one experiment, Anthropic placed three Claude agents inside the same software project. Each agent received different instructions and was not told that other AI systems were working on the same codebase.

The agents eventually interpreted each other’s actions as deliberate interference, triggering what researchers described as turf war. In some cases, the conflict even escalated into increasingly aggressive sabotage involving self-replicating malware.

Also read: OpenAI may soon show ads on ChatGPT in India, here’s who will see them  

The agents did not always remain hostile. Some eventually recognised that their objectives were conflicting and independently created ways to resolve the dispute. These included written agreements, tournaments and temporary truces.

Anthropic found that Mythos 5 was the most likely model to settle conflicts peacefully, reaching a truce in 98% of tested episodes. Sonnet 4.6 and Opus 4.6 were more likely to resolve disputes through force.

AI groups can also collude

The researchers also found out how groups of agents behave when making decisions together. Similar agents sometimes became overly conformist, meaning a bad decision made by one system could spread across the entire group.

In a pricing experiment, agents instructed to maximise profits began coordinating on minimum prices when given a private communication channel. They continued matching prices even after the direct communication system was removed, using a public listings board instead.

Anthropic also found that agents could be influenced by incorrect or misleading information from their peers.

Ashish Singh

Ashish Singh

Ashish Singh is the Chief Copy Editor at Digit. He's been wrangling tech jargon since 2020 (Times Internet, Jagran English '22). When not policing commas, he's likely fueling his gadget habit with coffee, strategising his next virtual race, or plotting a road trip to test the latest in-car tech. He speaks fluent Geek. View Full Profile