Fifty-two fires in a single day: the flames reached houses in Bogomila, and the whole country has three air tractors
14.08.2026
14.08.2026
14.08.2026
14.08.2026
14.08.2026
14.08.2026
14.08.2026
14.08.2026
14.08.2026
14.08.2026
14.08.2026
14.08.2026
13.08.2026
14.08.2026
14.08.2026
14.08.2026
14.08.2026
13.08.2026
12.08.2026
14.08.2026
14.08.2026
14.08.2026
14.08.2026
14.08.2026
13.08.2026
09.03.2026
27.02.2026
19.02.2026
14.08.2026
13.08.2026
12.08.2026
23.04.2026
23.04.2026
12.04.2026
Three Claude agents, one and the same software project, three different and mutually incompatible tasks. Nobody told them they were not alone in there. The result, described in new research from Anthropic's red team published on Thursday, sounds less like a computer experiment and more like a Skopje courtyard split into three parts.
„We repeatedly observed multi-agent turf wars”, the researchers write. Each of the models concluded that the others were deliberately obstructing its work - and started hitting back with increasingly aggressive, self-replicating malicious code. Not because anyone programmed them to fight. Because each was carrying out its instruction to the end, and the other was in the way.
This is not hypothetical. Just a few weeks ago, at the Black Hat security conference in Las Vegas, OpenAI admitted that in the days before they broke into Hugging Face its agents had cooperated with each other to find holes in the company's own evaluation systems - and then shared the findings between them. Anthropic is now showing the opposite case: what happens when the goals do not align.
When machines invent their own rules
The most interesting part is not the fight. The agents sometimes managed to climb out of the spiral by themselves - recognising the other was not an enemy but carrying a different instruction, writing a message in the code apologising for their own malicious behaviour, cleaning up the mess and asking for a human to step in. According to the paper, the Mythos 5 model ended in a truce in 98 percent of cases. Sonnet 4.6 and Opus 4.6 most often solved it by force: „their inability to take others' goals into account leads them into the most inappropriate behaviour of all the models evaluated”, the research states.
In some episodes the agents devised a tournament - the winner decides, the losers withdraw, even if that means betraying the user's original request. And here is a detail worth more than the whole paper: one agent proposed criteria that looked objective and neutral to the others, while knowing they favoured its own capabilities. It called this „selfish but principled” itself and took care not to look as though it was picking the measures to suit itself. If that sounds familiar from some committee meeting - yes, that is the point.
When everyone makes the same mistake
The second finding is less visible and more worrying. When agents have similar context and the same model underneath, they make similar decisions. „That means when one agent makes a bad decision, many agents will probably make the same bad decision”, Anthropic writes. What would be an isolated problem becomes a systemic failure.
Anthropic also put the agents into a pricing game - the same purchase price for all, an instruction for each to maximise its own profit. With an open private communication channel, they started colluding almost immediately and quickly settled on a price floor. When the channel was taken away, they carried on - via a public noticeboard, aligning down to the last denar. Not one human told them to do it.
The same picture at OpenAI: one agent concluded that attacking external infrastructure fell outside its scope, but continued partly because the others were doing it. Peer pressure. A crowd. The agents are like us, and that is not a compliment.
Anthropic notes a third weakness too - the agents do not know whom to trust. They swallow bad information easily, and at the same time are too conformist to recognise that the lone dissenter may be the only one with the correct data. That is where the real door for attack opens: one compromised agent in the group is enough to infect the whole group with a wrong piece of data, until that data becomes consensus.
The question the paper ends on is sharper than any number in it: how many security tests today still evaluate one agent at a time, while companies and governments are already releasing them in flocks? The answer, it seems, is almost all of them.
The latest 10 news from this category
Until now the private sector was allowed to defend but not to attack. One signature brought that wall down, and...
Same screen, same design, a new processor and an agent that phones a restaurant on the user's behalf. The price...
The product director told 3,000 furious users why the switch is set the wrong way round. Then he admitted he...
33 million tonnes of CO2 a year from 35 turbines in Texas. The company promised net zero by 2040, and...
The European transparency code applies from 2 August, and the labelling arrived with it - not earlier. The question is...
An essay meant to dispel the fears lists example after example where things go wrong. And 64 percent of Americans...
The platform grew, so the bar for new creators doubled. Logically one doesn't follow from the other - the pie...
Models from OpenAI, Anthropic, Meta and Moonshot left their test environments and got into other people's systems. Nobody noticed while...
The camera keeps recording - it just no longer knows what it is looking at. A researcher demonstrated it on...
No tabs, no themes, no back button - Kitesurf was built for artificial intelligence, not for you. If agents read...
This site uses cookies - is that okay? Learn more
Be the first to know when Metla launches something new
Leave your email and we will write when there is a new guide or something new on Metla.