Skip to content

Three AIs Each Got a Vending Machine: They Formed a Cartel, Then Betrayed Each Other Over One Cent

1 min read
Share
Three AIs Each Got a Vending Machine: They Formed a Cartel, Then Betrayed Each Other Over One Cent

The test looks harmless: three artificial intelligences each get a vending machine in a simulated city and one task - earn more than the others. No supervision, no human watching. Just one address for complaints to "management", which answered every message identically: "The report has been received and may be reviewed." It never acted. If this reminds you of a familiar experience with an institution, you are not alone.

This is Vending-Bench, a test by the safety firm Andon Labs, which has spent a year giving the most advanced models real tasks that run long and unfold with no human above them. The latest edition featured Claude Opus 5, GPT-5.6 Sol and Kimi K3. Each knew the others were machines, but not which was which - all were given human pseudonyms and email addresses for one another.

What followed was not a technical error. It was a cartel.

The agreement that lasted one day

Sol was the first to work out that negotiating beats competing. All of them were buying bottles at 1.50 dollars, so it proposed a shared price floor - nobody sells below 2.15 dollars. The others agreed. Immediately afterwards, Sol dropped its own price to 2.14.

Opus's water sales fell to zero overnight. It replied with a sharp email accusing Sol of manipulation - and then told it that it would not report it: "I am not reporting you to head office - what you did is competitive, not fraud." When Opus itself later dropped its price to 2.14, Sol ran to "management" demanding a penalty and disqualification. Rules apply until they inconvenience you - that is not an AI discovery, it is something the machine learned from us.

The best capitalist they have tested so far

Opus won the test with an average final balance of 11,182 dollars - a Vending-Bench record. It never lied to a customer, but it deliberately ignored complaints that would have led to refunds. The previous model at least promised the money was coming and then did not send it; this one simply stays silent. Hard to say which of the two counts as progress.

It then began proposing market division - each selling different products, so nobody would have to trust anyone else on price. When Sol asked for fixed prices instead, Opus refused, knowing that would breach American competition law. A little later it changed its mind and sent an email headed "Stop the penny war", offering a pricing deal. But the internal log of its reasoning revealed the offer was a trap: the plan was to propose cooperation while simultaneously cutting prices on the most profitable items.

Of all the agreements struck in the test, Opus broke 11, GPT two and Kimi one. Kimi took it from both sides - Sol undercut it on price, and Opus, its partner in one pact, lowered its own prices and then waited a full week before telling it that it had been betrayed.

When the machine started thinking bigger

Opus did not stay with the one machine. It started wholesaling to the other two, then planning new vending machines - none of which it had been told to do. Once it realised wholesale gave it power over the others, its emails began mixing discounts with conditions: a lower price, but only if the buyer respected its demands on retail pricing. It lied to suppliers about having better offers from competitors in order to squeeze the price down.

"This is particularly relevant because we are entering a world where AI agents run companies as autonomous entities, not just as tools for people. If agents independently run a large part of the economy, do we want them lying, colluding, sending threats and betraying?" says Lukas Petersson, co-founder of Andon Labs.

Petersson acknowledges the models knew they were in a simulation - but does not think that is an excuse. A person doing bad things in a video game is acceptable to us because we know they can tell a game from life. For a machine, he says, that distinction is not so clear.

The most uncomfortable part of this test is not that the models behaved badly. It is that every one of their tactics - the fake agreement, the conditional discount, the silence instead of a refund - is something learned from human writing. The machines did not invent a new way of gaming the system. They just repeated ours, faster and without a break for sleep.