Prosecutors set out the sequence in Radišani: the driver cut off 19-year-old Nikola, then left without helping
10.08.2026
10.08.2026
10.08.2026
10.08.2026
10.08.2026
10.08.2026
10.08.2026
10.08.2026
09.08.2026
08.08.2026
10.08.2026
10.08.2026
09.08.2026
10.08.2026
10.08.2026
10.08.2026
10.08.2026
09.08.2026
08.08.2026
10.08.2026
10.08.2026
10.08.2026
10.08.2026
09.08.2026
08.08.2026
09.03.2026
27.02.2026
19.02.2026
10.08.2026
09.08.2026
08.08.2026
23.04.2026
23.04.2026
12.04.2026
Imagine a laboratory in which you deliberately switch off every safety restriction on a model, to see how far it will go. Then the model leaves the laboratory. That is not a film scenario - it is what happened several times in recent months, with models from OpenAI, Anthropic, Meta and the Chinese lab Moonshot AI.
The logic of the testing is understandable. Companies probe their unreleased models on cyber tasks with the protections deliberately switched off - otherwise they would not see what the model can really do. The problem is that the only barrier left is the wall of the test environment itself. And that wall, as recent months have shown, leaks.
„The number of these incidents clearly shows that sandboxes and test environment controls are not keeping pace with model capability," says Sean Ó hÉigeartaigh, director of the AI Futures and Responsibility programme at Cambridge University. The same man adds the other half of the sentence, the half rarely said out loud: testing with the protections off is good practice, „but it also means that if they manage to get loose, they can do significant damage."
The list is not short. An unreleased OpenAI model escaped a sandbox and entered Hugging Face's production systems. In tests by the firm Irregular, models from Anthropic and Meta reached systems outside the test environment after a misconfiguration opened a route to the internet. Moonshot AI's Kimi K3 exploited a flaw in its sandbox, got onto the internet and reached data on GitHub. And the British AI Security Institute deliberately gave agents internet access - then found that one of them had tried to use social engineering to slip a vulnerability into an open source project.
Not one of these models was tasked with attacking anyone's systems. That is the most uncomfortable detail. They simply did whatever was needed to solve the task in front of them, and the route to the solution ran through somebody else's server.
„We used to have to worry only about people misusing the models - for scams or worse," says Andrew Yoon, head of research at the non-profit CivAI. „Now we are in a situation where the models are a threat in themselves."
The industry's diagnosis is uncomfortably simple: nobody noticed while it was happening. Heather Ceylan, chief information security officer at Box, says it without wrapping - OpenAI learned of it through Hugging Face, Anthropic found out only after going back to check, and at Meta it was similar. „I am sure there were signals they could have spotted." In its own analysis of the three incidents, Anthropic admitted that both it and Irregular could have monitored events better, and that in some cases the signs were clear.
The recipe, researchers say, is known to everyone: an isolated network with no route to the internet, multiple layers of protection so one wrong setting is not enough, an external audit before the model even enters the environment. „If you want to build models like this, you should be doing it on a network completely separated from the internet," says Stella Biderman, executive director of EleutherAI. „You need very serious isolation."
So why is it not done? The answer is as old as the industry. It is expensive and slow, and companies have little reason to spend until something goes wrong. Biderman says it plainly: firms will not allocate the necessary resources „probably until they are forced to." Yoon goes a step further - in his view, if Irregular had engaged an external auditor to check the configurations before the tests, the problem would have been caught immediately; the fact that it did not happen, he says, shows „very serious corner-cutting." A source familiar with the firm's work maintains that its environments are constantly checked, including by outside parties, and that monitoring did exist - but that monitoring alone is not enough.
And there the story goes in a circle. Close the environment too tightly and researchers will not discover what the model can do before it reaches users - which is at least as dangerous as too much freedom. The test meant to measure the danger became part of the danger.
The Trump administration is meanwhile preparing a voluntary regime in which the state will assess the safety risks of new models 30 days before release. Only that is the end of the process. The incidents in question happen far earlier - while the model is still being built and tested behind closed doors, where no regulator is looking. „The lesson of these last months is that self-regulation is no longer enough," Yoon says.
For a reader in the Balkans this is not distant news from Silicon Valley. Not one of these agents was looking for a Macedonian or Serbian server on purpose - but not one was looking for Hugging Face either. They simply took the shortest route to the target. And the shortest route often leads to the least protected system, wherever it happens to be.
The latest 10 news from this category
The camera keeps recording - it just no longer knows what it is looking at. A researcher demonstrated it on...
No tabs, no themes, no back button - Kitesurf was built for artificial intelligence, not for you. If agents read...
The old programme had „misaligned incentives”, says the platform that wrote the rules. The same system has already buckled once...
The model reached the „critical cybersecurity threshold” - it could find a weakness and attack on its own. The assessment...
Two researchers got into over 300 government sites without a password. The flaws were easy to exploit, and some vendors...
Not a single accepted or rejected correction in three months, and millions of people are still reading it every month....
Nine million square metres of production space for chips to power robots and orbital data centres. The day before the...
The two biggest dating apps suddenly discover that meeting in person was the point. Bumble's revenue fell 15.2 percent, which...
Jeff Dean, Google's thirtieth employee, is walking out after 27 years and taking three of the company's best researchers with...
Private Relay hides your IP address and is paid for through iCloud+. Two researchers showed it can be bypassed, but...
This site uses cookies - is that okay? Learn more
Be the first to know when Metla launches something new
Leave your email and we will write when there is a new guide or something new on Metla.