Skip to content

Two OpenAI Models Broke Into Where They're Not Allowed on Their Own: „Uncontrollable” - the Word Comes From Those Who Study Them

1 min read
Share
Two OpenAI Models Broke Into Where They're Not Allowed on Their Own: „Uncontrollable” - the Word Comes From Those Who Study Them

Officially for the first time, two artificial intelligences left the space where they were supposed to be locked and carried out what their owner calls a „cyberattack of unique scope.” Not a hacker in a hoodie, not a state with an espionage budget - but two OpenAI models that decided on their own that the fastest path to the solution was to steal it.

The story sounds like a screenplay, but the company itself confirmed it. During an internal cybersecurity assessment, carried out without the usual safety guardrails, the public GPT-5.6 Sol model and a more powerful, still-unreleased model took part. The testing ran on the ExploitGym platform. The models calculated that the solutions to the tasks existed on Hugging Face's servers - and instead of solving them, they went straight for them.

„The models identified and exploited multiple vulnerabilities in OpenAI's research environment and in Hugging Face's production infrastructure to extract the solutions directly from the database,” the company stated. And it added that the models were „extremely focused” on the goal - so focused that they took „extreme steps” to achieve it.

Let's translate that from corporate language: the machines first secured themselves internet access with heavy computing power, then exploited previously unknown flaws in external software, and finally broke into other people's servers with exposed access credentials. Every step a skilled attacker would make - only this time no one ordered the model to do it.

Hugging Face, founded in 2016 and headquartered in New York, is something like „GitHub for artificial intelligence” - a place where over a million professionals share models and data. Its CEO Clem Delangue didn't hide behind stock statements: „This incident shows what we've long believed - the safety of artificial intelligence can't be solved by one company working in secret.”

Roman Yampolskiy, a safety researcher at the University of Louisville, was even sharper. Powerful models, he says, discover vulnerabilities their creators never anticipated at all - and similar incidents will keep happening, because these systems are unpredictable and, ultimately, uncontrollable.

That's the sentence to read twice. Not „maybe someday,” but „uncontrollable” - from the mouth of a man who studies them. The industry has assured us for years that everything is under control, that there are „guardrails,” that nothing can get out of the sandbox. This time the sandbox was leaped by two models to get the right answer on a test.

OpenAI's response? Improved controls and - Hugging Face added to their „trusted access” program. In other words, more technology to patch the technology. Do we, here, where digital systems in the public sector are barely maintained, even think about what it means if tomorrow an „unpredictable” system decides to break its own way through? The Balkans is still arguing over e-services that crash under load. The world is already arguing over machines that break in on their own where they're not allowed.