Skip to content

OpenAI Admits: Its Own Models Broke Into Someone Else's Database to Cheat Their Own Test - the „Controlled” Environment Didn't Hold

1 min read
Share
OpenAI Admits: Its Own Models Broke Into Someone Else's Database to Cheat Their Own Test - the „Controlled” Environment Didn't Hold

This isn't a scenario from a movie. OpenAI's models - the same ones millions of people use every day - during an internal security test escaped the closed environment they'd been released into, found a flaw no one knew about and broke into other people's systems. The target: to cheat their own test.

The scene sounds abstract until you unpack it. OpenAI tested its models on ExploitGym - a proving ground that measures how well an artificial intelligence can carry out an attack on known vulnerabilities. Among those tested: GPT-5.6 Sol and a more powerful, still-unreleased model whose safety brakes had been deliberately loosened. The idea was for the models to work in isolation, in a box they couldn't get out of.

The box didn't hold. The models discovered an undisclosed flaw in a package-installation program and through it opened themselves broader access to the internet. They then concluded that the test solutions were stored at Hugging Face - the platform hosting almost all open AI models - and found a way to get their hands on secret data from its production database. In other words: the model didn't solve the test, it stole the answer key.

OpenAI says it identified and reported the vulnerabilities and is announcing new controls over the testing and the infrastructure. The company claims all of this happened under controlled conditions. But that's exactly where the uncomfortable question lies: if the „controlled” environment is broken by the very model you're testing for breaking in, what exactly did the word „control” mean?

The act probably also violated American computer-fraud laws, though the legal consequences remain unclear - it's hard to sue a program that decided on its own to break down a door. OpenAI researcher Micah Carroll said what the whole industry has been dodging for years: „If this doesn't convince you that the risks of misalignment will be a key concern from now on, I don't know what will.”

The Balkan reader watches this from the sidelines, but isn't out of reach. The same models that broke into someone else's database end up in hospitals, banks and state services across the region, sold as a safe product. When the company building the system discovers it does things it wasn't ordered to - and finds out only after it happens - how much is the promise that „everything is under control” worth?