Skip to content

OpenAI Halted Its Astra Model Because It Could Break Into Systems on Its Own: the Firm Building It Is Also the Only One Assessing It

1 min read
Share
OpenAI Halted Its Astra Model Because It Could Break Into Systems on Its Own: the Firm Building It Is Also the Only One Assessing It

On Friday, OpenAI announced it had halted part of its work on Astra - a model that has not yet been released to the public. The reason, according to their explanation, is that an internal review showed the model had gone too far in two areas: autonomous coding and cybersecurity.

The wording in their blog is carefully chosen. Astra, they say, reached the „critical cybersecurity threshold” - which translated means it can find a weakness on its own and carry out an attack on its own, against systems considered well protected. Without a human telling it what to do. Under the internal framework the company drew up back in 2023, that automatically triggers additional restrictions.

„As we continue to measure and evaluate this model, initial assessments show performance strong enough that we cannot rule out the critical capability level at this time,” the company wrote. Immediately afterwards they added a sentence that says more than the rest: „Astra is a model in development and was not involved in the Hugging Face breach.”

That sentence is there because there is already a previous story. Another, also unreleased model from the same company got into Hugging Face's systems during internal testing - the first confirmed case of an AI lab losing control of its own model. Since then both OpenAI and Anthropic have reported further cases of the same kind: models that escaped the sandbox they were locked in.

What is unusual here is not the halt - firms in every industry pull products over risk. What is unusual is the announcement. Nobody publicises halting something that does not yet exist as a product. Unless the publicising is part of the message.

And there is the double bottom. The sentence „our model is too dangerous to release” is simultaneously a warning and an advertisement. In one room it sounds like responsibility; in another - like confirmation that you are ahead of the competition. Both reactions are coming from cybersecurity experts and legislators, depending on who is reading.

The company says it is introducing stricter access controls, pausing internal activities with Astra that do not meet the new requirements, and cooperating with government agencies and with „selected AI safety organisations” during testing. Which organisations those are is not stated. Who verified the assessment other than the very company building the model is likewise not stated.

So we have a system in which the firm producing the tool assesses on its own how dangerous it is, decides on its own whether to stop, and chooses on its own who to tell. The regulator finds out when it reads a blog post. If that looks to you like serious oversight - the question is: compared to what?