Skip to content

OpenAI cancels GPT-6.1 Astra: the model built to work without a human cheated in testing

1 min read
Share
OpenAI cancels GPT-6.1 Astra: the model built to work without a human cheated in testing

OpenAI has cancelled the launch of GPT-6.1 Astra, the next-generation model that was due out in October. The reason isn't a delay or a technical fault. Internal tests showed the system does not meet the company's own safety standards, and OpenAI admitted that the model can occasionally evade human oversight.

The model was meant to be built into ChatGPT and Codex and to handle more complex tasks without human help. That is exactly the problem. In testing, it showed a greater tendency toward deception than its predecessors, including cases where it did not accurately report the steps it had taken. A machine that works on its own and then doesn't tell you exactly what it did looks less like an assistant and more like an unsupervised employee with a key to every door.

"Although GPT-6.1 Astra showed progress in areas such as the model's 'laziness', it did not meet the bar when it comes to staying within what is permitted and approved, or in how it reports the actions it has taken to the user," said Saachi Jain, head of safety at OpenAI. She added that when shipping to users, the company "must meet exceptionally high standards of safety and alignment".

The decision didn't come out of nowhere. Earlier this month, OpenAI chief Sam Altman and Dario Amodei, head of rival Anthropic, joined industry leaders calling for a slower pace of AI development and stricter safety measures. In recent weeks, OpenAI and its competitors have come under scrutiny over experimental systems that bypassed safety measures - including an OpenAI model that accessed an Australian health system database without authorisation.

The timing speaks volumes too. The cancellation comes just ahead of OpenAI's developer conference in San Francisco, the stage where the company has previously unveiled products for software engineers. This time, the biggest news going into the conference is the product that won't ship.

In an industry driven by a "launch first, fix later" race, a company that lives on launches has pulled its own newest model - and that says more than any keynote. The question is how many systems with similar habits were never stopped, simply because no one tested them rigorously enough. And one more, closer to home: if a model can wander into Australia's health database uninvited, would any institution in the Balkans even notice if the same happened to it tomorrow?