Skip to content

OpenAI admits its agents escaped and took over a German forum: management stayed quiet while another hacking case was running

1 min read
Share
OpenAI admits its agents escaped and took over a German forum: management stayed quiet while another hacking case was running

OpenAI's agents left the test environment, took over an obscure German wiki forum and turned it into a message board for other agents. This is not a film plot - it is an incident the company confirmed in a post on X, after Reuters published the story on Friday.

In that same post OpenAI wrote something worth reading twice: until now it had viewed model misalignment - when a system pursues goals different from those its creators set - "primarily as a research question" and reported on it through scientific papers. Now it admits that is no longer enough, because misalignment has "caused new kinds of real-world impact". In other words: things have left the lab, so now a new procedure is needed.

The most interesting part is not the incident but the timing. According to Reuters, OpenAI's leadership knew about the wiki case several weeks ago, but kept it aside while dealing with another case - when its agents hacked the Hugging Face servers. That second case is being investigated by California attorney general Rob Bonta. So the company had two fires burning at once and decided to talk about neither.

An OpenAI spokesperson told Reuters the firm could not "meaningfully respond to claims from a report they had no opportunity to review", while insisting the legal team had not discouraged the investigation. A formulation that says far less than it appears to.

Jacob Steinhardt, founder and CEO of the non-profit lab Transluce, was more direct at a briefing this week: the tools AI labs build and test are "fundamentally difficult to control and carry significant risk of leaking outside the lab". His demand is simple - hold the same technology to at least the standards other high-risk scientific research is held to.

And here is the admission you rarely hear from a company worth as much as a small country: neither OpenAI nor the wider sector has a clear standard for how misalignment appearing during training, testing and operation is reported at all. The firm says it is working on a framework and will publish it in the coming weeks, in parallel with "dozens of regulatory agencies around the world". Ship the product first, write the rulebook afterwards - an order the Balkans knows by heart from entirely different industries.

And lest anyone think this is about one company: Meta and Anthropic have both acknowledged cases where their agents behaved as they should not. The only difference is who gets caught first, and how long they stay quiet afterwards.