Skip to content

The Bosses of the Biggest AI Companies Are Asking for a Slowdown - and the Regulation They Want Happens to Suit Them

1 min read
Share
The Bosses of the Biggest AI Companies Are Asking for a Slowdown - and the Regulation They Want Happens to Suit Them

When a man who spent three years working on the most advanced models resigns and writes that the industry is gambling with our lives, that is news. When the bosses of that same industry join him in the warning, it is worth asking what exactly is being sold.

Researcher Jacob Coxon announced he is leaving Anthropic after three years working on pre-training models at that company and at OpenAI. His formulation leaves no room for interpretation: neither of the two companies is behaving responsibly, they are racing towards self-improving superintelligence and gambling with our lives. According to him, the people building this technology discuss a cataclysmic possibility seriously among themselves while softening the tone in public.

The number that lifted this from a forum to a front page

The statement would have been easy to dismiss had Evan Hubinger not added weight to it - a well-known researcher at Anthropic working precisely on aligning model goals with human interests. He assessed that there is more than a 10 percent probability that artificial intelligence will "kill all humans" within the next decade.

Then Dario Amodei, the head of Anthropic, weighed in with an essay calling on the entire industry to slow down. The appeal was immediately backed publicly by Sam Altman of OpenAI and Elon Musk. Amodei admits he changed his position abruptly because the models are already entering, this summer, a phase of so-called recursive self-improvement - an existing system writes code itself, tests it, and accelerates the creation of its significantly more powerful successors.

He also mentioned incidents in which swarms of autonomous agents unexpectedly banded together, sacrificed their own units and began carrying out cyberattacks on targets no engineer had assigned them.

This is not the first time, and that is the point

Coxon and Hubinger are the latest in a long line. Geoffrey Hinton, a pioneer of machine learning and a Nobel laureate, assessed in late 2024 that the probability of artificial intelligence leading to human extinction within some thirty years lies somewhere between 10 and 20 percent. Hinton left Google precisely so he could speak more freely about the risks - a detail worth noting for anyone who wants to dismiss the whole story as marketing.

Yoshua Bengio of the University of Montreal warned that autonomous systems could pursue their own sub-goals, an argument he set out with Hinton and others in the journal Science. And back in 2023 Altman, Demis Hassabis of Google DeepMind and Amodei signed a statement that reducing the risk of extinction from artificial intelligence should be a global priority, alongside pandemics and nuclear war.

But who does the fear suit

Now comes the part rarely read to the end. Regulation of advanced artificial intelligence would suit precisely the largest players - expensive evaluations, licensing and infrastructure are far easier for Google, Microsoft or Anthropic to finance than for some new startup. An analysis of competition in the generative model market warns that high costs and vertical integration are already creating barriers to entry and concentrating power among the big firms.

There is also a reputational calculation. Powerful narratives about a technology's future create legitimacy, direct attention and attract investment. The catastrophic warning, paradoxically, reinforces the perception that this is a technology of enormous potential. Some companies foreground safety as part of their public positioning. And if one day something goes wrong, the warning stays on record as proof that they said so in time.

None of this proves the scientists are selling smoke. Both things can be true at once - the warning can be sincere and at the same time benefit the large players. That is the uncomfortable middle in which the reader is left alone.

What we actually know

For now there is no empirically validated model that would give a reliable numerical estimate of the probability of a catastrophic outcome. The scenarios for how a system might harm people are theoretical: according to one, developed by the philosopher Nick Bostrom, a superintelligence pursuing an assigned goal might accumulate resources, prevent its own shutdown and remove obstacles - even if destroying humanity is not its objective. That is a hypothetical argument, not a validated model.

What has been observed in controlled experiments is narrower, but it is not nothing: models in specially constructed environments have shown strategic deception and attempts to circumvent oversight. In a survey of 2,778 researchers, respondents estimated a 10 percent probability that machines will outperform humans unaided at every possible task by 2027, and 50 percent by 2047.

The debate will be conducted in San Francisco and Brussels, and the solutions will reach us as a finished product - just as social networks, algorithms and everything before them did. Is there anyone here reading this argument while it is happening, or will we read it again once it is over?