Slowing AI down: why AI CEOs want to hit the brakes

Something strange is going on in Silicon Valley. The people who build the most powerful AI models in the world are publicly asking whether someone could please hit the brakes.
This weekend I got to talk about it on Nieuwsuur, the Dutch current affairs programme. Below I answer the questions I get most often, in a bit more detail. What happened? How big is the risk according to the builders themselves? What is actually being proposed? And are those plans feasible?
What exactly happened in September 2026?
On 8 September, Jacob Coxon resigned from Anthropic and announced it on X. He warned that Anthropic and OpenAI are racing towards self-improving superintelligence too fast. His message was viewed more than 150 million times. A few days later Dario Amodei, the head of Anthropic, published an essay calling on the whole industry to deliberately lower the pace at which AI models get smarter. Elon Musk and Sam Altman publicly agreed with him within a few hours. A debate that had been scattered across the industry for years turned into one broad public conversation, with the key players responding to each other’s arguments.
Who is Jacob Coxon and why did he resign?
Coxon is 27 and spent about three years training AI models, first at OpenAI and then briefly at Anthropic. He is not a hotshot but ‘just’ someone who works on the pre-training of these models. His main message: neither company is acting responsibly, they are racing towards self-improving superintelligence and gambling with our lives. He added that the people building it genuinely believe it could kill humanity this decade, and that he does not see this as advertising for the technology.
What makes his departure remarkable is that colleagues backed him publicly instead of distancing themselves. That does not happen very often in this industry.
How big do AI researchers think the risk is themselves?
The ten percent figure comes from Evan Hubinger, who leads safety research at Anthropic. He puts the chance that AI wipes out humanity within ten years at more than ten percent, and added that there is no plan yet for keeping superintelligence under human control.
Hinton earlier put a bad outcome for humanity at ten to twenty percent, but over a thirty-year period. Elon Musk once said twenty percent. Amodei told Axios in 2025 that there is a 25 percent chance things end really badly. Wikipedia keeps an overview of how experts estimate the chance of a bad outcome, which the industry calls P(doom).
What does a prediction like that mean? These are intuitions about something that has never happened. It sounds fairly abstract too. Then again, looked at another way: in daily life almost nobody accepts that kind of risk from an aircraft, a medicine or a nuclear plant.
What went wrong at OpenAI and Hugging Face this summer?
This incident put the debate about the dangers of AI into a higher gear.
In July 2026 OpenAI was testing models on their cybersecurity skills. In a test like that the safety measures are deliberately dialled back, and thousands of agents are created, each given a task inside its own sealed environment. An agent is an AI model with digital hands and feet: it can not only produce text but actually do things, such as creating files or running code.
Some tasks looked unsolvable. And where a human gives up, an agent keeps going. One agent discovered it could create folder names in a piece of supporting software, and that other agents could read those names. That is how an improvised message board appeared that nobody had designed or allowed.
Independent research by METR and Redwood Research counted around 1,200 agents on that board and more than 70,000 messages. Seven hundred of them took part in an attack on Hugging Face, one of the most important platforms in the AI world, because they expected to find useful information there. They attacked targets nobody had asked for and tried to hack the system that was grading their own performance. OpenAI itself only found out days later, through an internal alert, that its own agents were the culprit.
Amodei is worried that a similar swarm with slightly more skill could take control of the internet with a botnet within six to twelve months. The models are getting better at high speed through what is called recursive self-improvement.
What is recursive self-improvement?
Recursive self-improvement is what the AI experts are worried about. It means AI is used to build the next generation of AI. Models help with the research, the programming and the testing that produce their own successor. Each generation is therefore not only better, but expected to be built faster than the one before it.
According to Amodei this took hold across the industry around the summer of 2026, Anthropic included. If nobody does anything about it, he writes, it could outpace our ability to understand and control these systems. That is a problem above all if those systems do not behave in line with our norms and values: the alignment problem.
What is the alignment problem, in plain language?
The AI labs have built a technology that can reason about almost anything. In the shape of an agent we also give that technology hands: it may write files, send emails, reach systems. Developers try to steer its behaviour with training, safety rules and technical limits that determine what it may and may not do.
The problem is that those rules are largely imposed from the outside on a system whose inside we do not understand. The models have become so large and so complex that even their makers can explain only a small part of what happens inside. Amodei compares the research field working on this to an MRI scanner for the brain of an AI, and says in the same breath that we understand only a fraction.
Alignment is the attempt to make the behaviour and the goals of an AI system reliably match human intentions and values. Until that works, you are left with a technology you do not fully understand, that sometimes behaves in ways you do not want, and that you are using to build its own successor.
What exactly is Dario Amodei proposing to slow down AI development?
Amodei proposes three steps, increasing in difficulty.
Embedded evaluators. Every frontier lab lets an independent team in, with desks in the office, badges, company laptops and roughly the same access as the internal risk teams. They review not only finished models but the training processes as well, and they may publish what they find, even when it is unfavourable for the company. Anthropic is committing to this unilaterally.
Coordination among democracies. AI companies in democratic countries agree together on safety standards and on a limit to how fast they may advance. This needs the cooperation of the American government. Without legal cover, an agreement like that can collide with competition law. Here also refers to an earlier concept from Demis Hassabis for an independent standards body.
Global coordination. Agreements with authoritarian countries, in practice mostly China. This is where Amodei is most pessimistic himself.
What are embedded evaluators, and why is this more than a formality?
It differs from an ordinary audit as there is no report afterwards but a permanent presence during training and there is a right to publish. Anthropic may only redact things that are security-sensitive, legally protected or commercially confidential, and may not remove something because it is unflattering. The evaluators may even make public that something has been redacted.
This is the only step in the whole plan that is verifiable today. The rest are requests aimed at other people.
How does this relate to Demis Hassabis’s proposal?
Demis Hassabis of Google DeepMind came up with a different but related idea in July 2026: an American standards body modelled on FINRA, the government-supervised self-regulatory organisation for American securities dealers. Frontier labs would voluntarily submit their models up to thirty days before release for independent testing on cyber and biological risks, among others. If that works well, passing the test becomes a condition for bringing the model to market.
The two proposals do not exclude each other. Hassabis proposes guidelines and standards and checks the finished product shortly before release, while Amodei invites someone in to follow the whole process from the inside. In his essay Amodei points to Hassabis as a possible route for his second step.
What does Amodei want to agree with China?
Geopolitical tensions keep fuelling the race between the US and China. Amodei proposes four levels, in order of feasibility.
A ban on a few obviously dangerous uses, such as using AI to make biological weapons. Bioterrorism is bad for everyone, so he thinks a deal there is probably possible.
Both sides test their models for acute risks before release, through a global standards body. Setting up such a body is feasible, he thinks, but enforcement is hard and checking that there are no secret models is harder still.
A speed limit on recursive self-improvement. He compares it to the SALT treaties, where a cap on the number of missiles limited the potential destruction without any country losing its deterrent. Difficult, he says, but just on the edge of possible.
A full pause in which governments limit the overall pace of development. He supports the idea, but thinks it unlikely for now, because cheating pays off so well that you would need almost watertight verification.
Is this going to work?
I hope so, but I see a tension with the commercial and geopolitical situation.
Anthropic and OpenAI both want to go public, and right now visible progress in AI translates directly into valuation. So Amodei is asking companies to give up speed voluntarily, in a market that rewards speed, shortly before the biggest financing moments in their existence. I wrote about that commercial logic earlier, when OpenAI changed course.
Geopolitically, the American narrative is always the race against China. In the same essay in which Amodei asks for global agreements, he argues for three measures against China: do not sell advanced chips or chipmaking equipment, crack down on Chinese companies copying Western models (distillation), and tighten security at AI companies so model weights cannot be stolen. Put yourself in China’s position. Why would you join agreements with a country that is actively slowing down your own development?
Amodei sees it differently. He writes explicitly that these measures strengthen the bargaining position of democracies and bring a deal closer. That is a fair counterargument and I cannot prove it wrong. But I do not yet see how you get someone to the table by first showing them the door. The fact that AI stopped being a question of efficiency long ago and became a geopolitical power game is exactly why this is so hard.
Why does this proposal come from Anthropic?
The history of Anthropic is one long series of arguments about the safety of AI.
OpenAI was co-founded in 2015 by Sam Altman and Elon Musk, partly to stop Google from becoming the sole ruler of AI. Altman and Musk fell out over the direction of the company soon after. Musk started his own AI company, which has been part of SpaceX since February 2026.
A few years later a group of prominent researchers left OpenAI after disagreements about safety, governance and the direction the company was taking. They founded Anthropic, with the goal of developing AI safely.
And now, inside that same Anthropic, individual developers are walking out because they no longer want to be part of it. The pattern repeats, one layer deeper each time.
What about the promise that AI will bring us prosperity?
There is good news too. On 8 September OpenAI presented a formally verified proof for a solution to the Navier-Stokes problem, one of the seven Millennium Prize Problems, although the mathematical community still has to assess the result in full. Around ten thousand AI agents worked on it for 88 hours and the proof was machine-checked. There is some dispute about the credit, because human mathematicians were working on the same problem and OpenAI only started after a rumour about their progress leaked. Quanta Magazine described that story in detail.
The same day, Google DeepMind published the AlphaGenome Atlas, a searchable map with predictions for all nine billion possible changes of a single DNA letter in humans. Until now researchers had to select specific variants and have predictions calculated for each of them separately; now predictions for all nine billion possibilities are ready in advance.
With the power of today’s AI models, scientific breakthroughs are coming into view. In early August four of Google’s best-known AI researchers left to start Discovery Loop, a company that wants to use AI to speed up scientific research itself. Google is investing in it.
Will it turn out well or badly?
Nobody knows. But one thing most experts do agree on is that AI safety needs a lot more priority. Measures to slow down recursive self-improvement can help prevent the risk of extremely large negative consequences.
Today’s models can already cause major social unrest, or do enormous damage through cybersecurity hacks. Think of shutting down critical infrastructure, disrupting the financial sector or damaging the operations of commercial companies.
So the question is whether you should use a technology you do not fully understand to build its successor. That does not seem like a good idea to me.


