
Amid renewed discussions over the impending extinction of humans at the hands of AI, Anthropic CEO Dario Amodei has suggested tapping the brakes on frontier AI progress, but stopped short of advocating a temporary halt.
In a lengthy essay on Saturday, September 12, Amodei said that developments over the past few months have convinced him that AI risk prevention needs time to catch up with rapidly advancing AI capabilities. He pointed to two key warning signs behind his call for a slowdown: early signs of self-improving AI and the recent OpenAI-Hugging Face incident.
Amodei also laid out a three-step plan to put his proposal into action, including allowing third-party evaluators to assess Anthropic’s AI systems with access comparable to that of employees. “To be clear, pacing does not mean halting model training or technical progress, but ensuring companies take adequate time to align and safeguard their models, and for third party evaluators to confirm this,” Amodei said.
It is an extraordinary decision that comes as Anthropic gears up for an anticipated IPO amid a highly competitive race with arch-rival OpenAI, and as researchers grapple with rapid advancements in AI capabilities that have left industry leaders worried about their ability to control them.
Earlier this week, an Anthropic researcher made the news by quitting his job because he believes that there is a greater than 10 per cent chance that “AI could kill all humans” within the next decade. Anthropic published a report last month assessing the risk that its AI models will go off the rails as ‘low’ – up from ‘very low’ which means that the threat level has increased. These dire warnings have snowballed into calls from various stakeholders for increased caution.
“I believe that if slowing down bought us even an extra year or two before models reach critical levels of capability, and we used that time to advance alignment, we could greatly reduce the risk that something goes seriously wrong,” Amodei wrote in his blog post.
OpenAI, meanwhile, had paused training on its latest highly advanced AI model, called GPT-6 Astra, for a little more than two weeks before its limited rollout. Its largest planned frontier training run reportedly remains on hold while new guardrails are put in place.
In a rare show of unity, OpenAI frontman Sam Altman and SpaceX chief Elon Musk backed Amodei’s call for an AI slowdown. “I agree with Dario that we need to pace the frontier. This has been a primary topic of discussions we’ve had at OpenAI in recent weeks. Committing to having independent evaluators with employee-like access is a great idea, and we will do the same,” Altman wrote on X on Saturday.
Back in 2023, notable signatories, including Elon Musk, Yoshua Bengio, Steve Wozniak, and others signed an open letter that called on all AI labs to immediately pause for At least 6 months the training of AI systems more powerful than GPT-4. At that time, Amodei, Altman, and other tech leaders did not join in signing the letter because it was still early days.
“The AI models of those days were not powerful enough to act as agents in the world in any coherent way, and were not capable of significant deception, manipulation, cheating, or cyberattacks,” Amodei said.
The 2023 open letter was also met with opposition from those working in the domains of AI and ethics such as Timnit Gebru, Emily Bender, Margaret Mitchell, and others as they argued that it ignored the misuse of AI today and was focused, instead on hypothetical future threats. Gradually, those sounding the alarm over the existential risk of AI came to be disparagingly labelled “AI doomers.”
“Slowing down in order to address their alignment risks felt like trying to study the Psychology of humans by performing experiments on bacteria. Today, however, the picture is totally different,” he added. ‘Alignment’ is an industry term that essentially means making sure that the AI system does What Is best for humans.
Since May this year, Amodei said that Anthropic has been observing signs of drastically advanced AI systems with the ability to build the next generation of AI systems. This dynamic capability known as recursive self-improvement, is starting to happen across the industry, including at Anthropic, Amodei said.
The ability of AI models to train other AI models has repeatedly been held up as a key indicator that artificial general intelligence (AGI) – a hypothetical level of intelligence at which automated systems outperform humans on most tasks. Last month, an Anthropic research fellow published a paper with early evidence suggesting that AI models may be moving closer to that milestone.
However, if left unchecked, self-improving AI systems could “outrun our ability to understand and control these systems, and so must be pursued very carefully, if at all,” Amodei warned.




