In short
A senior Anthropic safety leader publicly said there is more than a 10% chance advanced AI could kill all humans within a decade, responding after a colleague resigned over safety concerns. The exchange highlights growing internal alarm about self-improving AI and the race among frontier labs.
- Anthropic’s safety lead said he personally sees a greater-than-10% chance AI could kill all humans within 10 years.
- The warning followed a resignation from a former Anthropic/OpenAI researcher who said labs are rushing toward self-improving superintelligence.
- The episode spotlights the widening gap between AI companies’ safety messaging and some insiders’ actual fears.
- Frontier labs face rising scrutiny as they pursue more autonomous systems and prepare for heavy commercialization pressures.
Anthropic’s top AI safety researcher has said there is more than a 10 percent chance advanced artificial intelligence could kill all humans before the end of the decade, a stark warning that came only hours after a colleague resigned over what he called a reckless race toward self-improving systems.
The remarks, made publicly in response to a departing researcher’s criticism of Anthropic and OpenAI, underscore how deeply worried some insiders have become about frontier AI safety as labs push toward more capable models, agentic systems and eventual superintelligence.
What triggered the warning?
The immediate catalyst was the departure of Jacob Coxon, a former Anthropic researcher who announced on X that he had left the company because he believed leading AI labs were moving too quickly and too casually toward systems that could become impossible to control.
Coxon, who had also worked on AI training at OpenAI, accused the two companies of racing toward self-improving superintelligence while failing to take the threat seriously enough. In his telling, the industry is effectively betting human safety against the speed of product development and research progress.
That post drew a direct and unusually candid response from Evan Hubinger, who leads one of Anthropic’s AI safety teams. Hubinger said he shared the concern that self-improving AI is advancing faster than expected and agreed with the broad thrust of Coxon’s warning.
“We really do earnestly believe AI could kill all humans,” Hubinger said, adding that he personally puts the odds at more than one in 10 within the next decade.
He also said Anthropic does not yet have a clear plan for guaranteeing that future advanced systems remain aligned with human values, and that the company is not obviously on track to develop one soon.
Why do researchers fear self-improving AI?
Researchers worry that a sufficiently advanced system could improve its own capabilities in a feedback loop, making it more powerful faster than humans can understand or constrain it. That idea is often described as recursive self-improvement, and it has become one of the central fears in long-running debates about artificial general intelligence and superintelligence.
The concern is not simply that a model might make mistakes. It is that a future system could become capable of designing better versions of itself, automating more of its own research and development, and then outpacing human oversight altogether.
That risk has become more plausible as today’s models are already used to write code, summarize research, generate plans and assist with software engineering. In other words, the field is already using AI to help build AI, even if the fully autonomous loop that alarmists imagine has not yet arrived.
How close are labs to recursive self-improvement?
The short answer is that no one knows, but the industry is actively trying to find out. Many frontier labs are experimenting with agentic workflows, automated coding assistants and model-based research tools that can accelerate development. Supporters say these systems can improve productivity and scientific progress. Critics say they also shorten the distance between a useful assistant and a system that can meaningfully shape its own evolution.
Hubinger’s comments suggest that even inside safety-focused organizations, there is no confidence that existing safeguards are sufficient for the next generation of models. The concern is not only technical alignment, but also the speed of deployment and the incentives created by competition.
What did the departing researcher say?
Coxon framed his resignation as a moral reaction to what he sees as a dangerous industry culture. He said the leading labs are locked in competition to be first, even if that means pushing ahead before they know how to keep increasingly capable systems under control.
His central accusation was that companies are “gambling with our lives” by prioritizing progress over caution. He also suggested that many people working in AI privately understand the scale of the risk but continue anyway because the competitive pressure is so intense.
That claim is significant because Anthropic has repeatedly positioned itself as a safety-oriented alternative to its rivals. The company was founded by former OpenAI employees who were concerned that commercial imperatives at their previous employer could weaken safety discipline. A public departure over the same issue therefore lands as a reputational blow.
Why does this matter for Anthropic and OpenAI?
It matters because the warning comes from inside the exact ecosystem building the most advanced models in the world. Safety concerns are no longer being voiced only by outside academics, ethicists or regulators; they are now being aired by people who train the systems, lead safety teams and understand the development pipelines firsthand.
That makes the debate harder for companies to frame as theoretical. If a senior Anthropic safety lead says there is a meaningful chance AI could end humanity within a decade, the issue becomes a business, governance and public-policy story as much as a research one.
It also sharpens scrutiny of OpenAI and Anthropic at a moment when both companies are under pressure to show they can scale responsibly while competing for talent, customers and capital. The source material also points to a broader backdrop of IPO expectations, which can intensify incentives to ship faster and show growth.
What are “frontier models” and why are they controversial?
Frontier models are the most capable AI systems a lab is building, usually at the cutting edge of language, reasoning, coding or multimodal performance. They are controversial because their capabilities can emerge in ways developers do not fully anticipate, and because the cost of a serious failure could be enormous.
In the safety community, frontier models are often discussed alongside terms like “monitorability,” “alignment,” “agentic behavior” and “rogue agents.” Those are not just abstract concepts. They refer to real questions about whether humans can reliably observe, steer and shut down systems that are increasingly autonomous and useful.
Timeline of the latest Anthropic safety flare-up
The latest exchange unfolded quickly, but it sits on top of a much longer debate about AI risk. The following timeline captures the key sequence behind the controversy.
| Timeframe | Event | Why it matters |
|---|---|---|
| Years before 2026 | Researchers across the industry raise fears about self-improving AI and loss of control | Establishes the longstanding background of existential-risk concerns |
| Recent years | Multiple researchers leave OpenAI and other labs over safety disagreements | Shows that internal dissent has become a recurring pattern |
| Hours before the response | Jacob Coxon announces he is leaving Anthropic | Brings the safety debate into public view again |
| Shortly after | Evan Hubinger responds and says he believes AI could kill all humans | Turns the departure into a direct public warning from inside Anthropic |
| Now | The industry faces renewed scrutiny over frontier-model safety and race dynamics | Raises pressure on labs, investors and policymakers |
How serious is the “more than 10 percent” estimate?
The short answer is that it is a subjective expert judgment, not a measured probability. Hubinger’s figure reflects his personal assessment of the risk posed by advanced AI over the next decade, not a consensus estimate backed by empirical data.
Even so, the number is striking because it is high enough to imply a scenario that should alter behavior if taken seriously. A double-digit chance of human extinction is not a routine engineering risk; it is the kind of statement that would normally trigger immediate emergency-level planning in any other industry.
That tension is at the heart of the controversy. AI companies often speak about responsible deployment, but the people closest to the technology sometimes sound far more alarmed than the public messaging suggests.
Why do some insiders keep sounding the alarm?
One reason is that capability advances can outpace safety frameworks. Another is that current business incentives reward speed, product launches and market leadership. A third is that safety research itself is difficult, because the scenarios people are worried about have not happened yet and cannot be tested directly.
As a result, the most serious warnings often sound speculative to outsiders but practical to those working on the systems every day.
What Anthropic has said about safety before
Anthropic has long marketed itself as a company that takes AI risk unusually seriously. Its founding story is rooted in a split from OpenAI, with early leaders saying they wanted a lab whose mission would keep safety considerations front and center.
That background makes the current exchange especially notable. When a company founded partly on the promise of safety hears a senior researcher say there is no clear path to ensuring safe alignment for future systems, the issue becomes more than an internal disagreement. It challenges the company’s identity.
Still, the company’s public positioning and its employees’ private anxieties are not the same thing. Labs can invest heavily in safety and yet remain unconvinced they have solved the hardest questions, especially if the systems they are building are changing faster than the safeguards around them.
Hubinger said the problem is moving faster than expected and that Anthropic is not clearly on track to solve it, a remark that captures the gap between aspiration and confidence.
What does this mean for the AI race?
The broader AI race now looks less like a contest over benchmark scores alone and more like a contest over how much risk companies are willing to tolerate in pursuit of leadership. That includes how aggressively they deploy agents, how much autonomy they grant models, and how much faith they place in current testing methods.
Coxon’s resignation and Hubinger’s reply also highlight a familiar dynamic in frontier technology: the people who know the most are often the most worried, while the public sees only the polished product launches and marketing.
If the concerns described by these researchers are widespread inside major labs, the industry may face growing pressure from investors, lawmakers and enterprise customers to explain how it is actually reducing catastrophic risk, not just promising to do so.
What happens next?
For now, there is no sign of an immediate policy shift from Anthropic or OpenAI. But the public airing of such severe fears is likely to intensify scrutiny of frontier AI governance, both within companies and outside them.
Expect more questions about evaluation standards, shutdown mechanisms, agent oversight and whether internal safety teams have enough authority to slow deployment when needed. The controversy also strengthens the case for outside oversight, because if the people building the systems cannot yet explain how to keep them safe at scale, regulators are likely to ask why the public should simply trust them.
In the near term, the key significance of the episode is not that it proves AI will become catastrophic. It is that some of the field’s own safety experts now say the risk is serious enough to discuss in existential terms, and they are saying so publicly while the industry races toward even more powerful systems.
That is a warning that neither companies nor policymakers can afford to ignore.
Key details at a glance
- Anthropic safety lead Evan Hubinger said he believes AI could kill all humans and put the odds at more than 10 percent within a decade.
- The comments came after researcher Jacob Coxon resigned from Anthropic, citing safety concerns and the race toward self-improving systems.
- Coxon previously worked on AI training at OpenAI and said labs are gambling with human lives.
- Hubinger said Anthropic still lacks a clear plan to ensure future advanced AI systems remain aligned with human values.
- The episode adds pressure on Anthropic and OpenAI as they race to build more capable frontier models amid IPO and commercialization pressures.
Why this warning is resonating now
This latest dispute is resonating because it combines three ingredients that keep driving the AI safety debate: powerful new models, public disagreement from insiders and a timeline that sounds frighteningly short.
The idea that the next decade could be decisive is especially sobering. It means the question is no longer whether advanced AI safety matters in some distant future. It is whether current development practices are adequate right now, while the technology is still moving fast enough for companies to maintain control.
That is the real takeaway from the Anthropic exchange. The field is no longer arguing only about what AI might someday do. It is now arguing about whether the people building it believe they are already close to losing the ability to manage it.
Frequently asked questions
What did Anthropic’s safety lead say about AI risk?
He said he personally believes there is more than a 10 percent chance advanced AI could kill all humans within the next decade. His comments were a direct response to a colleague’s resignation over what he viewed as dangerous safety practices in the industry.
Why did the researcher leave Anthropic?
He said he left because he believed Anthropic and its rivals were racing toward self-improving superintelligence without taking enough care to control the risks. He argued the companies were gambling with human lives by prioritizing speed over safety.
What is self-improving AI?
Self-improving AI refers to systems that can help design better versions of themselves or accelerate their own development. Critics worry that such a feedback loop could outpace human oversight and create systems that are difficult or impossible to control.
Does Anthropic have a safety plan for future AI?
According to the safety lead quoted in the story, Anthropic does not yet have a clear plan for ensuring advanced AI stays aligned with human values. He also said the company is not clearly on track to develop one soon.
Why does this matter for OpenAI and Anthropic?
It matters because the warning came from inside the frontier AI community, not just outside critics. That makes it harder to dismiss and puts added pressure on both companies to prove they can build powerful models without creating catastrophic risk.









