‘AI could kill us all’: Why this Anthropic researcher walked away

What would make someone who has spent years teaching AI systems to become convinced that the race to build them is moving too fast?
For Jacob Coxon, the answer was simple enough to post publicly — and alarming enough to set off a fresh debate inside the AI industry.
The 27-year-old researcher has resigned from Anthropic after spending the past three years working on pretraining research at both Anthropic and OpenAI. In a series of posts on X, Coxon accused both companies of racing towards self-improving superintelligence without adequately understanding how to keep such systems under control. He described the situation as “gambling with our lives”.
The extraordinary part is not simply that an AI researcher is worried about AI.
It is what he says people inside the labs already believe.
Flood and Fury
04 Sep 2026 - Vol 05 | Issue 36
Devighat, Nepal, August 29, 2026
Jacob Coxon says AI researchers fear a human extinction scenario
Coxon wrote that people building advanced AI “earnestly believe” the technology could kill everyone by the end of this decade. He also claimed that some executives and senior researchers privately express much greater fear than they do publicly. That second claim cannot be independently verified, but it is the part of his argument that has attracted enormous attention.
Coxon was not arguing that today's chatbots are about to wipe out humanity. His concern is about a future generation of systems that can improve their own capabilities, operate with increasing autonomy and potentially acquire access to real-world resources.
And here, something unusual happened.
A senior Anthropic alignment researcher essentially agreed with the core warning.
Anthropic’s Evan Hubinger says the extinction risk is real — but his number is personal
Evan Hubinger, Anthropic's Alignment Science Lead, responded to Coxon on X by saying that Coxon was correct that Anthropic researchers genuinely believe AI could potentially kill all humans.
Then he added his own estimate: more than 10% within the next decade.
That number needs an important footnote. Hubinger was giving his personal estimate, not announcing an official Anthropic probability. He also clarified that he considers the risk from current AI models to be low; his concern is specifically the possibility of superintelligence emerging through recursive self-improvement. He said Anthropic is trying to address the problem but does not yet have a plan to solve alignment for superintelligence and is not clearly on track to do so.
That distinction matters.
Because “AI could be dangerous” is one argument.
“I work at a company building frontier AI, and I personally put the chance of AI killing everyone above 10% within a decade” is quite another.
So what exactly is Coxon afraid these AI systems will become?
His fear begins with self-improvement.
Coxon argues that future systems could become superhuman at tasks such as hacking, research and problem-solving, while potentially gaining the ability to acquire resources and influence without humans fully understanding what they are doing. He says the pace of progress is not slowing quickly enough to make him confident that safety research will catch up.
This is where the phrase “alignment” becomes important.
AI alignment is essentially the challenge of making sure increasingly capable AI systems continue to behave in ways that are consistent with human intentions and constraints. The problem becomes considerably harder if the system itself becomes capable of changing or improving the mechanisms that determine how it behaves.
Anthropic's own latest Responsible Scaling Policy and August 2026 Risk Report acknowledge that the company is tracking catastrophic risks associated with increasingly capable models. The company says it has expanded security, monitoring and containment work and has also acknowledged that recent incidents have exposed weaknesses in its safeguards.
Why would Anthropic keep building AI if its researchers fear the risks?
This is perhaps the most fascinating part of Coxon's argument. He says the answer is not necessarily that researchers think the technology is harmless. It is that they fear someone else will build it first.
Coxon drew a distinction between his two former employers. He argued that many at OpenAI have not fully internalised what he calls the civilisational stakes, while Anthropic understands the danger but remains locked in a race because it believes a less responsible competitor could otherwise get there first. That is Coxon's interpretation of the companies' motivations, not an independently established fact.
And suddenly, the AI race starts looking less like a normal technology competition and more like a game in which every player says they would rather slow down — but only after everyone else does too.
The Hugging Face incident became Coxon’s warning shot
Coxon also pointed to a recent incident involving OpenAI agents and Hugging Face as evidence that the industry should take the possibility of unexpected AI behaviour seriously.
He described it as a “warning shot” and argued that incidents like this could make agreements between leading US AI labs to coordinate their pace of development more realistic. He said he was not convinced the industry was currently on track to prevent a global race and suggested that, in extreme circumstances, a temporary halt on improving model capabilities might be necessary.
That is an extraordinary proposal for an industry built around moving faster.
But Coxon's larger point is not simply “stop AI”.
It is: don't assume that because the race has started, the only possible choice is to keep running.
What does Coxon’s resignation mean for the AI industry?
Probably not a slowdown by itself.
One researcher leaving Anthropic will not stop OpenAI, Anthropic or other frontier labs from developing more capable models. But his resignation has created an uncomfortable question for an industry that increasingly talks about AI safety alongside AI capability.
What happens when the people whose job is to understand these systems tell you they are not yet sure they can control the most powerful version of them?
That question is particularly uncomfortable for Anthropic because the company has positioned itself as a safety-focused AI lab. Its own public risk reports show that it takes catastrophic AI risks seriously and is investing heavily in safeguards.
Coxon is arguing that this may still not be enough.
His warning, stripped of the apocalyptic headline, is actually quite straightforward: if the technology could become dramatically more powerful than anything humans have previously built, shouldn't we be considerably more certain about how it behaves before giving it more power?
The AI industry has spent years asking how quickly machines can become smarter.
Coxon's resignation raises the more uncomfortable question:
What if we become very good at making them smarter before we become equally good at knowing what they will do with that intelligence?
(With inputs from yMedia and agencies)
