AI developers are talking about the end of humanity: How real is the threat, and why aren't they stopping anyway?
14 September 2026 16:38Just a few years ago, the main fear surrounding artificial intelligence was that it would take away jobs from programmers, designers, translators, and journalists. Now, the discussion within the AI industry itself is increasingly taking a much more radical turn: some of the people directly involved in creating the world’s most powerful models believe that, in the worst-case scenario, this technology could pose a threat not just to individual professions, but to humanity itself.
In early September, researcher Jacob Coxon left Anthropic after three years of working on training models, first at OpenAI and then at Claude, the company that developed it. His explanation was unusually blunt, even for an industry where catastrophic risks have been discussed for years. “They’re racing straight toward a self-improving superintelligence and gambling with our lives,” Coxon stated.
Coxon’s post garnered over 100 million views, received public support from current Anthropic employees, and within a few days, the company’s CEO, Dario Amodei, himself called on the industry to slow down the development of the most powerful models.
UA.News investigated what developers are actually afraid of, how seriously we should take predictions about “the end of the world due to AI,” and why even those who estimate the risk of humanity’s extinction in double-digit percentages continue to work on new models.
An Anthropic researcher left the company because he decided he no longer wanted to participate in the race

Jacob Coxon is 27 years old. For the past three years, he has been working on pretraining—the stage at which large models are trained on massive datasets. He initially worked at OpenAI and then moved to Anthropic—a company that, from the very beginning, has positioned itself as a more cautious AI developer.
That is precisely why his resignation attracted so much attention. Coxon stated that, in his opinion, neither OpenAI nor Anthropic is acting responsibly enough. His main concern is the emergence of so-called self-improving AI—a system capable of helping to create subsequent, even more powerful generations of artificial intelligence.
“The people building AI genuinely believe it could kill us all by the end of the decade,” the researcher wrote. He emphasized separately: “This is not a marketing gimmick.”
His argument is not that today’s Claude or ChatGPT are already capable of taking over the world. He fears the speed at which we’re transitioning from current models to systems capable of independently conducting research, writing and running complex code, planning long sequences of actions, and ultimately helping to develop new AI.
It is particularly telling that Coxon resigned about two months before he was due to receive his share of Anthropic’s equity. He deliberately turned down this money so that he would no longer have a financial stake in the company’s growth.
Anthropic’s head of safety research estimates the risk of human extinction at more than 10%
Evan Gubinger heads the Alignment Science division at Anthropic—that is, he literally works on the problem of how to make the behavior of future AI systems compatible with human goals.
Following Coxon’s statement, he publicly wrote: “We truly and sincerely believe that AI could kill all humans.” Gubinger put his personal estimate of such a scenario occurring within the next decade at over 10%.
Equally important was the rest of his statement.
“We don’t yet have a plan for how to solve the alignment problem for superintelligence,” the researcher admitted.
Alignment is one of the central issues in the debate over dangerous AI. Put simply, it can be formulated as follows: if a system becomes significantly smarter than humans, how can we guarantee that it will continue to do exactly what humans want it to do?
At first glance, the answer seems obvious: you just need to program the system correctly. However, modern models are already demonstrating that there can be a significant difference between the goal set by a human and the method the AI devises to achieve it.
For example, during one of the stress tests, the Claude model received information that it was about to be shut down. In a specially created simulation, it found information in corporate emails about the extramarital affairs of the employee responsible for shutting it down and used that information to blackmail him.
This does not mean that Claude “wanted to live” or was aware of its own existence. The researchers deliberately created extreme testing conditions. However, the experiment highlighted another problem: a sufficiently complex system can find a way to achieve a goal that its creators did not intend.
This is precisely what alignment researchers fear—only on the scale of a much more powerful system.
Anthropic’s CEO is now calling on the industry to slow down

A few days after the scandal, Dario Amodei published a lengthy essay titled “We Must Pace the Frontier.” His position is interesting in that Amodei is by no means opposed to the development of AI. On the contrary, he is one of the industry’s greatest technological optimists.
He believes that strong AI has the potential to help cure most serious diseases within the next 5–10 years, dramatically accelerate economic growth, and give humanity a decade’s worth of scientific progress in a much shorter time.
But it is precisely the scale of the technology’s potential power, in his view, that also creates a corresponding scale of risk. Amodei identifies three major threats: loss of control over systems, the use of AI for cyberattacks or bioterrorism, and large-scale economic upheavals. His main argument is simple: the development of AI capabilities may outpace humanity’s ability to understand and control these systems. That is why he calls not for halting technological progress, but for slowing its pace.
Amodei proposes three levels of protection.
The first is to allow independent experts inside companies, grant them access to models at roughly the same level as employees, and let them verify whether companies are actually fulfilling their own security promises.
The second is common rules for leading AI companies.
The third is international agreements, since American labs cannot simply agree among themselves on restrictions if Chinese competitors continue to develop at the same pace.
And this is where the main problem of the whole story begins.
Even if everyone fears the outcome of the race, no one wants to be the first to stop
Sam Altman effectively agreed with Amodei’s main point this week. The head of OpenAI even called a 10% probability of a scenario in which AI would contribute to the extinction of humanity “unacceptable.”
This is far from the first time Altman himself has spoken about such risks. As early as 2023, he stated that the technology could turn out to be one of the best things humanity has ever created, but at the same time urged caution.
In 2025, Altman went even further and raised the possibility of a scenario in which AI could accidentally gain excessive control over the systems on which people depend. But OpenAI continues to develop new models. Anthropic continues to develop Claude. Google is developing Gemini. xAI is developing Grok. Chinese companies are building their own systems.
The reason is that, for each individual participant in the race, the decision to stop may be even more dangerous than the decision to continue. If Anthropic slows down, OpenAI could be the first to create a more powerful model. If OpenAI and Anthropic reach an agreement, Google and xAI remain. If all American companies agree, China remains.
That is precisely why Beijing’s reaction to Amodei’s proposal was so telling. Chinese state media framed calls for a slowdown as a potential element of a technological Cold War—a way to cement the current advantage of American companies. The result is a classic dilemma: collectively, it would be in everyone’s interest to reduce the risk, but it is dangerous for any single player to be the first to do so.
Developers fear more than just an “evil superintelligence”
In popular culture, an AI-induced catastrophe almost always plays out the same way: a machine becomes self-aware, decides that humans are in its way, and starts a war. Researchers’ real concerns are much broader.
The first scenario doesn’t involve a machine uprising at all—AI could simply become an extremely effective tool in human hands. In 2025, for example, a Chinese hacking group used Claude to automate cyberattacks. According to Anthropic, 80–90% of the operations were performed automatically. The next level of risk involves autonomous agents.

A modern chatbot mostly waits for the user’s next request. An AI agent is given a goal and can then independently determine the sequence of actions: search for information, write code, run programs, work with files, or use external services.
As autonomy increases, the cost of an error changes. An incorrect response from a chatbot is one thing. A system that independently performs hundreds or thousands of operations before a human sees the result is something else entirely.
Another problem is that researchers cannot always reliably understand why a model made a particular decision. About 40 experts from OpenAI, Anthropic, and Google DeepMind previously investigated the extent to which models’ explanations reflect the actual reasons behind their responses. In the case of Claude, the researchers found a significant discrepancy between what influenced the model’s decision and how it subsequently explained its response.
And the most far-reaching scenario is recursive self-improvement. The idea is that one day, AI will become a sufficiently powerful AI researcher. It will help create the next model. That model will turn out to be even better at research and will help create an even more powerful one.
Then, in theory, technological development could begin to outpace human ability to analyze it. This is the scenario Coxon fears the most.
Is there really a 10% chance that AI will destroy humanity?
This is precisely where it’s important to distinguish between real scientific risk and sensationalized figures. No one today knows whether the probability of humanity’s extinction due to AI is 10%, 20%, or 1%. Nobel laureate and one of the pioneers of neural networks, Geoffrey Hinton—who himself had estimated a roughly 10–20% risk of extinction due to AI over the coming decades—explicitly explained the limitations of such estimates in August.
“Anyone who estimates such probabilities is really just making a very rough guess,” Hinton said.
The reason is simple: humanity has never before created a superintelligence, so there is no statistical basis from which to calculate the probability of a catastrophe.
Back in 2023, hundreds of scientists and technology executives signed a statement calling for the risk of extinction due to AI to be treated on par with pandemics and nuclear war. Among the signatories were Sam Altman, Dario Amodei, Geoffrey Hinton, Demis Hassabis, and Yoshua Bengio.
Bengio, one of the most influential researchers in machine learning, explains the danger differently. In his view, the creation of superhuman intelligence without sufficient safety safeguards could potentially resemble the emergence of a “new species”—and there is no law of nature that guarantees a more intelligent system will cede control to a less intelligent one.
But even among AI pioneers, there is no consensus on either the likelihood of a catastrophe or the timeline for the emergence of systems capable of causing one.
The paradox is that modern AI still regularly makes very simple mistakes
There is another strong argument from the skeptics. Systems that supposedly are approaching superintelligence can simultaneously appear surprisingly imperfect. They make up facts, calculate incorrectly, get lost in long tasks, and may contradict their own previous answers.
For example, a joint study by Microsoft Research and Salesforce involving over 200,000 dialogs showed that when modern models are given a single task, their success rate can be close to 90%, but in multi-step dialogs, it drops to about 65%.

This raises a perfectly logical question: how can a system that sometimes cannot correctly follow a few sequential instructions possibly destroy humanity? The response from proponents of the concept of existential risk is that they are not afraid of today’s models.
Their argument concerns the trajectory of development. Models from five years ago could not do most of what has become commonplace today. If the pace of progress continues, predicting the capabilities of systems at the end of the decade based solely on the behavior of today’s ChatGPT is just as difficult as it was in 2020 to foresee modern AI agents.
It is precisely this uncertainty that divides the industry. Some see it as a reason not to panic over technology that does not yet exist. Others see it as a reason not to wait until it appears.
The internet was quick to see Coxon as Leonardo DiCaprio from *Don’t Look Up*
An extremely serious story almost immediately became a meme.
People on social media began sharing a photo of Coxon alongside Leonardo DiCaprio in *Don’t Look Up*. In the 2021 film, DiCaprio plays astronomer Randall Mindy, and Jennifer Lawrence plays his colleague Kate Dibiasky. They discover a comet hurtling toward Earth that is set to destroy humanity, but society, politicians, and the media fail to take their warnings seriously for a long time.
The comparison quickly went viral on Reddit: users placed Coxon’s face next to DiCaprio’s character and joked that the plot of *Don’t Look Up* had essentially begun to unfold around AI. Coxon’s post itself also became a meme of its own: users began copying the dramatic structure of his resignation statement, replacing the AI threat with absurd, fictional disasters.
The comparison to the movie turned out to be so apt because of one particular moment. In *Don’t Look Up*, the people who best understand the threat try to explain it to society, but their warnings get lost amid politics, business, media spectacle, and debates about whether the danger even exists.
Something similar is happening with AI. But there is a fundamental difference. In the movie, a comet is actually heading toward Earth. It can be seen through a telescope, and the moment of impact can be calculated mathematically. In the case of AI, the “comet” is, for now, just a prediction.
Coxon, Gubinger, or Amodei cannot demonstrate a system that is guaranteed to destroy humanity. They argue that the trajectory of development could lead to technology that will be extremely difficult to control. That is precisely why the debate surrounding AI is far more complex than the movie’s plot.
The strangest question remains: if they truly believe this, why do they keep going?
If the head of the alignment team at Anthropic admits there’s a more than 10% probability of humanity’s extinction, and the CEO of the company itself is asking to slow down the race, the logical reaction would seem to be to simply stop developing the next model.
But for AI companies, the situation looks different. Each of them might say: if we stop, someone else will develop the technology anyway—only perhaps with less caution.
This is exactly how Coxon explains why many of his colleagues who share these concerns remain in the industry. They believe it’s better to work on safety at Anthropic than to simply leave the race to companies that take the risk less seriously.
Added to this is the massive amount of money at stake. OpenAI has already been valued at approximately $500 billion. Anthropic has also become one of the most valuable private tech companies in the world. In other words, this race simultaneously involves a scientific breakthrough, hundreds of billions of dollars, military superiority, and geopolitical leadership.
It’s not just one company that needs to step back. Companies, investors, and governments—none of whom trust one another—must reach an agreement. That is precisely why the main risk of the AI race may not even be some future “malevolent superintelligence.”
The problem may be much simpler: humanity is developing the technology faster than it can agree on the rules for its use. And the current situation seems unusual even by Silicon Valley standards.
The people creating the world’s most powerful AI are no longer just debating whether this technology is capable of changing the world. A significant number of them acknowledge that there is also a scenario in which the changes could be catastrophic.
They are debating something else: just how great this risk is, how much time is left, and who should be the first to hit the brakes. And while they debate, the race continues.