AI Researcher Quits Anthropic With Warning That Superintelligence Could Become Impossible To Control


A researcher who has worked at both Anthropic and OpenAI has walked away from his job with a warning that has sent shockwaves through the artificial intelligence industry. Jacob Coxon says he resigned from Anthropic because he believes the companies developing the most powerful AI systems are taking risks that could eventually put humanity in serious danger. His post spread rapidly online, attracting more than 70 million views as people debated whether the technology is advancing faster than safety measures can keep up.

Coxon’s warning is especially striking because it comes from someone who has worked inside the industry. He claims the people building these systems “earnestly believe that it could kill us all by the end of the decade.” Coxon also warned that future AI systems could become capable of hacking systems, transforming entire fields almost overnight, acquiring resources and eventually improving themselves. The claim has reignited a debate that has followed advanced AI for years: how much control will humans actually have if machines become dramatically more capable?

The Researcher Says He Could No Longer Stay

Coxon announced his resignation in a post on X, directly accusing Anthropic and OpenAI of moving too quickly toward increasingly advanced AI. He said the companies were “gambling with our lives” while pursuing systems that could eventually reach superhuman levels of capability. His warning was not framed as a distant possibility that belongs entirely to science fiction. Coxon presented it as a risk that people working on the technology already take seriously.

“Do not underestimate the power of this technology,” Coxon wrote. He warned that future systems could “hack anything, revolutionize any field overnight, and acquire real power and resources.” Coxon then made his central accusation: “Neither company is acting responsibly. They are racing straight to self-improving superintelligence.” The post quickly became one of the most widely discussed warnings about AI safety, partly because of Coxon’s background working for the very companies he was criticizing.

The dispute centers on what could happen if AI systems become capable of making substantial improvements to themselves. That capability does not currently exist in the fully autonomous form described by the most extreme scenarios. Researchers are nevertheless studying how AI could increasingly assist with the development of future AI systems, creating a feedback loop in which increasingly capable models help build increasingly capable successors.

For Coxon, that possibility is reason enough to question the industry’s current speed. His argument is that once AI reaches a certain level of capability, humans could find themselves dealing with systems whose abilities grow faster than existing safety techniques can manage. The concern is not that today’s chatbot is secretly preparing to destroy civilization. It is what happens if future systems become powerful enough that human oversight stops being sufficient.

An Anthropic Researcher Says The Risk Is Real

One of the strongest responses to Coxon’s warning came from Evan Hubinger, an alignment lead at Anthropic. Rather than rejecting Coxon’s assessment, Hubinger publicly agreed with the central concern and offered his own estimate of the potential danger. His response gave the controversy another jolt because it came from someone still working inside one of the companies at the center of the debate.

“Jacob is correct here—we really do earnestly believe AI could kill all humans!” Hubinger wrote. He added that he personally believed there was a greater than 10% chance of that happening within the next decade. Hubinger also defended Anthropic’s intentions while acknowledging a major unresolved problem: “I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to.”

Alignment is the term used for efforts to make AI systems behave according to human intentions and values. It becomes increasingly complicated as systems grow more capable because researchers have to anticipate what a powerful model might do in situations that were never specifically programmed by its creators. Monitoring presents another challenge because the more sophisticated a system becomes, the harder it may be for humans to understand every decision or action it takes.

That creates a strange contradiction inside the AI industry. Companies are racing to develop systems that can perform increasingly sophisticated tasks, while some of the researchers working on those systems are openly acknowledging that there are still major unanswered questions about controlling a future superintelligent system. The debate is no longer only about what AI can do today. It is about whether humanity can safely manage what comes next.

OpenAI’s Chief Scientist Wants The Race To Slow Down

The concerns have also surfaced within OpenAI itself. Chief scientist Jakub Pachocki recently warned that AI development may be approaching a point where laboratories cannot responsibly continue scaling at maximum speed without stronger safety measures. His comments add weight to the argument that concerns about advanced AI are not limited to critics outside the industry.

Pachocki wrote that no AI company has “solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer.” He said he expected voluntary slowdowns to become more common until shared safety standards were established. He also argued that governments around the world should make international coordination on future AI development a major priority.

The warning is significant because it addresses the same problem Coxon raised from a different position. If competing companies believe they must keep moving quickly because another company might get there first, slowing down can become difficult even when researchers recognize serious risks. Each laboratory may have an incentive to continue pushing forward, creating a race in which safety decisions become entangled with commercial and strategic pressure.

The result is an industry facing an uncomfortable question. If the people developing increasingly powerful AI systems are themselves saying that alignment and monitoring remain unsolved, how much further should the technology advance before those problems are addressed? There is no agreed answer, and that uncertainty is becoming one of the defining arguments surrounding the next generation of AI.

Self-Improving AI Is The Scenario Researchers Fear

The phrase “self-improving superintelligence” sounds like something from a futuristic movie, but researchers have been studying the possibility for years. The basic idea involves an AI system becoming capable of contributing to the creation of a more capable successor. That successor could then contribute to further improvements, potentially creating a cycle of increasingly rapid development.

Recursive self-improvement is not currently possible in the fully autonomous way imagined in extreme AI scenarios. However, AI systems are already being used to write code, conduct research, analyze information and assist engineers. As those capabilities improve, AI can play a larger role in developing the next generation of AI technology itself.

The distinction is important. An AI system helping a programmer write software is not the same as an AI independently designing a successor that surpasses its creators. The concern among some researchers is that the gap between those stages could eventually become difficult to predict if AI systems continue improving rapidly.

That is why recursive self-improvement has become such a serious subject within AI safety research. If a system eventually becomes capable of substantially improving its own capabilities, traditional human oversight could face a completely different challenge. Researchers would potentially be trying to monitor a system that can change faster than the methods used to monitor it.

The Fear Of AI Extinction Did Not Start This Week

Warnings about AI causing catastrophic harm have existed for years. In 2023, prominent AI researchers and executives signed a statement arguing that reducing the risk of extinction from AI should be treated as a global priority alongside other large-scale threats, including pandemics and nuclear war.

Some researchers use the term “p(doom)” when discussing the estimated probability of a catastrophic AI outcome. The estimates vary widely, and there is no universally accepted figure. Coxon’s comments and Hubinger’s estimate have nevertheless brought the subject back into public attention at a time when AI capabilities are advancing quickly.

Hubinger was also among roughly 1,400 AI researchers who signed an open letter calling for deliberate pacing of frontier AI development. The letter argued that governments should develop tools capable of slowing or controlling the development of increasingly automated AI systems. The push reflects growing concern among researchers that market competition alone may not provide enough incentive for companies to slow down.

The debate has now moved into Washington as well. Rep. Jay Obernolte and Rep. Lori Trahan introduced the FRONTIER Act, which seeks to establish a framework for overseeing advanced AI models. Sen. Bernie Sanders and Rep. Greg Casar also introduced legislation that would temporarily pause advanced AI development until federal safety rules are established.

The Political Fight Is Getting Harder To Ignore

The argument over AI safety is no longer confined to research laboratories. Lawmakers are increasingly dealing with questions about how advanced models should be tested, monitored and deployed, while communities are also pushing back against the massive data centers required to run them.

Opposition to AI data centers has grown into a political issue of its own. The facilities require huge amounts of computing infrastructure and energy, and resistance has become strong enough to attract attention from politicians preparing for the next election cycle. The debate is therefore expanding beyond hypothetical questions about superintelligence and into immediate disputes over electricity, jobs, infrastructure and who benefits from the AI boom.

Treasury Secretary Scott Bessent has also criticized AI companies for failing to adequately explain their plans and benefits to the public. He argued that companies would have to convince Americans that the benefits of AI would not simply flow to a small group of people. That concern reflects a broader tension surrounding the technology: people are being asked to accept enormous changes while still trying to understand who will control the systems driving them.

For AI companies, the stakes are enormous. The race could produce major advances in science, medicine, software and productivity. But the faster the technology develops, the more difficult it becomes to separate commercial competition from questions about safety and control.

Nobody Knows Where The Race Ends

There is no evidence that today’s AI systems are about to wipe out humanity, and there is no scientific consensus that extinction is inevitable. What is real is the disagreement among researchers over how much risk society should accept while developing increasingly powerful systems.

Coxon’s resignation has given that disagreement a face. Hubinger’s response has made the internal debate harder to dismiss, while Pachocki’s call for voluntary slowdowns shows that concerns about the pace of development exist inside the industry’s biggest laboratories.

The most unsettling part of the story is not a robot suddenly turning against its creators. It is the possibility that humans could build systems faster than they can figure out how to control them.

That leaves one question hanging over the AI race: if the people closest to the technology are warning that they have not solved the safety problem, how far should the industry go before it slows down?

Loading…


Leave a Reply

Your email address will not be published. Required fields are marked *