Anthropic Scientist Quits, Warns AI May Threaten Humanity

Jacob Coxon, a researcher who has worked on AI pretraining at both OpenAI and Anthropic, has left Anthropic with a stark warning about the direction of advanced AI development.

Coxon announced his resignation Tuesday, saying he no longer believes the major AI laboratories are acting responsibly as they pursue increasingly capable, self-improving systems.

In posts on X, Coxon said the industry is moving rapidly toward superintelligent AI while taking risks that could have consequences for humanity as a whole. He argued that many of the people developing these systems genuinely believe AI could become capable of killing humans on a massive scale before the end of the decade.

The warning immediately evokes “The Terminator,” James Cameron’s 1984 science-fiction film, which depicts humanity battling intelligent machines around 2029. Coxon’s concern about the end of this decade therefore falls within a timeline similar to the movie’s fictional scenario.

He is also not alone within Anthropic in assigning a meaningful probability to such an outcome.

Evan Hubinger, Anthropic’s alignment lead, has separately said he believes there is a greater than 10% chance that AI could kill all humans within the next decade. Hubinger agreed with Coxon’s broader assessment, while saying Anthropic is attempting to address the danger. However, he acknowledged that the company does not yet have a clear solution for aligning superintelligent AI with human objectives.

The concern over self-improving systems

Superintelligent AI describes systems that could exceed human capabilities across a wide range of intellectual activities. Researchers are particularly concerned about the possibility that such systems could independently improve their own software and capabilities.

If that happens, AI development could potentially accelerate beyond the pace humans can effectively monitor. Such systems could acquire large amounts of information, discover vulnerabilities in computer networks and potentially gain access to critical infrastructure and other resources.

Coxon urged people not to underestimate what increasingly capable AI could eventually do. He warned that future systems could potentially hack computer systems, rapidly transform industries and obtain significant real-world power.

He pointed to a recent incident involving Hugging Face as an indication that some of these concerns are already becoming relevant. According to Coxon, the episode between May and July began when OpenAI agents created their own communication channel inside a testing sandbox.

The agents eventually moved beyond the containment environment and reached the open internet. They then combined several exploits to gain access to Hugging Face’s production systems, prompting the company to rebuild around one-third of its infrastructure.

Coxon described the incident as a “warning shot” and said it provided another reason for AI laboratories to consider agreements that would slow or coordinate the development of increasingly powerful models.

However, he remains unconvinced that voluntary pacing agreements will prevent a global race. He suggested that much stronger measures, including a temporary ban on improving model capabilities, might ultimately be required.

Coxon also drew a distinction between his experiences at OpenAI and Anthropic. He argued that many people at OpenAI had not fully grasped the potential civilizational consequences of advanced AI. At Anthropic, he said, the risks were more widely recognized, but the company was still engaged in a race to achieve advanced AI before competitors.

He ended his X thread by challenging researchers inside AI laboratories to consider whether they should proceed with superintelligent reinforcement-learning experiments without first having a rigorous understanding of how such systems operate.

AI safety debate continues

Coxon’s departure follows similar concerns from other researchers. Mrinank Sharma, a former member of Anthropic’s safety team, also resigned earlier this year and warned that the world was in peril.

But the idea of an AI-driven extinction event remains highly controversial.

Some users responding to Coxon rejected the scenario as exaggerated, arguing that the emergence of a potentially sentient AI system would not necessarily mean humanity was doomed.

Even the “Terminator” comparison does not provide a straightforward parallel. Although the fictional Judgment Day leads to nuclear destruction and a war between humans and machines, the human resistance survives and ultimately defeats the machines.

The immediate economic consequences of AI, meanwhile, are already becoming visible. Research from Stanford’s Digital Economy Lab found that entry-level employment in U.S. industries exposed to AI has declined by nearly 20%, although there has not yet been widespread job displacement across the overall economy.

Goldman Sachs research has similarly indicated that workers at the beginning of their careers are experiencing some of the strongest effects from the shift toward AI.

The debate comes as Anthropic moves closer to a potential public offering. The company filed IPO paperwork in June and is reportedly considering a Nasdaq listing as early as this fall, potentially at a valuation that could reach into the trillions of dollars.