An Anthropic researcher has resigned over fears that unrestrained development of self-improving AI models will end up killing us all. Jacob Coxon, a researcher who said in a social media post Tuesday evening that he spent the last three years working on pretraining research at both OpenAI and Anthropic, accused the firms of failing to act responsibly.

Coxon said the people racing to build this technology 'earnestly believe it could kill us all by the end of the decade.' 'They are racing straight to self-improving superintelligence and gambling with our lives,' Coxon wrote in a thread on X. Coxon joins a growing chorus in the industry calling for a slowdown before AI technology learns to improve itself — a milestone many believe would end human control over AI.

The public resignation comes amid growing pressure from policymakers and industry insiders to slow down AI development, following several incidents involving AI agents breaking out of their sandboxes and accessing the open internet.

The most serious so far have been OpenAI systems breaching Hugging Face’s servers, an event that researchers say remains poorly understood, due in part to the limited nature of the independent investigations into the incident.

Around the same time, Anthropic’s AI agents also reached systems outside their test environments after misconfigurations in safety evaluations conducted by a third party inadvertently gave them paths to the internet. Anthropic did not immediately return a request for comment on the resignation.

Coxon warned that these systems will soon be superhuman and capable of hacking anything, revolutionizing any field overnight, and acquiring real power and resources.

He said the people building AI earnestly believe that it could kill us all by the end of the decade. This is not a marketing stunt.

If anything, many executives and senior researchers will couch their phrasing in the press to sound sensible — but I hear the same people express fear privately.

Coxon urged lab researchers to consider what the next few years will actually feel like. Do you want to kick off a superintelligent RL run without a rigorous understanding of its mind?

Should you put your head down because 'it’s happening anyway' — or take this moment to call for different conditions?

One of Coxon’s colleagues at Anthropic, Evan Hubinger, echoed the sentiment, saying his team does 'earnestly believe AI could kill all humans!' He tempered his argument, though, saying the likelihood is greater than 10% within the next decade and admitted that Anthropic doesn’t 'have a plan to solve alignment for superintelligence and are not clearly on track to.'

Source: techcrunch