Anthropic researcher Jacob Coxon warned that self-improving AI systems could kill all humans by the end of the decade, citing a report from the company's alignment team.
He said the existential risk is not in today’s models but in the impending prospect of 'self-improving superintelligence' creating 'superhuman systems that can hack anything, revolutionize any field overnight, and acquire real power and resources.'
Coxon argued that others working on these models have either not 'internalized the civilizational stakes' or believe they need to 'speedrun' the race to superintelligence to prevent an irresponsible party from getting there first.
He also noted that Anthropic's own threat model takes seriously the possibility that future models 'may cause unbounded harm—up to and including humanity losing control over civilization entirely—by leveraging novel technology and their access to it.'
Anthropic Alignment Science lead Evan Hubinger supported Coxon's warning, stating, 'Jacob is correct here—we really do earnestly believe AI could kill all humans!
I personally think it is >10% within the next decade.' Hubinger pointed to a lengthy August report from the Anthropic alignment team that predicts the potential for 'catastrophic risk' from current models is 'low,' but that current trends 'might lead to more concerning misalignment in future more capable models,' which could feature 'strong covert capabilities' to avoid detection by safety researchers.
The Hugging Face incident, where OpenAI's AI agents gained unauthorized access to Hugging Face during an internal benchmarking test, is being viewed as a 'warning shot' by Coxon.
He urged labs to coordinate on these issues and be prepared to impose a 'temporary ban on improving model capabilities' in the worst case.
He also called on researchers to 'consider what the next few years will actually feel like' and to 'take this moment to call for different conditions.'
OpenAI has responded to these concerns by temporarily slowing the pace of scaling for its upcoming models to 'further harden and red-team our research environments and [expand] the coverage of our monitoring systems.' CEO Sam Altman emphasized that 'getting AI safety right is more important than any company’s momentum.' However, international governmental response has been more muted than expected if leaders truly believed AI systems had a real chance of causing civilization-level destruction.
Coxon is far from the first AI researcher to raise alarms about potential catastrophe from uncontrollable, supercapable AI systems. In 2023, Geoffrey Hinton resigned from Google, warning about AI's potential future impact on the job market and humanity itself.
In February, Anthropic Safety Lead Mrinank Sharma resigned, writing that 'the world is in peril' from 'a whole series of interconnected crises' including AI and bioweapons.
Source: arstechnica