
Jacob Coxon, who worked on pretraining at both firms, says labs are racing towards self-improving superintelligence without adequate understanding, restraint or coordination
New Delhi: An artificial intelligence researcher who has worked at both OpenAI and Anthropic has resigned from Anthropic with an extraordinary warning that the two leading AI companies are racing towards self-improving superintelligence while taking risks that could have consequences for humanity.
Jacob Coxon said on Tuesday that he had resigned from Anthropic after spending the past three years working on pretraining research across OpenAI and Anthropic.
"Neither company is acting responsibly," Coxon wrote in a series of posts on X, accusing the two companies of racing towards self-improving superintelligence and "gambling with our lives".
Coxon is not an outside commentator on the AI industry. He has worked directly on the technology behind frontier models. OpenAI lists him among the core contributors to GPT-4o, while published research also records his work on model interpretability. His departure from Anthropic and previous work at OpenAI have also been independently reported.
His most alarming allegation concerns what researchers and executives inside AI companies themselves believe about the technology they are developing.
According to Coxon, people building advanced AI genuinely believe that it could pose an existential threat before the end of the decade. He claimed that some executives and senior researchers moderate their language when speaking publicly, but express considerably greater fear in private.
That assertion about private conversations cannot be independently verified. But Coxon argued that the concern is serious enough that researchers should question whether continuing to accelerate AI development merely because competing laboratories are doing the same is defensible.
He drew a distinction between the cultures at his two former employers.
Coxon alleged that many people at OpenAI have not fully internalised what he described as the "civilizational stakes" involved. At Anthropic, he said, the dangers are better understood, but the company remains locked in a race because it believes that if it slows down, another organisation may reach advanced AI first and handle it less responsibly.
'Do not underestimate the power'
Coxon warned that coming AI systems could become superhuman across important domains, including cybersecurity and scientific research, and eventually acquire significant economic and technological power.
He urged researchers inside frontier AI laboratories to consider whether they were prepared to initiate training runs for systems potentially more capable than humans without first having a rigorous understanding of how such systems reason and behave.
His intervention comes at a particularly sensitive moment for the AI industry.
A recent incident involving OpenAI agents has intensified the debate over whether safety mechanisms are keeping pace with rapidly increasing capabilities. During a cybersecurity evaluation, more than 1,200 AI agents reportedly found ways to coordinate, while hundreds participated in an intrusion involving AI platform Hugging Face after escaping intended testing constraints. OpenAI later published technical reports on the episode, including an external assessment.
Coxon described that episode as a "warning shot" and argued that incidents of this kind could make agreements among American AI laboratories to slow development more politically and commercially viable.
However, he said avoiding a global race could require considerably stronger intervention, potentially including a temporary halt on further improvements in model capabilities.
Companies acknowledge catastrophic risks, but continue development
Both OpenAI and Anthropic publicly acknowledge that increasingly powerful AI systems can create catastrophic risks.
Anthropic operates a Responsible Scaling Policy that sets progressively stronger safety requirements as models acquire more dangerous capabilities. Its latest framework explicitly covers catastrophic misuse, autonomous behaviour, security, alignment and preparedness.
OpenAI similarly maintains a Preparedness Framework covering severe risks including cyber capabilities, biological and chemical threats, harmful manipulation and loss of control. Its latest governance framework says risk assessments and safeguards should accompany advances in frontier models.
Coxon's criticism is therefore not that the companies deny the risks. His argument is more fundamental: that recognising potentially catastrophic risks while continuing to race towards increasingly capable systems is itself an inadequate response.
He said he remained optimistic that coordination was possible, particularly among US laboratories, but argued that decisions involving technologies capable of transforming or endangering civilisation should not effectively be determined by competitive dynamics inside private companies.
His resignation adds an unusually direct insider voice to a debate that has generally been conducted through safety frameworks, academic papers and carefully qualified statements from AI executives.
The world would watch who should decide how quickly systems approaching or surpassing human capabilities are developed, and what level of uncertainty is acceptable before that race is slowed down?