
The resignation of an Anthropic employee has prompted chatter from other employees about the risk of AI destroying humanity.
The controversy erupted when a former AI researcher at Anthropic, Jacob Coxon, announced his resignation and warned that neither his former company nor OpenAI is "acting responsibly."
"They are racing straight to self-improving superintelligence and gambling with our lives," he wrote. Although many people use AI through chatbots, Coxon is particularly concerned about future AI models that won't require human software developers to improve, but could upgrade themselves at a superhuman rate.
"Do not underestimate the power of this technology," he added. "These will soon be superhuman systems that can hack anything, revolutionize any field overnight, and acquire real power and resources. We have all witnessed the progress in each of these domains, and progress is not slowing."
In another chilling warning, Coxon claims that "the people building AI earnestly believe that it could kill us all by the end of the decade. This is not a marketing stunt. If anything, many executives and senior researchers will couch their phrasing in the press to sound sensible -- but I hear the same people express fear privately. No other human activity poses this level of danger."
In a surprise, another Anthropic employee focused on AI safety work, Evan Hubinger, agreed with Coxon, rather than downplay it, as companies usually do. "Jacob is correct here -- we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade," he wrote in his own tweet.
Hubinger works on AI alignment, which involves ensuring the technology follows intended goals, such as upholding ethical values. He added: "Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to."
A greater than 10% chance is still relatively low. That said, Hubinger says he's worried about "superintelligence arising from recursive self-improvement," or where the AI model rewrites its own computer code, creating better versions of itself.
A third Anthropic employee, safety researcher Samuel Marks, also chimed in with his own bird's eye view, tweeting that "many AI developer staff desperately want to slow down to figure out how to build AI more safely." But commercial incentives, fierce competition with other AI companies, and fears that less responsible developers will obtain powerful AI, prevent anyone from hitting the brakes. However, Marks remains on the team, saying, "I hope my work will reduce the chance of these extinction-level bad outcomes."
Meanwhile, Coxon claims Anthropic is still barreling ahead because "they believe no one else will act responsibly, so they must do it themselves, despite the risk," he wrote.
So far, Anthropic hasn't responded to a request for comment. Although AI doomerism is nothing new, the fact that the latest warning comes from insiders at a leading artificial intelligence lab is fueling further fears about AI development. OpenAI staffer John Wolfe also weighed in, tweeting: "I don't know what my probabilities are on literal extinction, but I think there are a number of ways AI could go poorly for humanity, and at the current frankly terrifying pace, humanity will be quite lucky if we manage to find and stay on the narrow path between all the bad outcomes."
Microsoft co-founder Bill Gates has also warned that society is failing to create a plan to steer AI in a positive direction. "Even under the best circumstances, the transition to this new AI era will be one of the most turbulent times in human history," he wrote last month. That said, AI insiders have been wrong before or have overpromised on the technology's capabilities.