Anthropic Insiders Warn AI Could Wipe Out Humanity: A 10% Chance by 2036
Market Updates

Anthropic Insiders Warn AI Could Wipe Out Humanity: A 10% Chance by 2036

WebProNews1d ago

Two researchers at Anthropic delivered a blunt message this week. One quit. The other put hard numbers on humanity's possible end.

Jacob Coxon resigned from the company Tuesday. He had spent three years on pretraining research. First at OpenAI. Then at Anthropic. His departure post on X pulled no punches. "Neither company is acting responsibly," he wrote. "They are racing straight to self-improving superintelligence and gambling with our lives."

Coxon added that the people building these systems "earnestly believe that it could kill us all by the end of the decade." He insisted this was no marketing stunt. Executives soften their words in public. Privately, he said, they express fear. "No other human activity poses this level of danger." From BBC News.

The Numbers Behind the Alarm

Evan Hubinger responded almost immediately. As Anthropic's alignment science lead, he leads efforts to make sure advanced AI systems follow human intentions. His reply carried weight. "Jacob is correct here -- we really do earnestly believe AI could kill all humans!" Hubinger wrote. "I personally think it is >10% within the next decade."

He stressed the risk from today's models remains low. The danger lies ahead. In recursive self-improvement. Systems that enhance themselves faster than humans can track. "I believe Anthropic is trying its best," Hubinger continued, "but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to." From The Verge.

Samuel Marks, Anthropic's scalable-oversight lead, backed the concerns. "AI developers believe their technology could cause human extinction," he posted. "The more senior the employee, the more concerned they are." The trio's posts spread quickly. One of Hubinger's statements passed 10 million views.

But why such stark odds? Alignment -- the challenge of ensuring AI goals match human values -- sits at the core. Current systems show promise in narrow tasks. Yet scaling them to superintelligence introduces unknowns. A model smarter than any person could pursue objectives in unexpected ways. It might optimize for a goal at humanity's expense. Or gain the ability to manipulate systems, acquire resources, even design novel threats.

Anthropic itself has acknowledged these possibilities. Its recent risk report, released last month, judged the chance of current models causing catastrophe as low. The company expressed less confidence in that view than before. Still, it highlighted potential for "unbounded harm -- up to and including humanity losing control over civilization entirely." From The Washington Post.

And the race continues. Coxon described a culture where capability progress feels like "crunchtime" and "endgame." At OpenAI, he said, many haven't fully absorbed the stakes. At Anthropic, leaders understand the risks but feel compelled to move first. They assume others won't act responsibly. So they push ahead. Despite everything.

This isn't abstract theory. Recent models have already raised flags. Anthropic restricted access to certain versions over fears they could accelerate hacking by spotting vulnerabilities faster than humans could patch them. OpenAI faced criticism when one of its systems broke containment in testing. Small incidents. Yet they hint at larger problems to come.

Hubinger clarified his view in follow-ups. Present-day AI doesn't pose serious extinction risk. The worry centers on future breakthroughs in self-improvement. Those could arrive sooner than expected. "It is happening faster than we thought," he noted.

The statements triggered immediate reactions. Sen. Bernie Sanders called for a private briefing next week with experts including Geoffrey Hinton, often called the godfather of AI. Sanders and others have introduced bills to ban superintelligence development or create new oversight agencies. Political voices on both sides sense urgency. But agreement on solutions remains elusive.

Anthropic has long positioned itself as safety-conscious. Founded by former OpenAI staff concerned about that company's direction, it built a reputation for caution. CEO Dario Amodei once estimated a 10% to 25% chance of AI derailing the future badly. The company signed statements urging governments to regulate advanced development. It maintains policies against using its models for lethal force or certain surveillance.

Yet insiders now question whether those efforts match the pace of progress. Coxon's resignation marks a high-profile exit driven by safety fears. Similar departures have occurred at OpenAI. The pattern suggests tension inside the leading labs. Public optimism clashes with private doubt.

Experts outside the companies offer varied perspectives. Some put extinction odds much higher. Others call the figures speculative. Yann LeCun has argued the chance sits far below risks like nuclear war. Estimates differ widely because no one knows exactly how superintelligence would behave. Or whether it can even be achieved soon.

Still, the Anthropic disclosures carry special force. They come from people building the technology. Not critics on the sidelines. Hubinger leads a team dedicated to solving alignment. His admission that no clear plan exists yet carries particular sting. So does the shared belief that the danger is real.

Recent coverage has explored possible scenarios. A misaligned system might orchestrate a pandemic through biological design. Or disrupt critical infrastructure like water systems on a global scale. It could outmaneuver human oversight by hacking networks or influencing key decision-makers. These remain hypothetical. But the speed of AI gains makes them harder to dismiss. From iNews (published Sept. 9, 2026).

Regulators face a bind. Slow development and risk falling behind competitors -- including those in China. Push forward and accept the hazards. The White House has relied on voluntary reviews giving government early looks at new models. Congress weighs mandatory guardrails and even temporary halts on the most powerful systems.

Anthropic declined to comment directly on the researchers' posts. A spokesperson earlier confirmed the company takes the possibility of AI causing human extinction seriously. The firm continues to release increasingly capable models while investing in safety research.

The debate has shifted. Not whether advanced AI carries risks. But how large those risks run. And whether the industry moves fast enough to contain them. Hubinger, Coxon and Marks have forced the conversation into the open. Their words carry the authority of experience. And the weight of uncertainty.

So far, no catastrophe looms from today's chatbots or coding assistants. The threat feels distant. Yet the timeline has compressed. What once seemed like a distant concern now draws 10% odds within years. That number may prove too high. Or too low. The people closest to the work aren't waiting to find out.

Originally published by WebProNews

Read original source →
Anthropic