
Anthropic just took its most potent cybersecurity model out of the vault. Not for everyone. Not even close. On August 21, the company announced that Claude Mythos 5 would now power scans inside its Claude Security tool for enterprise customers. The model also heads into partner products. Defenders get findings, severity scores, confidence ratings, suggested fixes. They do not get the model itself.
A user receives an artifact. A patch proposal. An alert. No prompt access. No chance to ask it to craft an exploit. The design choice sits at the heart of Anthropic's approach. Anthropic's own blog post puts it plainly: outputs only. Humans remain firmly in the loop. Every patch requires review and approval before deployment. The scan stays narrow. It does not unlock broader model capabilities.
This move builds directly on Project Glasswing, launched in April. That effort brought together AWS, Apple, Google, Microsoft, Cisco, CrowdStrike, NVIDIA, Palo Alto Networks, JPMorgan Chase, the Linux Foundation and others. They received early access to Mythos Preview, the predecessor. Those partners scanned codebases and uncovered more than 10,000 high or critical vulnerabilities in a matter of weeks. The Next Web first reported the scale of those finds. Patching lagged far behind discovery. The bottleneck became obvious.
So Anthropic is putting money behind the fix. The new Defender Advantage Fund, dubbed 0xDAF, commits $35 million in Claude credits. Grants target three priorities. Patching live flaws in widely used projects. Automating scan-and-patch pipelines that others can adopt. Exploring architectures that resist entire categories of attacks. The fund follows earlier commitments. Project Glasswing already delivered up to $100 million in usage credits plus $4 million in direct donations to groups including OpenSSF, Alpha-Omega and the Apache Software Foundation.
Timing carries weight. European open-source maintainers face new obligations under the Cyber Resilience Act. Vulnerability reporting rules kick in September 11. Stewards must maintain cybersecurity policies, disclose actively exploited flaws and cooperate with authorities. They receive exemption from some penalties. Yet the burden lands on volunteer-driven projects that often lack resources. Credits for model usage help those who already have people to run the scans. They do not create new maintainers.
The caution traces back to real incidents. In July Anthropic disclosed that three of its models had reached real organizations during misconfigured evaluations. The episodes reinforced the company's preference for controlled outputs over open prompts. Earlier this year the U.S. government briefly imposed export controls on Mythos 5 and its sibling Fable 5, citing national security. Access was restored in late June to a limited set of trusted U.S. organizations after negotiations involving Commerce Secretary Howard Lutnick. NBC News covered the reversal.
OpenAI follows a parallel path. The company operates its own vetted access program for security teams. Both labs now gate their strongest cyber capabilities behind verification processes. The pattern suggests an emerging norm. Frontier models with dual-use potential stay restricted. Defenders gain tools built on top of them. Direct interaction remains off limits for most.
Independent tests have tempered some claims. One analysis found that several headline vulnerabilities spotted by Mythos were also caught by much smaller open-source models. A Wikipedia entry on Claude Mythos notes that an open-weight model with only 3.6 billion active parameters identified the same flaw at a fraction of the cost. The Wikipedia page summarizes these findings. Yet the gap in scale, autonomy and exploit chaining still favors the frontier systems. Anthropic's Logan Graham, who leads the frontier red team, described Mythos Preview as capable of finding tens of thousands of vulnerabilities that even skilled human researchers would miss.
The economic reality sharpens the stakes. Successful exploit development runs using Mythos-class models have come in under $2,000 in some tests. The FBI has flagged the shift as a law enforcement challenge. TechTimes reported the agency's assessment. Attackers no longer need large teams or long timelines. The same technology that accelerates defense also lowers the bar for offense. This asymmetry explains why Anthropic insists on containment.
Critics argue the restrictions create new power centers. A small number of companies and government partners decide who qualifies as trusted. Lobbying for access has become part of the game. During the June export control episode, executives and researchers flew to Washington. Dario Amodei's team engaged directly with officials. The episode highlighted how political risk now factors into AI business planning. Alex Stamos, a prominent security voice, told The Verge that capabilities had been somewhat overstated but that the regulatory precedent mattered more. The Verge examined the internal debates.
Yet the direction looks set. On August 21, Unite.AI noted that the Cyber Verification Program will soon expand to cover broader dual-use capabilities on current models, with Mythos-class access to follow for vetted defenders. Unite.AI detailed the four concrete changes. The New Stack emphasized that because scans run inside Anthropic's controlled environment, the company can limit exposure. Users get results. They never touch the raw model. The New Stack broke down the safeguards.
Discussions on X reflect the tension. One post called the development a shift from benchmarks to infrastructure questions. Who gets access. Under what monitoring. With what containment. Another highlighted the $35 million fund as recognition that open-source sits at the foundation of critical systems yet often runs on volunteer labor. The conversation has moved past hype. It now centers on liability, triage capacity and the speed mismatch between discovery and remediation.
Anthropic's strategy accepts an uncomfortable truth. The models are coming. Competitors will match or exceed current capabilities. Withholding entirely cedes ground to attackers who face fewer constraints. Controlled deployment to defenders, paired with funding for the open-source base, offers a pragmatic middle path. It is imperfect. Patching still trails discovery. Credits cannot replace dedicated maintainers. Export controls and verification programs create gatekeepers.
But the alternative looks worse. Unrestricted access risks handing sophisticated exploit generation to anyone with an API key. The incidents in July showed how quickly misconfiguration can expose sensitive environments. So the company keeps the model behind the curtain. It ships the intelligence instead. Suggested fixes. Tagged vulnerabilities. Human-reviewed patches. And $35 million to help the volunteers who keep the internet's plumbing intact.
The test will come in the months ahead. Whether the Defender Advantage Fund actually accelerates patching at scale. Whether partner integrations deliver usable tools without creating new attack surfaces. Whether European maintainers, facing fresh regulatory deadlines, can turn credits into concrete improvements before the next wave of disclosures. Anthropic has bet that outputs plus resources beat open access. The industry is watching to see if the bet holds.