Byte by Mahmud logoByteby Mahmud
The AI Too Dangerous to Release Found 129,000 Software Bugs. Anthropic's Answer: Give It to More People.
Technology6 min read6 views

The AI Too Dangerous to Release Found 129,000 Software Bugs. Anthropic's Answer: Give It to More People.

Mahmud Hasan

Mahmud Hasan

October 6, 2026

Anthropic's defensive AI program helped find 129,000 verified software vulnerabilities in three months. On Tuesday, the company merged its two cybersecurity access programs and expanded who gets to use its most powerful models — with fewer safeguards. Here's the honest version of why the lab that called these models too dangerous to release thinks wider access is the safety play.

What Anthropic announced on Tuesday

On October 6, Anthropic announced a revamped Cyber Verification Program, merging two programs it has run for six months: Project Glasswing, which gave organizations defending critical software access to Claude Mythos — Anthropic's most cyber-capable model family — and the original CVP, which gave vetted security teams reduced safeguards on Claude Opus and Sonnet.

The new program has three tiers, and all three get access to Claude Opus 5.5, Sonnet 5.5, Mythos 5.1, and future models. The Defense tier covers incident response and malware analysis; security teams, critical-infrastructure operators, open-source maintainers, and researchers with a record of reported vulnerabilities can apply. The Red Team tier adds authorized penetration testing and red-teaming, but only organizations can apply. The Specialized tier has the fewest restrictions and is reserved for organizations testing safety-critical systems like power grids, flight systems, and interbank transfer infrastructure.

Anthropic vets each member together with the US government, and existing Glasswing members move straight into the Specialized tier. This is not an API key you buy with a credit card.

129,000 vulnerabilities in three months

The evidence Anthropic offered for why this expansion is working: Glasswing partners found at least 129,000 verified vulnerabilities between April and July. Anthropic's own open-source scanning found 5,500 more between April and October. More than 33,000 have been rated critical or high severity. The company says these figures, drawn from a survey of limited partners, are likely an undercount — true impact at least five times higher.

Those numbers need context to land. In April, when Anthropic unveiled Claude Mythos Preview, it scored 83.1% on CyberGym, a benchmark for cybersecurity vulnerability reproduction — against 66.6% for Anthropic's previous flagship. Its red team documented working exploits produced in hours — work expert testers said would take weeks — with the model finding nearly all its flagged vulnerabilities entirely on its own. The UK's AI Security Institute confirmed it independently: two years earlier the best models "could barely complete beginner-level cyber tasks"; Mythos Preview could run multi-stage attacks and autonomously find and exploit vulnerabilities.

Glasswing itself started in April with $100 million in compute credits and a coalition of heavyweights — AWS, Apple, Microsoft, Google, NVIDIA, CrowdStrike, the Linux Foundation, plus roughly 40 more organizations. In June, 150 more organizations across 15+ countries joined — many of them software vendors whose code sits inside other organizations' products. Claude Security, Anthropic's product on this work, uses Mythos 5 to scan code and return findings, with every patch requiring human review before it ships.

Why "fewer safeguards" is the safety argument

Anthropic's position is that it cannot release Mythos-class models publicly — the safeguards needed to release one safely "do not yet exist anywhere," in its own words — and that withholding the capability entirely would still end in the same place. The company's stated expectation: within six to twelve months, rival developers will field models with comparable cyber capabilities, possibly released without safeguards at all.

The logic of the Cyber Verification Program is the dual-use problem, solved by identity instead of refusal. The same request — "explain how this exploit chain works" or "build a realistic attack for testing" — is routine work for a red team and a starting point for a criminal. A content filter can't tell those two apart; vetting the organization can. So the program gives verified defenders reduced safeguards on precisely the dual-use work the default filters block, while prohibited uses like ransomware development stay blocked for everyone, verified or not. Approval is per organization, and the same person working from a personal workspace hits the default safeguards again.

The honest version of the bet: Anthropic thinks the next few years are a race between defenders and attackers armed with the same tools, and the only lever it has is making sure the defenders get there first. Its own summary of where this leaves us is not exactly reassuring: "the transitional period may be tumultuous regardless."

China built its own answer this week

The race framing isn't hypothetical anymore. At the ISC.AI 2026 conference in Beijing this week, 360 Security Technology founder Zhou Hongyi unveiled "Yitian Tulong" — a suite of two AI security tools, named for the legendary Heavenly Sword and Dragon Saber, pitched as a direct answer to Mythos. One component, Tulongfeng, is described as China's version of Mythos for automatically hunting vulnerabilities; 360 claims it has already flagged 3,432 vulnerabilities, with 105 officially confirmed by Chinese authorities. The other, Yitianzhen, is built to automate cyber defense and incident response.

Zhou's framing is worth quoting because it names the fear plainly: he warned of a dangerous "one-way transparency" if US entities could scan global software systems with Mythos-class models while Chinese firms lacked the same capabilities. The US government, for its part, went so far as to block Anthropic from exporting even a scaled-down version of Mythos to foreign destinations and nationals.

Zhou also admitted something revealing: tight US chip export controls have left domestic Chinese models with a self-admitted 20 to 30 percent gap in raw computing power. 360's workaround is organizational, not computational — instead of "the strongest chips and the strongest computing power," it layers security databases, automated tools, and human expertise over its existing models via an AI-agent strategy. "If the US route is to cultivate a genius hacker," Zhou said, "360's route is to organise a professional attack-and-defence team." When the side facing export controls stops trying to out-compute you and starts trying to out-organize you, the race just got harder to handicap.

What it means for everyone else

If you maintain an open-source project with a CVE track record, the Defense tier is open to you — that's the actionable line buried in the announcement. But the broader shift matters even if you never apply. One Glasswing observation that rarely gets airtime: the bottleneck changed. Finding vulnerabilities became easier than verifying, disclosing, patching, and deploying fixes. 129,000 verified bugs in three months is a defensive triumph and an operational nightmare, because every verified-but-unpatched finding is a head start for whoever finds it next.

So the practical takeaways are unglamorous, and that's the point:

  • Treat patch latency as the main security metric, not scanner coverage. AI is now flooding the front end of the vulnerability pipeline; the back end — disclosure, patching, deployment — is where defenses win or lose.
  • Inventory your dependencies before something else does. Anthropic's own scanning covered more than a thousand open-source projects. Assume any internet-facing dependency is being fuzzed by models as good as or better than Mythos right now.
  • Verify claims about "AI-found" vulnerabilities before you panic. Independent assessments found 90.6 percent of a sampled set of Mythos findings were valid true positives — impressive, but not 100. Slop reports are already a crisis in bug bounties; don't let the same flood rot your own triage.

The uncomfortable bottom line: the same week Anthropic handed its most powerful models to more security teams with fewer safeguards, China unveiled its own version, and the US is export-blocking the original. Everyone has accepted that the tools will be in both hands soon. The only open question is whether the patching pipeline of the world's critical software can move faster than the discovery pipeline — and for the first time in this industry's history, the honest answer is that nobody knows.

References

Comments

Leave a comment

Your email stays private — only your name is shown.

More in Technology