News · Policy · Global
Anthropic's AI Found Thousands of Software Flaws. Fixing Them Is Another Matter.
Mythos, the name given by Anthropic to its highly capable, unreleased cybersecurity-focused frontier model because the word represents stepping "beyond the boundaries of the known," has proved exceptionally good at…
Mythos, the name given by Anthropic to its highly capable, unreleased cybersecurity-focused frontier model because the word represents stepping "beyond the boundaries of the known," has proved exceptionally good at hunting vulnerabilities in critical software. The company is keeping it locked away.
Open-source software maintainers typically don't get many complaints. Most work nights and weekends, fielding the occasional bug report, keeping code running for millions of users who don't know their names.
Then Anthropic started pointing an AI at their repositories.
Over the past several months, the company has sent more than 1,100 vulnerability reports to the volunteers who maintain some of the internet's foundational open-source projects, Anthropic said.
The bugs were real. The patches were not materializing fast enough. Some maintainers asked Anthropic to stop sending them.
The AI responsible is called Claude Mythos Preview, and Anthropic is not making it available to the public. Last month, the company published results from the first month of Project Glasswing, an effort to use Mythos to find and repair critical software flaws before AI-assisted attacks become routine. What Anthropic and roughly 50 partners found in those weeks was enough to prompt an unusual disclosure: the company said one of its own AI models is too dangerous to release.
"No company — including Anthropic — has developed safeguards strong enough to prevent such models from being misused," Anthropic said in its Project Glasswing update.
Across all of Glasswing's partners and its separate open-source scanning effort, Mythos found what Anthropic says are more than 10,000 vulnerabilities rated high- or critical-severity, an aggregate that the company acknowledged is still growing. The challenge it now faces is that the security industry was not built to process findings at that speed.
Cloudflare, one of roughly 50 Glasswing partners, reported 2,000 bugs across its critical-path systems, 400 rated high- or critical-severity. The false positive rate, Cloudflare said, was better than human testers. Mozilla found and fixed 271 vulnerabilities in Firefox 150 while testing the model, more than ten times what it found in Firefox 148 using Claude Opus 4.6, an earlier Anthropic model available to the public.

Some of what Mythos found wasn't just theoretically exploitable.
In wolfSSL, an open-source cryptography library used by billions of devices, the model built a working exploit that would let an attacker forge digital certificates, making a fake bank website or email service appear genuine to any ordinary user. The flaw had sat in code its developers marketed for its security. It was assigned CVE-2026-5194 and has since been patched.
At one Glasswing partner bank whose name Anthropic did not publish, the model flagged and helped stop a fraudulent $1.5 million wire transfer. A threat actor had compromised a customer's email account and made spoofed phone calls to complete the scheme. Mythos caught it.
At one Glasswing partner bank whose name Anthropic did not publish, the model flagged and helped stop a fraudulent $1.5 million wire transfer. A threat actor had compromised a customer's email account and made spoofed phone calls to complete the scheme. Mythos caught it.
External tests tracked similar results.
Britain's AI Security Institute said Mythos Preview is the first model to complete both of its cyber ranges, which simulate multi-step attacks, from start to finish. Security platform XBOW called it a "significant step up over all existing models" on its web exploit benchmark. Two new academic benchmarks, ExploitBench and ExploitGym, rank Mythos first on exploit development. Anthropic's Frontier Red Team blog goes into the methodology.
The bottleneck, it turns out, is not finding the bugs.
For decades, vulnerability discovery set a natural ceiling on how fast hackers and defenders could move. That ceiling is gone. Palo Alto Networks put out more than five times its normal number of patches in its most recent update. Microsoftsaid its patch volumes will "continue trending larger for some time." Oracle is fixing flaws across its products multiple times faster than before.
The open-source side is harder.
Anthropic scanned more than 1,000 projects with Mythos Preview and identified an estimated 6,202 high- or critical-severity vulnerabilities out of 23,019 total findings. Of roughly 1,752 independently reviewed by outside security firms, 90.6 percent checked out as genuine. Nearly two-thirds were rated high- or critical-severity. On average, a confirmed bug takes two weeks to patch. Once a flaw is disclosed, that two-week window is open to anyone.
"The relative ease of finding vulnerabilities compared with the difficulty of fixing them amounts to a major challenge for cybersecurity," Anthropic said.
"The relative ease of finding vulnerabilities compared with the difficulty of fixing them amounts to a major challenge for cybersecurity," Anthropic said.
The company's partial response has been to share what it learned without sharing Mythos itself. Three weeks ago, Anthropic launched Claude Security in public beta for enterprise customers, powered by Claude Opus 4.7, a less capable but publicly available model. The tool scans codebases and proposes fixes. In three weeks, it patched more than 2,100 vulnerabilities. The company has also released to qualifying enterprise security teams the internal tooling it and its Glasswing partners used with Mythos Preview: custom skill sets, a codebase-mapping harness, and a threat model builder that prioritizes targets by likely attack surface.
A Cyber Verification Program now lets credentialed penetration testers and vulnerability researchers use publicly available Anthropic models for legitimate security work without triggering the misuse safeguards. The company partnered with the Open Source Security Foundation's Alpha-Omega project to help maintainers work through the incoming queue of reports. Cisco open-sourced its Foundry Security Spec, giving other defenders access to its internal evaluation framework.
Glasswing's vulnerability data is now public through a dashboard on Anthropic's red team site that tracks each stage of the disclosure process: identified, triaged, reported, patched. The chart shows a steep drop-off at each step. That is where the people run out of hours.
Mythos Preview itself stays off-limits. Anthropic says general access will come only after the company develops "far stronger safeguards," with no timeline given. The company has also warned that comparable models are coming from other companies regardless. If one of those companies releases without adequate protections, the window Glasswing is trying to exploit — a period when defenders hold an advantage over attackers — closes.
Next, Anthropic plans to expand Glasswing to additional partners, including United States and allied governments. On Friday, Japanese Finance Minister Satsuki Katayama told reporters that U.S. Treasury Secretary Scott Bessent had offered access to the model for use by the Japanese government and its companies within two weeks starting May 12, according to The Jiji Press, Ltd., a a news agency in Japan.
Back in the open-source world, the maintainers who asked Anthropic to slow down are still working through their queues.
The author is the head of Research and Analysis at Icarus Asia, a Hong Kong-based risk and advisory business.
Further Reading & Citations
Primary Sources -- Anthropic
Assessing Claude Mythos Preview's Cybersecurity Capabilities — Nicholas Carlini, Newton Cheng, Keane Lucas et al., Anthropic Frontier Red Team, April 7, 2026. Technical deep-dive on Mythos Preview's zero-day discovery, exploit construction, and autonomous attack capabilities across major operating systems and browsers.
Project Glasswing: One Month Update — Anthropic, May 2026. Aggregate results across 50 partners and open-source scanning; the primary source for the 10,000+ vulnerability figure.
Exploit Evaluation Benchmarks — Anthropic Frontier Red Team, 2026. Methodology behind ExploitBench and ExploitGym performance assessments.
Open-Source Vulnerability Disclosure Dashboard — Anthropic, 2026. Live tracker of triage, disclosure, and patch status across scanned open-source repositories.
Coordinated Vulnerability Disclosure Policy — Anthropic. The 90-day disclosure framework governing Glasswing's reporting process.
Claude Security (Public Beta) — Anthropic, 2026. Enterprise codebase scanning tool powered by Claude Opus 4.7.
Cyber Verification Program — Anthropic Support, 2026. Program allowing credentialed security professionals to use Anthropic models without standard misuse safeguards.
Partner & External Reports
Cloudflare: Cyber Frontier Models — Cloudflare, 2026. Details on 2,000 bugs found across critical-path systems.
Mozilla: AI Security Zero-Day Vulnerabilities — Mozilla Security Blog, May 2026. The 271-vulnerability Firefox 150 finding.
Behind the Scenes: Hardening Firefox — Mozilla Hacks, May 2026.
wolfSSL: How Claude Mythos Preview Helped Harden wolfSSL — wolfSSL, 2026. Technical narrative on CVE-2026-5194 and the certificate-forging exploit.
XBOW: Mythos Offensive Security Evaluation — XBOW, 2026.
UK AI Security Institute: How Fast Is Autonomous AI Cyber Capability Advancing? — AISI, 2026.
Palo Alto Networks: Defenders Guide — Frontier AI Impact on Cybersecurity — Palo Alto Networks, May 2026.
Microsoft MSRC: A Note on Patch Tuesday — Microsoft Security Response Center, May 2026.
Oracle: Accelerating Vulnerability Detection and Response — Oracle Security Blog, 2026.
Cisco: Announcing the Foundry Security Spec — Cisco, 2026.
Benchmarks & Academic Research
ExploitBench — Independent benchmark for measuring AI exploit development capabilities.
ExploitGym — arXiv, 2026. Academic benchmark for exploit development; Mythos Preview ranked first.
Ecosystem & Governance
Linux Foundation / OpenSSF Alpha-Omega: $12.5M Grant Announcement — OpenSSF, March 2026.
NIST Cybersecurity Framework — National Institute of Standards and Technology.
NCSC: 10 Steps to Cyber Security — UK National Cyber Security Centre.
Classics
The Poetics of Aristotle — Aristotle (c. 335 BCE), trans. S. H. Butcher. Project Gutenberg, public domain. The foundational Western text on mimesis, narrative structure, and the mechanics of storytelling.
Mythos: The Greek Myths Reimagined — Stephen Fry. Chronicle Books, 2019. A modern retelling of the Greek myths -- and the source of the name Anthropic chose for its most capable model to date. (Editor's note - Anthropic named a model after the body of myth that Fry spent a book reimagining, and now that model is finding security holes in the software that runs the modern world.)