Anthropic's Mythos Model Discovers Security Vulnerabilities Faster Than Teams Can Patch
What's new in v2
The HAWK finding has moved from 'mathematical weakness, currently impractical' to 'algorithm broken, developer withdrew from NIST standardization' — a concrete real-world consequence that supersedes the previous synthesis's cautious framing [5]. Matthew Green is a new voice with a cautiously optimistic take on AI cryptanalysis timing during the PQC transition [6]. The vulnerability count has grown from 10,000 to 23,000 across 1,000 OSS projects [2], though the item lacks published claims beyond its title.
What
Anthropic's Claude Mythos model, deployed under Project Glasswing, has broken a NIST post-quantum cryptography candidate: Mythos found a flaw in the HAWK digital signature algorithm that caused its developer to withdraw it from NIST's standardization process after HAWK had passed two evaluation rounds [5]. Separately, the program has reported 23,000 potential vulnerabilities across 1,000 open-source software projects [2], and Microsoft has held emergency engineering meetings as Mythos discovers flaws faster than teams can patch them [1]. Cryptographers are now debating whether AI-assisted cryptanalysis during the PQC transition period is net positive or a dual-use liability [6][1].
Why it matters
The HAWK withdrawal shows that AI-assisted cryptanalysis can produce concrete consequences for security standards — a NIST PQC candidate that survived two human-led evaluation rounds was broken and removed from contention before widespread deployment. The dual-use concern runs through every part of this: the same capability that caught a flaw in a candidate standard could be turned toward breaking standards already in use.
Open questions
How will NIST and PQC evaluators adjust processes now that AI can break a candidate that survived two rounds of expert human review? [5]
Matthew Green argues AI cryptanalysis during PQC standardization is net positive [6] — does the cryptography community broadly agree, or are there substantive dissenting views?
The vulnerability count has grown from 10,000 [3] to 23,000 across 1,000 OSS projects [2] — are Microsoft and other Project Glasswing partners keeping pace with patching, or is the backlog growing?
Microsoft engineers internally questioned whether Mythos matched Anthropic's stated capabilities [1] — has the HAWK withdrawal changed assessments inside partner organizations?
Narrative
Anthropic's Claude Mythos model has been deployed under Project Glasswing to find software vulnerabilities in critical systems at a rate that has exceeded human patching capacity. Microsoft was among the first partners; the company convened emergency engineering meetings to handle vulnerabilities Mythos was uncovering faster than its teams could address [1]. Anthropic frames the program as a race to find and fix flaws before adversarial governments — China is specifically named — deploy equivalent AI tools [1]. By mid-2026, Mythos had reportedly identified over 23,000 potential vulnerabilities across 1,000 open-source software projects [2], growth from the 10,000 reported in May [3].
In parallel, Anthropic researchers ran Mythos through a focused 60-hour cryptographic research session at approximately $100,000 in API costs, targeting mathematical weaknesses in the HAWK post-quantum signature standard and a reduced-round version of AES-128 [4]. A distinctive feature of the session was that human interventions were primarily motivational rather than technical: researchers repeatedly prompted Mythos not to conclude problems were unsolvable, urging it to find something publishable [4]. The model produced genuine mathematical analysis, finding weaknesses that had not surfaced during years of expert evaluation.
The HAWK findings proved more consequential than initially characterized. HAWK had passed two rounds of NIST's post-quantum cryptography evaluation before Mythos identified a flaw severe enough to render the algorithm broken [5]. Following Anthropic's public disclosure, HAWK's developer withdrew the algorithm from NIST's standardization process entirely [5]. The result is concrete: a candidate standard designed to secure communications against future quantum computers was removed from consideration by AI-assisted cryptanalysis.
The results have drawn reactions from the security and cryptography communities. Matthew Green, quoted by Simon Willison, argues the timing is close to ideal — the PQC transition period, before any of these algorithms are widely deployed, is exactly when catching weaknesses is most valuable, and AI cryptanalysis arriving now is net positive for defenders [6]. That view sits alongside the investigative reporting's dual-use concern: the same capability that helped remove a flawed standard also tells adversarial programs what to build [1]. Internally, at least one Microsoft engineer questioned whether Mythos lived up to Anthropic's capability claims [1], a gap between the company's public cooperation with Project Glasswing and private assessments of the model's performance that remains unresolved.
Timeline
- 2026-04: Microsoft MSRC publishes a blog describing how it is evolving security processes for AI-driven vulnerability discovery. [9]
- 2026-05-06: Barracuda Networks publishes a CISO-level analysis on how Mythos will change the vulnerability discovery landscape. [11]
- 2026-05-26: Anthropic's Project Glasswing update reports Mythos has identified over 10,000 software flaws. [3]
- 2026-07-28: Anthropic publishes research showing Mythos found mathematical weaknesses in HAWK and reduced-round AES-128 in a 60-hour, ~$100,000 session requiring primarily motivational rather than technical guidance. [4]
- 2026-07-29: ProPublica and Ars Technica report that Microsoft held emergency engineering meetings under Project Glasswing as Mythos discovered bugs faster than teams could patch them, with internal skepticism about Mythos's claimed capabilities. [1][10]
- 2026-07-29: Ars Technica (Dan Goodin) reports Mythos found a flaw that rendered HAWK broken; HAWK's developer withdrew the algorithm from NIST's PQC standardization process. [5]
- 2026-07-29: Cryptographer Matthew Green, quoted by Simon Willison, argues AI cryptanalysis coming online during the PQC standardization period is net positive for security. [6]
- 2026-07: SecurityWeek reports Mythos has detected 23,000 potential vulnerabilities across 1,000 open-source software projects. [2]
Perspectives
Anthropic
Deploying Mythos as a defensive tool under Project Glasswing, framing the program as a race to find and fix vulnerabilities before adversarial governments deploy equivalent AI tools.
Evolution: Consistent throughout the thread; no public acknowledgment of the patching-speed tension Microsoft's situation revealed, or of internal partner skepticism.
Matthew Green (cryptographer)
Cautiously optimistic: argues the PQC transition period is the best possible time for AI cryptanalysis to come online, because catching weaknesses in candidate standards before widespread deployment is the ideal outcome.
Evolution: New voice in this pass; single-point commentary with no prior stance to compare.
Simon Willison
Impressed by the cryptographic research results; identifies the motivational prompting dynamic as the most revealing aspect of how Mythos actually operates; quotes Green approvingly on the PQC timing argument.
Evolution: Consistent; expanded to include endorsement of Green's framing.
Microsoft
Publicly cooperative via the MSRC blog; internally under strain, with emergency meetings to handle discovery volume and at least one engineer questioning whether Mythos matched Anthropic's capability claims.
Evolution: Public posture aligns with Anthropic's framing; internal reporting reveals skepticism and operational pressure not visible in official communications.
ProPublica / Ars Technica
Investigative: the story is that discovery pace has outrun patching capacity, and that demonstrating Mythos's capabilities publicly informs adversarial governments about what to build.
Evolution: Ars Technica's Goodin added a concrete data point — HAWK's withdrawal — to the investigative framing; consistent dual-use concern throughout.
Security industry analysts (Barracuda, Dynatrace, Cloud Security Alliance)
Treating Mythos as a meaningful operational development; focused on what AI-scale vulnerability discovery means for organizational security processes rather than the dual-use question.
Evolution: Uniformly analytical; no strong dissent from the broader framing.
Tensions
- Mythos discovers vulnerabilities faster than Microsoft's engineering teams can patch them, creating an accumulating backlog of known-but-unpatched exposure. [1]
- Microsoft internally questions whether Mythos matched Anthropic's stated capabilities, while Anthropic publicly frames Project Glasswing as a clear success. [1]
- Matthew Green argues AI cryptanalysis entering during PQC standardization is net positive because it catches weaknesses before deployment [6]; ProPublica and Ars Technica treat the same dual-use capability as a risk that benefits adversaries as much as defenders [1]. [6][1]
- Anthropic frames Project Glasswing as protecting against adversarial exploitation, but each public disclosure about Mythos's cryptanalytic capabilities tells adversarial programs exactly what tools they should build. [1][5]
- The $100,000 cryptographic research session was initially characterized as producing impractical findings, but the HAWK withdrawal shows the results were consequential enough to remove a NIST candidate from consideration. [4][5]
Status: active and growing
Sources
- [1] Anthropic is finding bugs faster than Microsoft can fix them — Ars Technica AI (2026-07-29)
- [2] Anthropic: Mythos Detected 23,000 Potential Vulnerabilities Across 1,000 OSS Projects - SecurityWeek — reactive:anthropic-mythos-vulnerability-discovery
- [3] Anthropic: Claude Mythos identified 10,000+ software flaws - Help Net Security — reactive:anthropic-mythos-vulnerability-discovery
- [4] Discovering cryptographic weaknesses with Claude — Simon Willison (2026-07-28)
- [5] Mythos attack on 3rd-round PQC algorithm candidate puts it out of commission — Ars Technica AI (2026-07-29)
- [6] Quoting Matthew Green — Simon Willison (2026-07-29)
- [7] Project Glasswing: An initial update — reactive:anthropic-mythos-vulnerability-discovery
- [8] Anthropic's coordinated vulnerability disclosure dashboard — reactive:anthropic-mythos-vulnerability-discovery
- [9] Strengthening secure software at global scale: How MSRC is evolving with AI — reactive:anthropic-mythos-vulnerability-discovery
- [10] Microsoft Struggling with AI-Discovered Security Bugs — reactive:anthropic-mythos-vulnerability-discovery (2026-07-29)
- [11] From the desk of the CISO: How will Anthropic’s Mythos change vulnerability discovery? | Barracuda Networks Blog — reactive:anthropic-mythos-vulnerability-discovery
- [12] Anthropic Claude Mythos is reshaping the vulnerability landscape — reactive:anthropic-mythos-vulnerability-discovery
- [13] Claude Mythos: AI Vulnerability Discovery and Containment Failures — reactive:frontier-ai-cyber-capabilities