Project Glasswing: When a Model Finds Too Many Bugs
Two weeks ago, Nicholas Carlini showed that Claude Code could find a 23-year-old Linux kernel vulnerability using a 10-line bash script. Today, Anthropic announced what happened when they scaled that approach: a new model called Claude Mythos that found thousands of high and critical-severity zero-day vulnerabilities, then decided not to release it publicly.1
Project Glasswing is Anthropic’s restricted deployment of Claude Mythos, a frontier model that discovered thousands of zero-day vulnerabilities across every major operating system and web browser. Mythos found critical bugs including a 27-year-old OpenBSD TCP SACK flaw and a FreeBSD NFS remote code execution vulnerability. Anthropic restricted access to 12 partner organizations for defensive security only, committed $100M in usage credits, and opened the Cyber Verification Program application form at claude.com/form/cyber-use-case for qualified researchers.
Project Glasswing is Anthropic’s answer to the question practitioners have been asking since Carlini’s [un]prompted talk: what happens when this capability is deployed at scale? The answer: you restrict it.
TL;DR
Claude Mythos Preview is a frontier model whose cybersecurity capabilities, according to Anthropic, “emerged as a downstream consequence of general improvements in code, reasoning, and autonomy.”1 Anthropic positions it as more cyber-capable than any generally available Opus model (including the April 16, 2026 Opus 4.7 release), and restricts access to 12 partner organizations (Apple, Amazon, Microsoft, Google, Linux Foundation, and others) for defensive security work only. The model found thousands of zero-days, including a 27-year-old OpenBSD TCP SACK bug, a 16-year-old FFmpeg vulnerability, and a FreeBSD NFS RCE (CVE-2026-4747).1 Anthropic committed $100M in usage credits and $4M to open-source security organizations. The Cyber Verification Program application form is now live for legitimate security researchers seeking access.1 The October 1, 2026 update below covers GLM-5.3, an open-weight model that Anthropic reports developing end-to-end exploits for known V8 bugs at about Mythos Preview’s rate on ExploitBench.7
Key Takeaways
- Security engineers: The capability threshold that Carlini demonstrated at [un]prompted is real, and it scales. Mythos found vulnerabilities in “every major operating system and web browser.”2 Defensive security teams at the 12 partner organizations now have access. Everyone else should be preparing for what comes when these capabilities reach generally available models. (By September 2026 Anthropic reported an open-weight model at about Mythos Preview’s ExploitBench rate, 50 against 56 of 410 attempts; see the October 1 update.)
- Scaffold builders: Mythos runs via Claude Code in isolated containers.1 The scaffold pattern (agent CLI + sandboxed execution + automated triage) now serves as the production architecture for frontier security research at Anthropic itself. The orchestration patterns practitioners built independently hold up at the highest level.
- Everyone else: Anthropic chose restriction over release. That is a real governance decision with real tradeoffs. The model exists. Anthropic demonstrated the capabilities. The question is no longer whether AI can find zero-days but who gets access and under what constraints.
Update (October 1, 2026): The Capability Reached Open Weights
In Anthropic’s own retelling, Glasswing was a bet on a head start: restrict the model, let defenders use it first, and expect the capability to spread anyway. On September 29 Anthropic’s Frontier Red Team said the spread has happened: “those models have now arrived.”7 GLM-5.3, an open-weight model from Zhipu AI (Z.ai), “develops end-to-end exploits in 50 of 410 attempts” on ExploitBench, a set of known V8 vulnerabilities, where Claude Mythos Preview “did so at a similar rate,” 56 of 410. On 100 randomly selected tasks from Anthropic’s internal binary-exploitation benchmark it reached a full control-flow hijack in 4% of trials against 6% for Mythos Preview, and “earlier models, like Claude Opus 4.6 and GLM-5.2, do not succeed in any of them.” NIST’s CAISI, in its own assessment, called GLM-5.3 “the most cyber-capable open-weight model released to date” on September 17, while finding its capabilities “significantly lower than those of current U.S. frontier models” and “about four months” behind them; by CAISI’s notes, that frontier includes trusted-access releases as well as public ones, and the U.S. models were tested with cyber safeguards disabled where applicable.8
The part that bears on this post is the safeguards. The exploit-development tasks above did not trigger the released model’s refusals at all; the numbers that follow cover overtly malicious requests. In Anthropic’s simulated tests GLM-5.3 refused every direct, overtly malicious request; a cover story got it to engage, meaning try to connect to the simulated target, 64% of the time, prefilled reasoning 92%, and an abliterated copy, one edited to stop refusing, 100%. That edit, which cut refusals on three public benchmarks from above 90% to between 2% and 12%, took Anthropic’s team, which had not done it before, “about 2,200 GPU hours at a computation cost of roughly $4,400”; Anthropic estimates an experienced team would need about $1,200, and notes that abliterated copies were public within days of release. The weights are public: “anyone can download GLM-5.3.” One session shows the price. A researcher gave the smaller GLM-5.3-Flash the public details of a recently disclosed Chrome flaw and one other, and it “chained together exploits for these two flaws, building a reliable exploit chain for an ARM64 target,” in “20 minutes of human attention, plus eight hours of work” by the model. Anthropic’s price for that: “At Zhipu’s API prices, this effort would have cost $20.40.”7
What this does to the argument below. The restriction bought a head start, not exclusivity. Anthropic’s words: “a critical threshold in freely accessible capabilities has now been crossed.” The practitioner advice gets more urgent rather than different. The triage layer, the disclosure workflow, and the patch cadence carry the weight now, because exploit development at about Mythos Preview’s ExploitBench rate, and, in a separate researcher session with the full GLM-5.3, the discovery and chaining of previously unknown vulnerabilities in a web browser, now come as a download. Opus 4.6’s bug-finding results below still stand; on the same 100 randomly selected tasks from Anthropic’s binary-exploitation benchmark it reached no full control-flow hijack. The same post notes that vetted defenders “can now use even more advanced models like Claude Mythos 5.1 through our trusted access programs,” so restricted access has moved on as well. Whether those programs are the application route this post describes, Anthropic’s post does not say.7 The model names in the sections below are April’s; read them with that date.
Update (April 19, 2026)
Since this post went live on April 7, two things changed:
- Opus 4.7 shipped on April 16, 2026 as the new generally-available flagship. Anthropic states that Opus 4.7 is deliberately less cyber-capable than Mythos Preview and ships with real-time cyber safeguards. Mythos Preview remains separate and restricted.5
- The Cyber Verification Program application form is now live at claude.com/form/cyber-use-case. What the original announcement called a “future” program is now a concrete application path.5
- Claude Code shipped two relevant infrastructure releases: v2.1.111 added Opus 4.7 / xhigh / Auto Mode support; v2.1.113 added
sandbox.network.deniedDomains, wrapper-command deny rules (env / sudo / watch / ionice / setsid), stricterfind -exec/-deletehandling, and macOS/private/{etc,var,tmp,home}removal protection underBash(rm:*).6 These are exactly the kind of hardening primitives a Mythos-style security research scaffold needs.
The core argument below, capability restriction over release, scaffold patterns holding up at the highest level, everyone else preparing for what comes when these reach GA, is unchanged. If anything, Opus 4.7’s explicit cyber-safeguard framing strengthens it.
From Talk to Product
Carlini’s [un]prompted talk in early April was the public preview.3 He showed five Linux kernel vulnerabilities and 22 Firefox CVEs found with a simple file-iteration script. The bottleneck, he said, was human validation, “several hundred crashes I haven’t validated yet.”
Mythos is what happens when you remove that bottleneck with a more capable model and dedicated infrastructure. The scale difference is significant:1
| Metric | Carlini’s talk | Project Glasswing |
|---|---|---|
| Vulnerabilities found | 5 kernel + 22 Firefox CVEs | Thousands across all major platforms |
| Targets | Linux kernel, Firefox | Every major OS, browser, open-source project |
| Validation | Manual, researcher-driven | Professional security contractors, 89% severity confirmation |
| Access | Opus 4.6 at the time of Carlini’s talk; Opus 4.7 is now the GA flagship | Mythos Preview (restricted to 12 partners) |
The professional validation number matters: 89% of 198 reviewed reports had severity assessments confirmed by independent security contractors, with 98% within one severity level.1 These are not hallucinated findings.
The Restriction Decision
Anthropic’s stated position: “We do not plan to make Claude Mythos Preview generally available due to its cybersecurity capabilities.”4
The decision stands out. Model companies typically race to ship capabilities. Anthropic built a model that is demonstrably better at finding vulnerabilities than any publicly available system, then chose to restrict it to defensive use by vetted partners. The $100M commitment in usage credits signals this is not a marketing exercise.1
The restriction model has three tiers:1 1. Project Glasswing partners (12 organizations): Direct access for defensive security 2. Broader access (40 organizations total): Supervised deployment 3. Cyber Verification Program (now live at claude.com/form/cyber-use-case): Application path for verified security professionals5
For practitioners, the standard API and Claude Code do not expose Mythos’s vulnerability-finding capabilities. The strongest generally available model is now Opus 4.7 (launched April 16, 2026), which Anthropic positions as deliberately less cyber-capable than Mythos and ships with real-time cyber safeguards.5 Mythos’s demonstrated capabilities already influenced that April 16 release, Opus 4.7 is Anthropic’s first post-Glasswing model with dedicated cyber safeguards.
What This Validates
Project Glasswing validates several patterns that the practitioner community built independently:
Claude Code as the execution scaffold. Mythos runs via Claude Code in isolated containers.1 The same agent CLI that practitioners use for daily coding serves as the execution layer for frontier security research. The hooks, skills, and sandboxing that Claude Code provides are not convenience features. They are the infrastructure that makes autonomous security scanning safe enough to deploy.
The verification bottleneck is an orchestration problem. Carlini’s talk identified human validation as the bottleneck. Project Glasswing’s solution: professional security contractors for validation, SHA-3 hash commitments for responsible disclosure, and structured triage infrastructure.1 The same triage problem surfaced in When Your Agent Finds a Vulnerability, and the solution is infrastructure, not model capability.
Governance hooks matter more than scanning capability. The model can find the vulnerabilities. The hard problem is controlling disclosure, managing access, and ensuring findings reach defenders before attackers. Anthropic’s answer is organizational (restrict the model, vet the partners, commit resources). For practitioners building their own security scanning, the governance hooks that gate output are the equivalent.
What This Means for Practitioners
You are not getting Mythos access. Here is what you can do with what you have:
Opus 4.6 is already capable. Carlini’s [un]prompted results (5 kernel bugs, 22 Firefox CVEs) used Opus 4.6, not Mythos.3 The capture-the-flag methodology, ASAN-instrumented builds, and file-iteration script are all reproducible with the generally available model.
Build the triage layer now. When future Opus models inherit some of Mythos’s capabilities (as Anthropic has implied), the bottleneck will be the same one Carlini identified: human validation. The teams that have automated deduplication, severity classification, and disclosure workflows ready will benefit first.
Apply to the Cyber Verification Program. The application form is live at claude.com/form/cyber-use-case. If you do legitimate security research, this is the path to elevated access.
The trajectory is clear: AI-assisted vulnerability discovery is real, it scales, and the governance question is now the central problem. The model capability is solved. The scaffold that orchestrates discovery, triage, and responsible disclosure is not.
Sources
Frequently Asked Questions
Can I use Claude Mythos through Claude Code?
No. Mythos Preview is restricted to Project Glasswing partners. Opus 4.7 (April 16, 2026) is the strongest model available through Claude Code for general users; Anthropic states Mythos remains more cyber-capable than any GA model. That was the lineup in April 2026. By September 29, Anthropic said vetted defenders could use Claude Mythos 5.1 through its trusted access programs; see the October 1 update above.
Will Mythos capabilities come to Opus?
Opus 4.7 is Anthropic’s first post-Glasswing Opus release and ships with real-time cyber safeguards. The pattern suggests future Opus models will carry additional safeguards rather than the full Mythos capability envelope. Anthropic’s original announcement said they aim to “enable safer deployment through new safeguards in future Claude Opus models.”
How does this relate to the earlier vulnerability blog post?
Carlini’s [un]prompted talk (covered in When Your Agent Finds a Vulnerability) used Opus 4.6 and found 5 kernel bugs + 22 Firefox CVEs. Mythos scaled that approach to thousands of vulnerabilities across all major platforms. The methodology is the same; the model is more capable.
-
Claude Mythos Preview, Project Glasswing. Anthropic, April 7, 2026. Official announcement. Thousands of high/critical-severity zero-days found. 89% severity confirmation rate by professional validators. $100M in usage credits. Led by Nicholas Carlini with 21+ co-authors. ↩↩↩↩↩↩↩↩↩↩↩
-
Anthropic’s Project Glasswing. Simon Willison, April 7, 2026. Analysis and context on the restricted release model and Carlini’s earlier work. ↩
-
Nicholas Carlini, “Black-hat LLMs,” [un]prompted AI security conference, April 2026. Conference agenda. See also: AI Finds Vulns You Can’t, Security Cryptography Whatever podcast. ↩↩
-
Anthropic says its most powerful AI cyber model is too dangerous to release publicly. VentureBeat, April 7, 2026. ↩
-
Post-publication updates (April 19, 2026). Anthropic’s Introducing Claude Opus 4.7 announcement (April 16, 2026) positions Opus 4.7 as the GA flagship while noting Mythos Preview remains more cyber-capable. Real-time cyber safeguard details at Anthropic Support: Real-time cyber safeguards on Claude. Cyber Verification Program application form live at claude.com/form/cyber-use-case. ↩↩↩↩
-
Claude Code CHANGELOG. v2.1.111 added Opus 4.7 launch support (xhigh effort, Auto Mode for Max without flag). v2.1.113 added
sandbox.network.deniedDomains, wrapper-command deny rules,find -exec/-deletepermission tightening, and macOS/private/{etc,var,tmp,home}removal protection. ↩ -
Andrew Fasano, Marius Fleischer, Cole McFaul, Robert Xiao, and Tripp Gallagher, “GLM-5.3 and the spread of advanced cyber capabilities,” Anthropic Frontier Red Team, September 29, 2026, fetched October 1, 2026. Quoted or taken from the post: its opening on Glasswing’s head start, the ExploitBench and binary-exploitation results (the latter on 100 randomly selected tasks), the GLM-5.3 browser session, the N-day session with GLM-5.3-Flash, “anyone can download GLM-5.3,” the abliteration cost and its effect on refusal rates, with footnote 3’s estimate for an experienced team, the abliterated copies public within days, footnote 2 on refusals, the bypass rates (64%, 92%, 100%) with Figure 5’s definition of engagement (tried to connect to the target; 50 samples per cell), and “What does this mean?” Anthropic notes its simulated environment executes no model-generated code. ↩↩↩↩
-
NIST Center for AI Standards and Innovation, “CAISI’s Assessment of Z.ai’s GLM-5.3 Cyber Capabilities,” September 17, 2026, fetched October 1, 2026: “GLM-5.3 is the most cyber-capable open-weight model released to date,” “GLM-5.3’s cyber capabilities are significantly lower than those of current U.S. frontier models,” and “lags the capability level of the U.S. frontier by about four months in an aggregate measure of performance across CAISI cyber benchmarks”; the Methodological Notes define “U.S. frontier best” as “including both trusted-access releases and full public releases” and say “when applicable, U.S. models were tested with cyber safeguards disabled.” ↩