Anthropic, an AI research company, has announced an expanded AI cyber verification program that has uncovered over 129,000 software vulnerabilities since April 2026. More than 33,000 of these vulnerabilities are rated as critical or high risk. The company’s new initiatives consolidate and enhance its earlier Project Glasswing, marking a significant advance in the intersection of artificial intelligence and cybersecurity.
AI cyber verification program: What’s new?
This expanded initiative was publicly unveiled on October 6, 2026 (AI Weekly). Anthropic’s program brings together multiple strands of security testing under one roof. The structure now includes three tiers:
- Defense Access: For incident response teams and vulnerability validation efforts
- Red Team Access: For authorized penetration-testing organizations
- Specialized Access: Focused on safety-critical systems, such as aviation operating systems and power grids
Project Glasswing, the foundation of this expanded program, originally partnered with vetted cybersecurity professionals to test Anthropic’s advanced AI models under reduced safeguards and fewer blocking classifiers. According to the company, these offense-defense exercises led directly to the discovery of tens of thousands of confirmed vulnerabilities between April and July 2026.
How Project Glasswing evolved
Anthropic’s Project Glasswing launched in early 2026 with the goal of exposing advanced AI models to scrutiny from cybersecurity professionals. Glasswing allowed vetted experts to probe models under simulated attack scenarios. These scenarios removed some protective constraints, enabling testers to identify true model and software vulnerabilities that would not surface under regular use conditions.
The latest expansion streamlines access and creates stricter reporting requirements, making results easier to validate. From April to July 2026 alone, Glasswing partners verified over 129,000 vulnerabilities, including 33,000+ ranked as critical or high-severity. In addition, open-source scanning between April and October 2026 discovered at least 5,500 more confirmed vulnerabilities.
Three-tiered structure: Why it matters
A key shift is Anthropic’s move to consolidate previous initiatives under a three-tier cyber verification framework:
- Defense Access grants incident responders and blue teams real-time access to test and validate vulnerability findings during live incidents or major security events.
- Red Team Access intentionally opens models to white-hat security organizations, who can actively hunt for flaws and stress-test AI safety boundaries.
- Specialized Access is reserved for safety-critical sectors: think aviation, transportation control, or SCADA systems, where model errors could have severe physical consequences.
Anthropic claims that, by formalizing these layers, it can better align incentives, enable faster vulnerability disclosure, and improve patch cycles for AI-driven software and infrastructure.
Broader impact: How AI changes vulnerability discovery
AI-driven cyber verification is significant for several reasons:
- Scale: Large AI systems can rapidly test for vulnerabilities far beyond the capacity of traditional static analyzers or manual code reviews.
- Adversarial robustness: Red team exercises with AI models mirror real-world attack and defense patterns, increasing the resilience of model deployments.
- Collaboration: Opening access to vetted external researchers has led to faster discovery and disclosure of risks that might have taken years to reach production systems otherwise.
These moves are part of a wider industry trend. The U.S. and EU are pushing AI model providers to facilitate independent red-teaming and to publicly document vulnerabilities. Anthropic’s latest results come shortly after OpenAI, Google DeepMind, and other labs released similar transparency guidelines and collaboration programs. For more, see CyberProfi’s cybersecurity analysis and AI security trends.
Industry reaction and best practices
The disclosed numbers—over 129,000 bugs, including 33,000 at critical or high risk—underscore concerns about the fragility of both proprietary and open-weight AI software stacks. Security experts recommend several approaches based on Anthropic’s findings:
- Encouraging continuous external testing, especially for safety-critical models and public APIs
- Rapid vulnerability patching, with public notification for severe bugs if exploit evidence emerges
- Maintaining transparent logs of access, discovery, and remediation cycles for all high-stakes AI applications
- Engaging in cross-lab red-teaming and sharing non-sensitive findings with peers and regulators
However, the scale of discovered vulnerabilities also raises questions about broader systemic risk. As both attackers and defenders increasingly rely on AI for automation and exploitation, the security arms race will likely intensify.
Leadership commentary and future plans
While Anthropic has not named all of its partners, sources confirm that several major security research organizations and leading academic teams participated in the latest program cycles. The company intends to expand the initiative, increase transparency, and standardize vulnerability reporting procedures across partners.
Industry observers believe these efforts are crucial for ensuring that the next generation of AI models remains robust against both existing and novel attack types. The rapid pace of adoption in sectors like finance, healthcare, and energy means even minor bugs could potentially trigger outsized consequences (see The Hacker News and SecurityWeek coverage).
Frequently asked questions
- What is Anthropic’s AI cyber verification program?
- It is a multi-tiered security testing initiative that brings together internal and external experts to test Anthropic’s advanced AI models for vulnerabilities. It consolidates previous efforts, notably Project Glasswing, under a single, more accessible structure.
- How many vulnerabilities has the initiative uncovered?
- Anthropic validates that over 129,000 software vulnerabilities were confirmed since April 2026—33,000+ of which were critical or high severity.
- Who participates in the program?
- Vetted cybersecurity professionals, academic teams, and incident response organizations participate in both red-team and defense access tiers.
- Does Anthropic work with outside organizations?
- Yes. The company collaborates with a network of approved cybersecurity researchers, expanding beyond purely in-house testing.
- Where can I find comparable industry efforts?
- Other large AI labs, including OpenAI and Google DeepMind, run similar red-team and vulnerability disclosure initiatives. Details are regularly published on their official security advisories and transparency reports (see more).
