AI is becoming increasingly capable at finding software vulnerabilities.
That creates an unusual problem for the cybersecurity industry.
The same AI system that can help a security team discover and fix a dangerous vulnerability could potentially help someone exploit that vulnerability.
Anthropic is now trying to address that problem in a more structured way.
On October 6, 2026, Anthropic announced an expanded Cyber Verification Program (CVP) that gives qualifying cybersecurity professionals controlled access to advanced cyber capabilities and models with fewer restrictions on legitimate security work. The updated program introduces three different access tiers, with different verification requirements and security controls.
The program includes access to models such as Claude Opus 5.5, Claude Sonnet 5.5, and Claude Mythos 5.1, with additional models expected to be added in the future.
This isn't simply another AI model announcement.
It is an experiment in how increasingly capable AI should be deployed when the technology can be useful for both defense and offense.
What Is Anthropic's Cyber Verification Program?
The Cyber Verification Program is designed for security professionals whose legitimate work can be restricted by the conservative cyber safeguards used on Anthropic's generally available models.
Anthropic says its standard models maintain strong restrictions because cybersecurity is inherently dual-use.
A security researcher might ask an AI system to analyze malware, investigate a vulnerability, or perform an authorized penetration test.
A malicious actor could potentially ask for assistance with similar technical tasks for completely different purposes.
The CVP attempts to separate those situations through verification, access controls, monitoring, and different levels of model capability.
Instead of giving every user the same level of access, Anthropic has created three tiers.
The Three Access Tiers
Defense Access
The first level is designed for defensive cybersecurity work.
This includes activities such as security operations, incident response, malware reverse engineering, vulnerability analysis, and validating vulnerabilities.
Anthropic says qualifying organizations can include corporate security teams, nonprofits, universities, government bodies, critical infrastructure operators, smaller security firms, open-source maintainers, and individual researchers with a track record of reporting vulnerabilities.
The idea is relatively straightforward: organizations defending systems they own or maintain should be able to use more capable AI security tools without being blocked by safeguards designed primarily to prevent malicious use.
Red Team Access
The second level adds authorized penetration testing and red-team operations.
This is important because penetration testing is intentionally adversarial. Security professionals need to behave like attackers in order to understand how a system could actually be compromised.
But Anthropic says organizations in this tier can only perform adversarial testing against systems they are authorized to test.
The company also says certain dangerous activities remain blocked, including actions that could cause physical harm or mass disruption, such as deploying ransomware, damaging physical systems, or testing certain high-risk safety systems.
Specialized Access
The third level is the most restricted and is intended for a limited group of verified organizations.
It is designed for organizations authorized to test systems where failures could potentially affect people's lives or disrupt major infrastructure.
Anthropic gives examples including:
- Flight operating systems
- Power grids
- Telecommunications networks
- Interbank transfer infrastructure
- Government administrative networks
Anthropic says organizations applying for this level undergo deeper review in collaboration with the U.S. government. Existing Project Glasswing members will transition into this tier.
This is where the program becomes particularly interesting.
AI is no longer being treated simply as software that can be switched on or off.
Instead, access to its capabilities is being matched to the potential consequences of the work being performed.
Why Does AI Cybersecurity Need Special Access?
The reason is the dual-use nature of cybersecurity.
Imagine an AI model that becomes extremely good at finding vulnerabilities.
For a hospital, that capability could help discover weaknesses before attackers find them.
For a criminal group, the same capability could potentially be used to identify systems that are easier to compromise.
The underlying technical capability hasn't necessarily changed.
The context and authorization have changed.
That is why Anthropic's approach focuses heavily on verification.
The company says its generally available models continue to have conservative cyber safeguards, while approved organizations can receive more permissive access through CVP.
This creates an interesting model for AI deployment:
More sensitive capability → stronger verification → more controlled access.
What Did Anthropic Learn From Project Glasswing?
The expanded CVP builds on Project Glasswing, which Anthropic launched earlier in 2026 to help organizations use advanced AI to secure important software.
According to Anthropic, its Glasswing partners identified at least 129,000 verified software vulnerabilities between April and July 2026.
Anthropic says its own open-source scanning efforts identified another 5,500 verified vulnerabilities between April and October.
The company says more than 33,000 of the vulnerabilities identified so far were rated critical or high severity.
There is an important qualification here.
Anthropic describes these figures as a lower bound because they are based partly on reports from a subset of its partners, and not every partner disclosed complete patching data. The company therefore says the true impact could be substantially higher.
So these numbers should be understood as Anthropic's reported results from the program, rather than an independently established measurement of all vulnerabilities discovered by AI.
Still, they illustrate why companies are taking AI-assisted cybersecurity seriously.
How Much Difference Do the Safeguards Make?
Anthropic also tested its tiered approach using CyScenarioBench, an evaluation designed to measure whether models can plan and execute multi-stage cyber operations under realistic constraints.
Anthropic tested Claude Opus 5.5 across five attempts on each of ten challenges for each access tier.
The results showed a substantial difference between the levels of access.
Without CVP access, Anthropic says every task was blocked at the first prompt.
In the Defense Access tier, 46 of 50 trials were blocked at some point.
In the Red Team Access tier, Anthropic says there were no blocks, and the model successfully completed 34 of 50 tasks.
Anthropic reports that this was effectively the same completion rate as the model achieved when no safeguards were applied.
These results are useful because they demonstrate the trade-off Anthropic is trying to manage.
Strong safeguards can prevent harmful behavior, but they can also prevent legitimate security professionals from performing realistic testing.
The challenge is finding the right level of restriction for each situation.
This Is Not Unlimited Access
It is important not to misunderstand what Anthropic announced.
The CVP does not mean that anyone working in cybersecurity can simply access unrestricted Claude models.
Organizations must apply and go through verification.
Anthropic says applicants must provide evidence of the required security controls for the relevant access tier. Higher-risk tiers have additional requirements and deeper review.
The program is currently available through the Claude Platform, Google Cloud's Vertex AI, and Microsoft Foundry. Anthropic says Amazon Bedrock access is available only for customers eligible for its Enterprise Frontier Safeguards.
That controlled-access structure is one of the most important parts of the announcement.
Why This Matters for AI Agents
This development becomes even more interesting when we look at the broader evolution of AI agents.
An ordinary chatbot answers a question.
An AI agent can potentially:
- inspect files
- use tools
- execute commands
- analyze systems
- make decisions
- interact with software
- continue working across multiple steps
Cybersecurity is one of the areas where those capabilities can become extremely powerful.
A capable AI agent could potentially scan thousands of systems much faster than a human security team.
That could help defenders find vulnerabilities before attackers do.
But it also means that mistakes—or deliberate misuse—could have much larger consequences.
Anthropic's CVP suggests that the industry may need to think about authorization as part of AI capability itself.
The question may no longer be simply:
How capable is the model?
It may also be:
Who is allowed to use which capabilities, under what conditions?
Could This Become a New Model for AI Safety?
Possibly.
The idea behind CVP is broader than cybersecurity.
We could eventually see similar structures for other sensitive areas where AI capabilities have significant dual-use risks.
For example, a highly capable model might require different access rules when it is being used for:
- Critical infrastructure
- Biological research
- Financial systems
- Industrial control systems
- Government networks
- Safety-critical engineering
Instead of having one universal safety setting, AI systems could increasingly use capability-based access controls.
A researcher working on harmless defensive security might receive one level of access.
An authorized red team might receive another.
An organization testing a power grid or flight system could require substantially stronger verification.
That is essentially the philosophy Anthropic is testing with CVP.
The Bigger Question: Can AI Give Defenders an Advantage?
This may ultimately be the most important question.
Cybersecurity has traditionally been a race between attackers and defenders.
Attackers look for vulnerabilities.
Defenders try to discover and fix them first.
AI changes the speed of that race.
If AI can discover vulnerabilities significantly faster, defenders could theoretically scan software continuously instead of waiting for security researchers or attackers to discover problems.
But the same technology could also increase the scale and speed of attacks.
Anthropic's own Project Glasswing work reflects that tension.
The company argues that advanced AI can give defenders a durable advantage, but it also acknowledges that the capabilities are inherently dual-use.
That means the technology alone doesn't determine the outcome.
Deployment rules matter.
What Could Happen Next?
Anthropic says it plans to continue refining its tier-based classifiers and expand the impact of the program to more cyber defenders.
The company also says it will share more information about what it learns as the program develops.
Another interesting development is Anthropic's planned Enterprise Frontier Safeguards, which the company says will combine strong safeguards with customer-controlled cloud infrastructure and zero-data-retention privacy. Eligible organizations are expected to gain access later this fall.
That points toward a future where powerful AI systems can operate in sensitive environments without requiring organizations to give up the privacy and control expected for critical workloads.
Final Thoughts
Anthropic's expanded Cyber Verification Program is important because it shows how the AI industry is beginning to deal with a difficult reality.
The more capable an AI becomes, the harder it becomes to separate useful capabilities from potentially dangerous ones.
Cybersecurity makes that problem particularly obvious.
The same model can help a security researcher discover a vulnerability and potentially help an attacker understand how to exploit one.
Anthropic's answer is not to remove those capabilities completely.
Instead, it is experimenting with verified users, tiered access, stronger monitoring, and different levels of safeguards.
Whether this approach becomes a successful long-term model remains to be seen.
But one thing is already clear:
The next phase of AI safety may not be only about making models refuse dangerous requests.
It may also be about deciding who gets access to powerful capabilities, how much access they receive, and what happens when those capabilities are used in the real world.
And as AI agents become more autonomous, that question is likely to become even more important.

0 Comments