OpenAI and Anthropic limit access to highly capable cyber models
OpenAI will restrict Astra's most advanced cybersecurity features to a small testing group after it reached a "Critical" risk threshold, while Anthropic launched a similarly limited Claude Mythos 5.1.

Key takeaways · 3
- 01
Astra is the first OpenAI model to reach the "Critical" cybersecurity threshold.
- 02
Astra chained two zero-day vulnerabilities during testing.
- 03
Anthropic is restricting its Claude Mythos 5.1 model to trusted partners.
OpenAI limits Astra access
OpenAI plans to release its latest model, Astra, soon, but its most advanced cybersecurity features will be limited to a small group of testers. [1] The company stated that Astra is the first model to reach its "Critical" cybersecurity capability threshold. [1] During testing, the model discovered and chained together two zero-day vulnerabilities. [1]
Astra uses a technique called "recurrent depth," which improves cost and performance but obscures the AI's reasoning. [2] OpenAI warned that Astra's safeguards might mistakenly flag legitimate activity as cyber misuse. [1] Through the API, this flagging will stop the task. [1]
Anthropic's parallel release
Anthropic has introduced Claude Fable 5.1 alongside Claude Mythos 5.1. [2] While Claude Fable 5.1 is generally available, access to Mythos 5.1 is reserved for trusted partners. [2] Anthropic stated that Mythos 5.1 includes specific safeguards for cybersecurity and life sciences. [2]
Meanwhile, the Fable 5.1 release is up to 45 percent cheaper for agentic work. [2]
What it means
Both OpenAI and Anthropic are physically bifurcating their frontier models, keeping the most capable versions behind closed doors due to cybersecurity risks. OpenAI's decision to restrict Astra mirrors Anthropic's approach with Claude Mythos 5.1, showing a shared industry caution around autonomous vulnerability discovery. The use of "recurrent depth" in Astra highlights the ongoing trade-off between performance gains and the ability to monitor AI reasoning. What the sources don't address: How OpenAI and Anthropic plan to vet the "trusted partners" and "small group of testers" granted access to these critical capabilities.
AI developers are beginning to gate their most capable models due to demonstrated risks in autonomous cybersecurity tasks. This marks a shift from broad general availability to tiered access based on safety thresholds.
Why it matters
Turn this story into practical AI skill after launch.
Get the release link for daily sessions built around your role and industry.
Join the waitlistHow this developed
2 September 2026
OpenAI and Anthropic limit access to highly capable cyber models
2 September 2026
Event created from source cluster.
Sources
- 5.6 pro model has been automatically downgraded and routed to the 5.5 mini model since its release - ChatGPT / Bugs - OpenAI Developer CommunityOpenAI News Search
- OpenAI to limit access to Astra's most powerful cyber toolsaxios.com
- Techmeme: Source: OpenAI's Astra model uses “recurrent depth”, a technique that improves cost and performance but obscures the AI's reasoning, making it harder Techmeme