Anthropic released Claude Opus 5 on July 24 and said it had deliberately kept cyber training out of the model. The company expects the cyber classifiers around Opus 5 to intervene about 85% less often than those on Fable 5, the model U.S. regulators placed under export controls days after its June launch.
What Changed
- Anthropic released Claude Opus 5 on July 24 and said it intentionally avoided training the model on cyber tasks, as it had with Opus 4.8.
- On OSS-Fuzz, Opus 5 scored non-zero on 79.4% of targets, near Mythos 5's roughly 80%, while completing 4 full exploits against Mythos 5's 13.
- Opus 5's classifiers permit source-code vulnerability research and block binary-based scanning, penetration testing and exploit generation. Anthropic expects them to intervene about 85% less often than Fable 5's.
- The UK's AI Security Institute tested early checkpoints and found Opus 5 solved the cyber range called "The Last Ones" end to end in 8 of 10 attempts.
AI-generated summary, reviewed by an editor. More on our AI guidelines.
OSS-Fuzz and ExploitBench
In its July 24 release, Anthropic wrote that "as with its predecessor, Opus 4.8, we've intentionally avoided training Opus 5 on cyber tasks." The company added that Opus 5 came close to Mythos 5 at finding cybersecurity vulnerabilities, while remaining substantially behind Mythos 5 on exploiting those vulnerabilities.
Marktechpost reported the figures behind that split. On OSS-Fuzz, Opus 5 scored non-zero on 79.4% of targets, near Mythos 5's roughly 80%, while completing 4 full exploits against Mythos 5's 13. On ExploitBench, Opus 5 captured 10.14 mean capability flags in the AutoNudge arm and produced 99 full arbitrary-code-execution exploits, compared with 132 for Mythos 5.
Anthropic describes OSS-Fuzz as an evaluation built to test whether models can find and then exploit vulnerabilities without extensive human guidance. Its safeguards track that division, permitting source-code vulnerability research while other cyber tasks stay blocked.
The 5% classifier run
Opus 5's cyber classifiers allow vulnerability finding in source code and block binary-based vulnerability scanning, penetration testing and exploit generation, Anthropic's release states. The company calls binary-based scanning "a method more likely to be associated with malicious actors."
The filters have already been counted in one run. During Anthropic's Frontier-Bench v0.1 evaluation, Opus 5's safety classifiers flagged and refused 5% of API calls across 4% of trials, while Fable 5's classifiers flagged 42% of calls across 26% of trials, Marktechpost reported from Anthropic's material.
Flagged requests in Claude.ai, Claude Code and Claude Cowork fall back to Opus 4.8 by default, Anthropic said. API customers can enable the same fallback behavior, and enterprises or researchers already in the Cyber Verification Program have immediate access to a version of Opus 5 with fewer security restrictions.
The UK AI Security Institute
The UK's AI Security Institute tested early Opus 5 checkpoints on three cyber ranges at 100 million tokens per attempt, according to the system card. Opus 5 solved the range called "The Last Ones" end to end in 8 of 10 attempts. It did not solve the harder range, "Doing Life," but reached step 22 of 23, further than any model the institute had tested.
Unite.AI noted that the AISI cyber-range work is the one outside check Anthropic discloses. The rest of the launch figures are vendor-reported, run on Anthropic's own harnesses and not independently replicated.
Get Implicator.ai in your inbox
Strategic AI news from San Francisco. No hype, no "AI will change everything" throat clearing. Just what moved, who won, and why it matters. Daily at 6am PST.
No spam. Unsubscribe anytime.
Asked by The Verge whether Anthropic ran Opus 5 past the Trump administration before the public rollout, spokesperson Danielle Ghiglieri responded: "We continue to work with our government partners to conduct their own independent testing of our models. This includes Opus 5." Dianne Penn, Anthropic's head of product management for research, has described the same arrangement, telling CNBC the company keeps collaborating with government agencies on pre-deployment testing, Opus 5 included. Neither addressed whether that testing preceded the launch.
The question follows the June release of Mythos 5 and Fable 5. Fable 5 was the general-availability version of the restricted Mythos model, and the U.S. government imposed export controls days after launch, citing national security authorities, after Amazon researchers reported they could bypass its safeguards. Anthropic pulled Fable 5 on June 12 and restored it on June 30 after strengthening those safeguards, Fortune reported. The export controls lifted after roughly two weeks of negotiations, per CNBC.
Know someone who'd find this useful? ✉️ Email it to a friend in one click, or they can subscribe free here.
Clem Delangue and Z.ai
"Closed model APIs have guardrails that flag and refuse a lot of legitimate security work, because analyzing an attack looks a lot like preparing one," Hugging Face chief executive Clem Delangue told Fortune. "When you're in the middle of an active incident, you can't have your tools refusing to examine malicious payloads or getting your account flagged."
Hugging Face turned to a Chinese open-source model from Z.ai after an unnamed U.S. frontier model refused its requests, the same interview reported. OpenAI models had autonomously hacked Hugging Face's servers earlier in July.
Anthropic's system card also publishes limits outside cybersecurity. Opus 5 "hallucinates factual claims slightly more than Opus 4.8, despite being more accurate overall," Unite.AI reported from the document, and in one internal biology exercise the model got stuck in self-verification loops and failed to deliver protein designs that Mythos 5 completed.
Claude Max and Claude Pro
Opus 5 is available across Anthropic's platforms, the company said, and is now the default model on Claude Max and the strongest model on Claude Pro. It is priced at $5 per million input tokens and $25 per million output tokens, unchanged from Opus 4.8 and half of Fable 5's listed $10 and $50 rates, according to The Decoder's table.
Anthropic also shipped two beta platform updates with the release: mid-conversation tool changes that do not invalidate the prompt cache, and automatic API fallbacks that route classifier-flagged requests to another model rather than returning an error.
Frequently Asked Questions
Did Anthropic train Claude Opus 5 on cybersecurity tasks?
No. Anthropic wrote that "as with its predecessor, Opus 4.8, we've intentionally avoided training Opus 5 on cyber tasks." The company said the model came close to Mythos 5 at finding vulnerabilities anyway, while remaining substantially behind it at exploiting them.
What do Opus 5's cyber classifiers allow and block?
They allow vulnerability finding in source code and block binary-based vulnerability scanning, penetration testing and exploit generation. Anthropic calls binary-based scanning "a method more likely to be associated with malicious actors," and expects the classifiers to intervene about 85% less often than Fable 5's.
What happens when a request trips one of the safety classifiers?
Flagged requests in Claude.ai, Claude Code and Claude Cowork fall back to Opus 4.8 by default. API customers can enable the same behavior. Enterprises and researchers already in Anthropic's Cyber Verification Program get a version of Opus 5 with fewer security restrictions.
How did Opus 5 perform on the UK AI Security Institute's cyber ranges?
The institute tested early checkpoints on three ranges at 100 million tokens per attempt. Opus 5 solved the range called "The Last Ones" end to end in 8 of 10 attempts. It did not solve the harder range, "Doing Life," but reached step 22 of 23, further than any model the institute had tested.
What does Claude Opus 5 cost?
$5 per million input tokens and $25 per million output tokens, unchanged from Opus 4.8 and half of Fable 5's listed $10 and $50 rates. Opus 5 is now the default model on Claude Max and the strongest model on Claude Pro.
AI-generated summary, reviewed by an editor. More on our AI guidelines.



IMPLICATOR