Breaking News / Artificial Intelligence
The model scores 43.3% on Frontier-Bench agentic coding, more than double Opus 4.8’s 21.1%, while matching its predecessor’s price. It is now the default on Claude Max and the strongest model available on Claude Pro.
Frontier-Bench Score
43.3% (new SOTA)
Pricing vs Fable 5
Half the cost
ARC-AGI-3 Score
30.2% (3x next best)
Availability
Available now
Anthropic released Claude Opus 5 on July 24, 2026, positioning it as a daily-use model that approaches the capability of its flagship Claude Fable 5 at roughly half the price. The company announced the model on its official blog, stating it is available immediately through the Claude API and consumer tiers.
Opus 5 is now the default model on Claude Max and the strongest model available on Claude Pro. Anthropic describes it as “thoughtful and proactive,” designed for software engineering and knowledge work rather than experimental or research-only use cases. The company explicitly notes that Opus 5 remains behind Mythos 5 on cybersecurity tasks, a limitation that reflects Anthropic’s continued caution around releasing its most capable security-oriented capabilities to general users.
Benchmark Performance
Anthropic published a comprehensive benchmark comparison pitting Opus 5 against Claude Fable 5, Claude Opus 4.8, and OpenAI’s GPT-5.6 Sol. The results show Opus 5 establishing new state-of-the-art results on coding and knowledge work evaluations while trading blows with Fable 5 on multidisciplinary reasoning.
| Benchmark | Opus 5 | Fable 5 | Opus 4.8 | GPT-5.6 Sol |
|---|---|---|---|---|
| Agentic terminal coding (Frontier-Bench v0.1) | 43.3% | 33.7% | 21.1% | 34.4% |
| Knowledge work (GDPval-AA v2) | 1861 | 1747 | 1593 | 1736 |
| Novel problem-solving (ARC-AGI-3) | 30.2% | – | 1.5% | 7.8% |
| Agentic search (BrowseComp) | 90.8% | 87.4% | 84.3% | 90.4% |
| Computer use (OSWorld 2.0) | 70.6% | 66.1% | 55.7% | 62.6% |
| Business workflows (AutomationBench) | 26.0% | 17.4% | 17.0% | 18.1% |
| Multidisciplinary reasoning (Humanity’s Last Exam, no tools) | 56.3% | 56.5% | 49.8% | – |
| Multidisciplinary reasoning (Humanity’s Last Exam, with tools) | 64.7% | 63.9% | 57.9% | – |
| Agentic coding (DeepSWE v1.1) | 68.8% | 69.7% | 59.0% | 72.7% |
| Agentic coding (FrontierCode v1.1) | 53.4% | 53.5% | 46.5% | 47.5% |
| Legal (Legal Agent Benchmark) | 11.7% | 13.3% | 10.4% | 2.5% |
| Health (HealthBench Professional) | 59.8% | 66.0% | 57.4% | 60.5% |
| Biology (BioMysteryBench, hard) | 49.4% | 46.5% | 42.4% | – |
| Biology (BioMysteryBench, human solved) | 90.1% | 89.0% | 88.5% | – |
Source: Anthropic official announcement, July 24, 2026. Dashes indicate the model was not evaluated on that benchmark or data was not provided.
Cost Efficiency and the “Effort” Setting
Anthropic emphasized that Opus 5 delivers its performance gains without a price increase over Opus 4.8. The company published cost-performance curves showing that on Frontier-Bench v0.1, Opus 5 not only surpasses all competitors but more than doubles Opus 4.8’s score at a lower cost per task. On CursorBench 3.2, Opus 5 at maximum effort performs within 0.5% of Fable 5’s peak score while costing roughly half as much per task.
The model introduces an adjustable “effort” setting that lets users optimize for either maximum intelligence or token conservation. At lower effort levels, Opus 5 still passes more tasks on Zapier’s AutomationBench than any other model at equivalent cost, according to Anthropic. On OSWorld 2.0, a computer-use benchmark, Opus 5 outperforms every competitor at any given cost point and surpasses Fable 5’s best result at approximately one-third of the cost.
The ARC-AGI-3 result is particularly striking. At 30.2%, Opus 5’s score on this novel problem-solving evaluation is roughly three times higher than GPT-5.6 Sol’s 7.8% and twenty times higher than Opus 4.8’s 1.5%. Anthropic frames this as evidence that Opus 5 has made genuine progress on tasks requiring reasoning outside its training distribution, not just memorization or pattern matching.
Where It Falls Short
The benchmark table also reveals clear limitations. On the Legal Agent Benchmark held-out set, Opus 5 scores 11.7%, trailing Fable 5’s 13.3%. On HealthBench Professional, it scores 59.8% versus Fable 5’s 66.0%. On DeepSWE v1.1, GPT-5.6 Sol leads at 72.7% compared to Opus 5’s 68.8% and Fable 5’s 69.7%.
Most notably, Anthropic states that Opus 5 remains behind Mythos 5 on cybersecurity tasks. Mythos 5 is Anthropic’s restricted-tier model, available only to government and critical infrastructure partners through Project Glasswing. That gap suggests Anthropic is still withholding its most capable security-oriented model from general release, a decision the company has framed around safety rather than product segmentation.
Pricing and Positioning
Anthropic has priced Opus 5 at the same level as Opus 4.8, which costs $5 per million input tokens and $25 per million output tokens. That is half the price of Claude Fable 5, which runs at $10 per million input tokens and $50 per million output tokens. Fable 5 was suspended from general API access on June 12, 2026, and remains available only to select partners, making Opus 5 the most capable publicly accessible Claude model as of this release.
The pricing strategy places Opus 5 in direct competition with OpenAI’s GPT-5.6 Sol on cost and significantly below Fable 5’s tier. For developers running agentic coding workflows, the cost difference is material. Anthropic’s own data shows that on complex multi-step tasks, Opus 5’s higher per-task success rate can offset its per-token price when compared to models that fail more frequently and require more attempts.
What This Means for the Market
The release tightens the race at the top of the large language model market. OpenAI shipped GPT-5.6 Sol earlier this year, and Google’s Gemini 3.1 Pro has been competitive on software engineering benchmarks. Anthropic’s move to deliver near-Fable performance at Opus pricing pressures both competitors to justify their own flagship pricing or risk losing enterprise coding contracts.
For everyday users, the shift is simpler. Claude Pro and Max subscribers now have access to a more capable default model without a price increase. The adjustable effort setting also gives users more control over their token spend, which matters for anyone hitting usage limits on subscription plans.
Whether Opus 5’s benchmark gains translate into consistently better real-world performance is a question that will be answered over the coming weeks as developers integrate it into production workflows. Anthropic’s track record suggests the gap between benchmark scores and practical utility is narrowing, but the only reliable test is sustained use at scale.
Sources
- Anthropic: “Introducing Claude Opus 5” (July 24, 2026)
- Anthropic Platform Docs: Claude Pricing (July 2026)
- VanceIQ: “Claude Opus 5: Frontier Intelligence at Half Cost” (July 25, 2026)
- Data Dynamics: “What’s Different About Claude Fable 5” (June 12, 2026)
- Finout: “Claude Fable 5 and Mythos 5: Pricing, API Costs, and Benchmark Comparison” (June 10, 2026)
- MorphLLM: “Claude Benchmarks (2026): Fable 5 Hits 95% SWE-bench Verified” (June 9, 2026)
This article was published on July 25, 2026. Benchmark scores are self-reported by Anthropic unless otherwise noted. Real-world performance may vary by use case.