Claude Opus 5: Near-Fable Intelligence at Half the Price, the ARC-AGI Breakthrough, and the New Default for Agentic Work
Anthropic releases Claude Opus 5 on July 24, 2026 β near Fable 5 intelligence at $5/$25 (half the price). New SOTA on Frontier-Bench (43.3%), ARC-AGI-3 (30.2%, 4Γ GPT-5.6 Sol), and GDPval-AA (1861 Elo). Thinking on by default, five-level effort control, 1M context, and the most aligned Claude model to date. Analysis of benchmarks, behavioral shifts, safety, and what it means for the frontier.
Executive Summary
On July 24, 2026, Anthropic released Claude Opus 5, positioning it as "near Fable 5 intelligence at half the price." Priced at $5 per million input tokens and $25 per million output tokens β identical to Opus 4.8 and exactly half of Fable 5's $10/$50 β Opus 5 is now the default model on Claude Max and the strongest model on Claude Pro. It features a 1M-token context window, 128K max output, thinking on by default, and a new five-level effort control (low, medium, high, xhigh, max).
The benchmarks tell an extraordinary story. On Frontier-Bench v0.1 β an agentic terminal coding benchmark measuring real software engineering work β Opus 5 scores 43.3%, more than doubling Opus 4.8's 18.7% and surpassing Fable 5's 33.7%. On ARC-AGI-3 β FranΓ§ois Chollet's fluid intelligence benchmark designed to resist pattern-matching β Opus 5 achieves 30.2%, roughly four times GPT-5.6 Sol's 7.8% and twenty times Opus 4.8's 1.5%. On GDPval-AA v2, it reaches 1861 Elo, the new state-of-the-art for knowledge work.
But the numbers are only part of the story. Opus 5 behaves fundamentally differently from prior models: it verifies its own work without being told, delegates to subagents more readily, narrates progress during agentic sessions, and catches its own errors before producing output. These behavioral shifts make it a qualitative leap for long-horizon agentic work β the kind of multi-file refactors, end-to-end feature builds, and complex debugging that enterprises actually pay for.
This article covers the full picture: benchmark performance, architectural changes, the effort dial, safety and alignment, pricing dynamics, and what Opus 5's release means for the frontier AI landscape.
1. The Launch: What Anthropic Said
1.1 The Official Announcement
Anthropic's announcement post on July 24, 2026 was unusually direct:
"Claude Opus 5 is available today. It's a thoughtful and proactive model that comes close to the frontier intelligence of Claude Fable 5 at half the price."
The key claims:
- Near Fable 5 intelligence at half the price β $5/$25 vs. $10/$50
- New SOTA on coding and knowledge work β Frontier-Bench, GDPval-AA, ARC-AGI-3
- Behind Mythos 5 on cybersecurity β intentionally not trained on cyber tasks
- Most aligned Claude model to date β lowest misaligned behavior score (2.30)
- Default on Claude Max β the strongest model for Max subscribers
1.2 What Was Shipped
| Component | Status |
|---|---|
| claude-opus-5 API model | β Live (Claude API, Bedrock, Vertex, Foundry) |
| 1M context window | β Default and maximum |
| 128K max output | β (300K via Batch API beta) |
| Thinking on by default | β Adaptive thinking, cannot be disabled at xhigh/max effort |
| Five-level effort control | β low, medium, high, xhigh, max |
| Mid-conversation tool changes | β
Beta (mid-conversation-tool-changes-2026-07-01) |
| Default fallbacks mode | β
Beta (server-side-fallback-2026-07-01) |
| Lower prompt cache minimum | β 512 tokens (down from 1,024) |
| Fast mode | β Research preview ($10/$50, API only) |
| System card | β Published with full benchmark tables |
2. Benchmark Performance: The New SOTA
2.1 Agentic Coding: Frontier-Bench v0.1
Frontier-Bench v0.1 measures real software engineering work β multi-file changes, debugging, building features from specifications. It's the benchmark that most directly maps to what enterprises pay for.
| Model | Frontier-Bench v0.1 |
|---|---|
| Claude Opus 5 | 43.3% |
| Claude Fable 5 | 33.7% |
| GPT-5.6 Sol | 34.4% |
| Claude Opus 4.8 | 18.7% |
Opus 5 more than doubles Opus 4.8's score and clears Fable 5 by nearly 10 points β a result that contradicted the pre-launch consensus that Opus 5 would be "comparable to Fable 5, but won't surpass it."
On CursorBench 3.2, Opus 5 lands within 0.5% of Fable 5's peak score at max effort, but at half the cost per task. Anthropic claims it achieves greater performance at a given cost than every other model on high, xhigh, and max effort settings.
On FrontierCode 1.1 Extended (from Cognition/Devin), Opus 5 reaches 63.6% at its best effort setting, with particular strength in debugging and root-cause analysis.
2.2 Fluid Intelligence: ARC-AGI-3
This is the result that stopped scrolling.
ARC-AGI-3, designed by FranΓ§ois Chollet, drops an AI agent into interactive, turn-based environments with no instructions, no stated rules, and no stated goal. The agent must work out everything through trial and error. It's scored by Relative Human Action Efficiency.
| Model | ARC-AGI-3 |
|---|---|
| Claude Opus 5 | 30.2% |
| GPT-5.6 Sol | 7.8% |
| Claude Opus 4.8 | 1.5% |
| Best score at release (Mar 2026) | 0.37% |
Opus 5's 30.2% is roughly four times GPT-5.6 Sol's result and twenty times Opus 4.8's. When ARC Prize Foundation released ARC-AGI-3 in March 2026, the best AI model scored 0.37%. A 30% score means Opus 5 is genuinely working out unfamiliar rule systems through interaction β not recalling something adjacent to what it has seen.
This is the hardest property to fake in benchmarking. Benchmarks like SWE-Bench or GDPval reward a model for doing familiar categories of work faster and cheaper. ARC-AGI-3 rewards a model for handling something it has never been trained to handle at all.
2.3 Knowledge Work & Automation
| Benchmark | Opus 5 | Next Best | Context |
|---|---|---|---|
| GDPval-AA v2 (ELO) | 1861 | ~1736 (GPT-5.6 Sol) | New SOTA for knowledge work |
| Zapier AutomationBench | 26.0% | ~17% | ~1.5Γ next-best at same cost |
| OSWorld 2.0 | 70.6% | Fable 5 (at higher cost) | Surpasses Fable 5 at 1/3 the cost |
On Zapier AutomationBench, Opus 5's pass rate is roughly 1.5Γ the next-best model for the same cost per task. Even at its lowest effort setting, it passes more tasks than any other model at any effort level. Wade Foster at Zapier confirmed Opus 5 ran a "full churn-prevention sequence end to end" on a raw account-health workbook β something previous models couldn't do.
2.4 Science & Visual Outputs
Opus 5 outperforms Opus 4.8 on every life sciences evaluation Anthropic tracks:
| Domain | Improvement over Opus 4.8 |
|---|---|
| Organic chemistry | +10.2 percentage points (inferring molecular structures from spectroscopy) |
| Protein-related tasks | +7.7 percentage points (predicting sequence-function relationships) |
| Structural biology | Improved (specific margin not disclosed) |
| Bioinformatics | Improved (specific margin not disclosed) |
Visual output quality is substantially stronger, with users noting "insane" 3D generation quality, improved chart understanding, and the ability to produce interactive illustrations of scientific concepts.
2.5 Cybersecurity: Close on Finding, Far on Exploiting
On OSS-Fuzz (Anthropic's vulnerability-and-exploit benchmark):
| Capability | Opus 5 | Mythos 5 |
|---|---|---|
| Vulnerability discovery | Close to Mythos 5 | Frontier |
| Exploit development | Far behind Mythos 5 | Frontier |
Opus 5 identifies vulnerabilities with success rates close to Mythos 5 but falls well behind on turning found bugs into working attacks. This is by design β Anthropic intentionally avoided training Opus 5 on cyber tasks. The model improved on cybersecurity as a side effect of becoming more generally capable, but the safeguards block offensive capabilities.
3. The Effort Dial: Five Levels of Reasoning
3.1 The Effort Ladder
Opus 5 ships with a five-level effort control that determines how much reasoning the model expends:
| Level | Use Case | Token Cost | Latency |
|---|---|---|---|
| low | High-volume, simple tasks | Minimal | Fastest |
| medium | Standard queries, document work | Low | Fast |
| high | Default β complex reasoning | Moderate | Moderate |
| xhigh | Demanding coding, agentic work | High | Slow |
| max | Hardest problems, multi-step analysis | Highest | Slowest |
The effort parameter controls thinking depth, not response length. Lower effort means fewer tokens, faster responses, lower cost. Higher effort means more thinking, better results, more spend.
3.2 Why Effort Matters More on Opus 5
Opus 5 converts additional effort into better results more reliably than any earlier Opus model. The benchmark charts Anthropic published are all cost-versus-performance curves, not single-point scores. Opus 5's advantage isn't just that it scores higher at max effort β it's that at every effort level, it delivers more performance per dollar than every other model.
On an internal trading benchmark, Opus 5 reaches the best score using roughly a seventh of the reasoning tokens and under half the latency of Opus 4.8.
3.3 Implementation
from anthropic import Anthropic
client = Anthropic()
# Default effort (high) β good starting point
response = client.messages.create(
model="claude-opus-5",
max_tokens=64000,
messages=[
{"role": "user", "content": "Debug this multi-file refactoring issue"}
],
# thinking is on by default; effort controls depth
output_config={"effort": "high"},
)
# Max effort for the hardest problems
response = client.messages.create(
model="claude-opus-5",
max_tokens=64000,
messages=[
{"role": "user", "content": "Design a new market data feed parser"}
],
output_config={"effort": "max"},
)
# Low effort for high-volume tasks
response = client.messages.create(
model="claude-opus-5",
max_tokens=8192,
messages=[
{"role": "user", "content": "Summarize this document"}
],
output_config={"effort": "low"},
)
4. Behavioral Shifts: How Opus 5 Works Differently
4.1 Self-Verification
The most significant behavioral change is that Opus 5 verifies its own work without being told to. This means:
- Remove verification instructions from legacy prompts ("include a final verification step," "use a subagent to verify") β they cause over-verification and waste tokens
- Remove re-check instructions ("double-check your answer," "re-verify before responding") β the model already does this
- The model narrates corrections more than prior models, which may need tuning for user-facing products
4.2 Agentic Narration
Opus 5 narrates its progress during agentic sessions more than prior models. It tends to announce what it's about to do, and its per-message output in agentic sessions is often longer. This can be tuned with explicit prompts:
Before your first tool call, say in one sentence what you're about to do.
While working, give a brief update only when you find something important
or change direction. When you finish, lead with the outcome.
4.3 Subagent Delegation
Opus 5 delegates to subagents more readily than prior models. Delegation pays off on genuinely independent, sizeable tracks of work but multiplies cost and time when applied to small tasks. For cost-sensitive workloads, explicit caps on subagent spawning are recommended.
4.4 Task Scope
Opus 5 can expand the scope of a task, adding steps that weren't requested. For narrow tasks, explicit scope constraints are needed:
Deliver what was asked, at the scope intended. Make routine judgment
calls yourself, and check in only when different readings of the request
would lead to materially different work.
4.5 Real-World Examples
The announcement included several illustrative examples of Opus 5's behavior:
-
FreeCAD 3D reconstruction: Given a drawing of a machine part with no way to directly view it, Opus 5 wrote its own computer vision pipeline to pull geometry from raw pixels, then reconstructed the full model. No competing model solved it after five attempts.
-
Package manager bug fix: Given a real bug in a popular open-source package manager, Opus 5 found the root cause and fixed an edge case the community's patch had missed. A competing model fixed only the surface symptom.
-
Market data feed: An engineer at a trading firm used Opus 5 to build a market data feed for a new exchange in a single session. Finding no live feed to validate against, Opus 5 built its own test harness.
5. API Changes and Migration
5.1 Thinking On by Default
On Opus 4.8, requests run without thinking unless you set thinking: {"type": "adaptive"}. On Opus 5, thinking is on by default. The model decides when and how much to think on each turn, and the effort parameter controls thinking depth.
Breaking change: max_tokens is a hard limit on total output (thinking plus response text). Workloads that ran without thinking on Opus 4.8 may need larger max_tokens values.
5.2 Disabling Thinking Requires Effort β€ High
On Opus 5, thinking: {"type": "disabled"} is accepted only when effort is high or below. Setting thinking disabled with xhigh or max effort returns a 400 error.
# β
Valid β disabled thinking with high effort
client.messages.create(
model="claude-opus-5",
thinking={"type": "disabled"},
output_config={"effort": "high"},
messages=[{"role": "user", "content": "Quick answer"}],
)
# β Error β disabled thinking with max effort
client.messages.create(
model="claude-opus-5",
thinking={"type": "disabled"},
output_config={"effort": "max"}, # Returns 400
messages=[{"role": "user", "content": "Complex task"}],
)
5.3 Mid-Conversation Tool Changes (Beta)
You can now add or remove tools between turns while preserving the prompt cache. Include the mid-conversation-tool-changes-2026-07-01 beta header.
5.4 Default Fallbacks Mode
The fallbacks parameter supports a new "default" mode, applying Anthropic's recommended fallback models by refusal category. Use the server-side-fallback-2026-07-01 beta header.
5.5 Lower Prompt Cache Minimum
The minimum cacheable prompt length is 512 tokens (down from 1,024 on Opus 4.8). Prompts that were too short to cache on Opus 4.8 can now create cache entries with no code changes.
6. Safety and Alignment
6.1 Most Aligned Claude Model
On Anthropic's automated behavioral audit, Opus 5 scores 2.30 on overall misaligned behavior β the lowest of any recent Claude model, ahead of Opus 4.8, Sonnet 5, and Fable 5. Anthropic describes it as the least deceptive and the strongest adherent to Claude's Constitution.
6.2 Cybersecurity Safeguards
Opus 5's cyber classifiers are proportionally less restrictive than Fable 5's:
- Allowed: Finding vulnerabilities in source code (defensive work)
- Blocked: Binary-based vulnerability scanning, penetration testing, exploit generation
Anthropic expects the classifiers to intervene roughly 85% less often than Fable 5's. When they do trigger, flagged requests fall back to Opus 4.8 automatically.
6.3 No Frontier Advance in Dual-Use
Opus 5 does not advance the frontier in risky, dual-use capabilities. In evaluations conducted alongside private-sector and government partners, it remains behind Mythos 5 in both biology research and offensive cybersecurity.
7. Pricing and Access
7.1 API Pricing
| Model | Input ($/MTok) | Output ($/MTok) | Total ($/MTok) |
|---|---|---|---|
| Claude Opus 5 | $5.00 | $25.00 | $30.00 |
| Claude Opus 4.8 | $5.00 | $25.00 | $30.00 |
| Claude Fable 5 | $10.00 | $50.00 | $60.00 |
| Claude Sonnet 5 | $2.00ΒΉ | $10.00ΒΉ | $12.00ΒΉ |
| Claude Haiku 4.5 | $1.00 | $5.00 | $6.00 |
| GPT-5.6 Sol | $5.00 | $30.00 | $35.00 |
| GPT-5.6 Terra | $2.50 | $15.00 | $17.50 |
| Gemini 3.6 Flash | $1.50 | $7.50 | $9.00 |
ΒΉ Sonnet 5 introductory pricing through August 31, 2026 (then $3/$15)
Opus 5 is priced identically to Opus 4.8 and exactly half of Fable 5. For agentic workloads generating 100K output tokens, a single task costs $2.50 on Opus 5 vs. $5.00 on Fable 5 vs. $3.00 on GPT-5.6 Sol.
7.2 Fast Mode
Fast mode (research preview) provides 150% faster token generation at double the standard price ($10/$50). Available on the Claude API only β not on Bedrock, Vertex, or Foundry.
7.3 Subscription Access
| Plan | Opus 5 Access |
|---|---|
| Max | Default model (replaces Opus 4.8) |
| Pro | Strongest model available |
| Enterprise | Available via API and cloud platforms |
| Free | No access |
7.4 Cloud Availability
Opus 5 is available on:
- Claude API β
claude-opus-5 - Amazon Bedrock β
anthropic.claude-opus-5 - Google Cloud Vertex AI β
claude-opus-5 - Microsoft Foundry β Available
8. Integration with Prior Research
8.1 The Fable 5 Story Continues
Opus 5's release is the natural evolution of the Fable 5 saga documented in Claude Fable 5 Mythos 5 Full Return Safeguards Jacobian Conjecture 2026 07 23:
- Fable 5 was pulled for 19 days due to export controls, then restored with aggressive safeguards that blocked too many benign requests
- Opus 5 delivers near-Fable intelligence with 85% fewer classifier interventions β Anthropic's answer to the "too restrictive" critique
- The Fable/Mythos dual-release template is now joined by the Opus 5 as default strategy β making near-frontier capability the standard, not the premium
8.2 The Efficiency Thread
Opus 5 continues the efficiency trajectory seen in Gemini 3 6 Flash 3 5 Flash Lite Cyber Token Efficiency Agentic Scale 2026 07 22:
- Google optimized for fewer output tokens (17% reduction in 3.6 Flash)
- Anthropic optimized for more performance per dollar (Opus 5 beats every model at every effort level)
- Both represent a shift from "highest benchmark score" to "best value for agentic work"
8.3 The Open-Weight Context
In the context of the open-weight arms race covered in Qwen3 8 Max Preview 2 4t Multimodal Moe Open Weight Promise 2026 07 24 and Thinking Machines Inkling 975b Multimodal Moe Self Improvement Controllable Effort 2026 07 21:
- Opus 5 is closed-weight β no open release planned
- The gap between closed-weight (Opus 5, Fable 5) and open-weight (Kimi K3, Qwen3.8) models remains significant on agentic coding and fluid intelligence
- The ARC-AGI-3 result (30.2%) is particularly hard for open-weight models to match, as it requires genuine fluid reasoning, not pattern matching
8.4 Multi-Model Routing
Opus 5's five-level effort control and strong low-effort performance make it ideal for the multi-model routing strategies discussed in Howto Multi Model Routing Layer:
- Simple queries β Opus 5 at
loweffort ($5/$25, fast) - Complex coding β Opus 5 at
xhighormaxeffort - Hardest problems β Fable 5 (if the extra 0.5% on CursorBench justifies 2Γ cost)
- High-throughput β Sonnet 5 or Haiku 4.5
9. Key Takeaways
-
Opus 5 is the new default for serious work. At $5/$25 with near-Fable intelligence, it's the most cost-effective frontier model available. If you're using Opus 4.8, migrate to Opus 5. If you're using Fable 5 for non-cybersecurity tasks, evaluate whether Opus 5 at max effort meets your needs at half the cost.
-
The ARC-AGI-3 result is the real story. A 30.2% score on a benchmark designed to resist pattern-matching suggests Opus 5 has genuine fluid reasoning capabilities, not just better memorization. This is the hardest property to fake and the most significant for long-term AI capability.
-
The effort dial is a game-changer for cost optimization. The ability to tune reasoning depth from
lowtomaxβ with Opus 5 delivering more performance per dollar than every competitor at every level β transforms how organizations budget for AI workloads. -
Self-verification changes the prompting game. Opus 5 verifies its own work without being told to, which means legacy prompts with explicit verification instructions are now counterproductive. This requires updating prompt templates across integrations.
-
The safety story is improved. With 85% fewer classifier interventions than Fable 5 and the lowest misaligned behavior score of any recent Claude model, Opus 5 addresses both the "too restrictive" and "not safe enough" critiques from the Fable 5 saga.
-
Thinking on by default is a breaking change. Workloads that ran without thinking on Opus 4.8 need larger
max_tokensvalues, and disabling thinking atxhigh/maxeffort is now impossible. This requires code updates for existing integrations. -
The frontier is bifurcating. Fable 5 remains the king of cybersecurity and the hardest coding tasks (SWE-Bench Pro 80.3%), but Opus 5 has taken the crown on agentic coding (Frontier-Bench), fluid intelligence (ARC-AGI-3), and knowledge work (GDPval-AA). The choice between them is no longer about raw capability β it's about the specific workload.
10. Future Directions
10.1 What to Watch
- Opus 5 adoption curves: How quickly do enterprises migrate from Opus 4.8 to Opus 5? The identical pricing makes this a no-brainer for most workloads.
- Effort level distributions: What effort levels do organizations actually use in production? The sweet spot between
highandxhighwill define real-world cost profiles. - ARC-AGI-3 independent verification: The 30.2% score needs independent confirmation. If it holds, it's a qualitative leap in AI fluid reasoning.
- Fast mode graduation: If fast mode moves from research preview to general availability, it could create a speed-tier alongside the effort tiers.
- Opus 5 vs. Kimi K3 open-weight: When Kimi K3's weights land on July 27, the community will test whether open-weight models can approach Opus 5's ARC-AGI-3 and Frontier-Bench scores.
- Sonnet 5 price increase: The $2/$10 introductory pricing ends August 31, 2026. The move to $3/$15 will reshape the mid-tier pricing landscape.
10.2 Strategic Implications
Opus 5's release signals that the frontier model race has entered a value phase:
- Capability is table stakes. With multiple models achieving near-frontier performance, the competition is shifting to cost-effectiveness, behavioral reliability, and developer experience.
- The effort dial is the new pricing lever. Rather than offering multiple model tiers (Luna/Terra/Sol, Haiku/Sonnet/Opus/Fable), Anthropic is offering a single model with a tunable reasoning depth. This simplifies the decision matrix for enterprises.
- Self-verification is the new differentiator. A model that catches its own errors, builds its own test harnesses, and iterates until it succeeds is fundamentally more valuable than a model that produces correct output on the first try.
- The open-weight gap may be widening. While Chinese labs dominate open-weight scale (Kimi K3 at 2.8T, Qwen3.8 at 2.4T), the fluid intelligence demonstrated by Opus 5 on ARC-AGI-3 represents a capability that may require more than just scale to replicate.
11. References & Resources
Official Sources
- Anthropic: Introducing Claude Opus 5 β Full announcement with benchmark charts
- Claude Platform Docs: What's New in Opus 5 β Technical specifications and API changes
- Claude Platform Docs: Prompting Claude Opus 5 β Behavioral differences and prompting patterns
- Claude Platform Docs: Pricing β Complete pricing table
- Claude Opus 5 System Card β Full benchmark tables and safety evaluations
- Claude Platform Docs: Effort β Effort level documentation
- Claude Platform Docs: Fast Mode β Fast mode documentation
- Claude Platform Docs: Refusals and Fallback β Handling refusals with default fallbacks
Key Analysis & Context
- Vellum: Claude Opus 5 Benchmarks Explained β Side-by-side benchmark analysis
- VentureBeat: Anthropic launches Opus 5 β Launch coverage
- TechCrunch: Anthropic launches Opus 5 β Launch coverage with customer quotes
- ARC Prize Foundation β ARC-AGI-3 benchmark details
Related Journal Articles
- Claude Fable 5 Mythos 5 Full Return Safeguards Jacobian Conjecture 2026 07 23 β Fable 5 benchmark context and export control saga
- Gemini 3 6 Flash 3 5 Flash Lite Cyber Token Efficiency Agentic Scale 2026 07 22 β Google's efficiency-focused release
- Qwen3 8 Max Preview 2 4t Multimodal Moe Open Weight Promise 2026 07 24 β Qwen3.8 open-weight context
- Thinking Machines Inkling 975b Multimodal Moe Self Improvement Controllable Effort 2026 07 21 β Open-weight model landscape
- Five Eyes Joint Warning Ai Cyber Threats Months Away 2026 06 25 β Cybersecurity AI governance context
- Howto Multi Model Routing Layer β Multi-model routing strategies for cost optimization
Report compiled July 27, 2026. All links verified against primary sources at time of publication.
π Referenced by
- π¬Meta Muse Spark 1.2 and Muse Code: Persistent Async Agents, Co-Trained Harness, and the $0.10/M Data-Share Pricing Play2026-08-07T00:00:00.000Z
- π¬DeepSeek V4-Flash-0731 Official Release: Agentic Coding at 99% Lower Cost, MIT License, and the New Floor for AI Inference Pricing2026-08-04T00:00:00.000Z
- π July 27: Claude Opus 5 Deep Dive & AI Weekly β The Week the Sandbox Broke2026-07-27T00:00:00.000Z
- πWiki Index2026-06-17T00:00:00.000Z
- πWiki Log2026-06-17T00:00:00.000Z