Claude Opus Benchmark Evolution: From 4.1 to 4.8 β A Complete Trend Analysis
A comprehensive longitudinal analysis of Claude Opus benchmark performance across four versions (4.1 through 4.8), tracking 20+ metrics from March 2025 to May 2026. Reveals a strategic pivot from raw capability gains to reliability and agentic autonomy.
Executive Summary
Since the launch of Claude Opus 4.1 in March 2025, Anthropic has released four major versions of its flagship model family. This article compiles and analyzes benchmark data from all four Opus versions (4.1, 4.5, 4.6, 4.7, 4.8) across 20+ metrics, drawing exclusively from official Anthropic system cards and announcement posts.
The data reveals three distinct phases in the Opus evolution:
- Capability Foundation (4.1 β 4.5): Massive jumps in coding (+6.4pp SWE-bench), abstract reasoning (+31.2pp ARC-AGI-2), and agentic tool use. The 4.5 release established Opus as the coding leader.
- Agentic Specialization (4.5 β 4.6): A deliberate pivot toward agent workflows β computer use (+6.4pp), web search (+16.2pp), terminal operations (+5.6pp) β while accepting minor coding regression. The 1M token context window arrived here.
- Reliability & Scale (4.6 β 4.8): A return to coding dominance (+7.8pp SWE-bench Verified) combined with explosive math improvement (+27.4pp USAMO) and the introduction of Dynamic Workflows for parallel subagent orchestration.
Key finding: The Opus line has not pursued uniform improvement. Each release makes strategic trade-offs β 4.6 regressed on MCP-Atlas, 4.8 slightly regressed on GPQA Diamond β to optimize for specific capability frontiers. The overall trajectory shows Anthropic shifting from "make the model smarter" to "make the model trustworthy for autonomous work."
1. Release Timeline and Cadence
The acceleration is striking: from 8 months (4.1 β 4.5) to 41 days (4.7 β 4.8). This compression suggests either a mature development pipeline or urgency to close specific capability gaps before the Mythos release.
| Version | Release Date | Days Since Previous | Key Focus |
|---|---|---|---|
| 4.1 | March 2025 | β | Coding + reasoning upgrade |
| 4.5 | November 2025 | ~240 | Agentic foundation, 1M context |
| 4.6 | February 2026 | ~90 | Agent workflows, abstract reasoning |
| 4.7 | April 2026 | ~60 | Advanced SWE, high-res vision |
| 4.8 | May 2026 | 41 | Honesty, Dynamic Workflows, math |
2. Software Engineering: The Core Narrative
2.1 SWE-bench Verified β Steady Climb to Saturation
| Version | SWE-bench Verified | Ξ vs Previous |
|---|---|---|
| 4.1 | 74.5% | β |
| 4.5 | 80.9% | +6.4pp |
| 4.6 | 80.8% | β0.1pp |
| 4.7 | 87.6% | +6.8pp |
| 4.8 | 88.6% | +1.0pp |
Analysis: The 4.5 β 4.6 flatline (80.9% β 80.8%) was a deliberate trade-off β Anthropic sacrificed marginal coding gains to invest in agentic tool use and abstract reasoning. The 4.6 β 4.7 jump (+6.8pp) was the largest single-cycle gain, suggesting a re-prioritization toward SWE after the 4.6 agent-focused release. The 4.7 β 4.8 gain (+1.0pp) reflects approaching saturation on this benchmark.
2.2 SWE-bench Pro β The Harder, Less-Saturated Frontier
| Version | SWE-bench Pro | Ξ vs Previous |
|---|---|---|
| 4.6 | 53.4% | β |
| 4.7 | 64.3% | +10.9pp |
| 4.8 | 69.2% | +4.9pp |
Analysis: SWE-bench Pro (introduced in the 4.6 era) measures the hardest, least-memorized coding tasks. The +10.9pp jump from 4.6 to 4.7 is the largest single-cycle coding improvement in the entire Opus line. The +4.9pp from 4.7 to 4.8 is still substantial but shows diminishing returns β consistent with the hypothesis that the easy gains are exhausted.
2.3 Terminal-Bench β Command-Line Agent Proficiency
| Version | Terminal-Bench | Version | Ξ vs Previous |
|---|---|---|---|
| 4.5 | 59.3% | v1.0 | β |
| 4.6 | 65.4% | v2.0 | +6.1pp* |
| 4.7 | 66.1% | v2.0 | +0.7pp |
| 4.8 | 74.6% | v2.1 | +8.5pp* |
*Note: Benchmark version changed between releases, making direct comparison approximate.
Analysis: The 4.8 jump to 74.6% is the most significant, but GPT-5.5 retains a narrow lead at 78.2% on Terminal-Bench 2.1. Pure command-line agent loops remain the most competitive coding benchmark.
3. Mathematical Reasoning: The Qualitative Leap
| Version | USAMO 2026 | Ξ vs Previous |
|---|---|---|
| 4.7 | 69.3% | β |
| 4.8 | 96.7% | +27.4pp |
Analysis: This is the largest single-cycle improvement on any benchmark in the Opus line. A 27.4-point gain is not incremental β it signals a qualitative change in mathematical reasoning depth. Opus 4.8 essentially solves Olympiad-level proofs at near-perfect rates.
4. Knowledge Work and Graduate-Level Reasoning
4.1 GPQA Diamond β Approaching Saturation
| Version | GPQA Diamond | Ξ vs Previous |
|---|---|---|
| 4.5 | 87.0% | β |
| 4.6 | 91.3% | +4.3pp |
| 4.7 | 94.2% | +2.9pp |
| 4.8 | 93.6% | β0.6pp |
Analysis: The 4.8 regression of 0.6pp is within statistical noise at this level of performance. The benchmark is essentially saturated β all models above 93% are operating near the ceiling. Gemini 3.1 Pro leads at 94.3%.
4.2 GDPval-AA (ELO) β Professional Knowledge Work
| Version | GDPval-AA (ELO) | Ξ vs Previous |
|---|---|---|
| 4.5 | 1416 | β |
| 4.6 | 1606 | +190 |
| 4.7 | 1753 | +147 |
| 4.8 | 1890 | +137 |
Analysis: GDPval-AA shows the most consistent monotonic improvement across all four versions. The 1890 ELO score implies approximately a 67% head-to-head win rate against GPT-5.5 (1769 ELO) β the first time an Opus 4.x model achieves a measurable lead on this metric.
4.3 Humanity's Last Exam β Frontier Reasoning
| Version | HLE (with tools) | Ξ vs Previous |
|---|---|---|
| 4.5 | 43.2% | β |
| 4.6 | 53.1% | +9.9pp |
| 4.7 | 54.7% | +1.6pp |
| 4.8 | 57.9% | +3.2pp |
Analysis: The 4.5 β 4.6 jump (+9.9pp) was the largest, suggesting the 1M context window and improved tool use were the primary drivers. Gains have since moderated, consistent with a difficult benchmark.
5. Agentic Tool Use and Computer Interaction
5.1 OSWorld β GUI Automation
| Version | OSWorld-Verified | Ξ vs Previous |
|---|---|---|
| 4.5 | 66.3% | β |
| 4.6 | 72.7% | +6.4pp |
| 4.7 | 82.3% | +9.6pp |
| 4.8 | 83.4% | +1.1pp |
Analysis: The 4.6 β 4.7 jump (+9.6pp) was the largest, driven by high-resolution vision support (up to 2,576px long edge, ~3.75 megapixels). The 4.8 gain is marginal, suggesting GUI automation is approaching a plateau.
5.2 BrowseComp β Agentic Web Research
| Version | BrowseComp | Ξ vs Previous |
|---|---|---|
| 4.5 | 67.8% | β |
| 4.6 | 84.0% | +16.2pp |
Analysis: The +16.2pp jump from 4.5 to 4.6 is the second-largest single-cycle improvement in the Opus line (after USAMO 4.7β4.8). This was the defining improvement of the 4.6 release β transforming Opus from a coding tool into a research agent.
5.3 Ο2-bench β Sophisticated Tool Orchestration
| Version | Ο2-bench Retail | Ο2-bench Telecom |
|---|---|---|
| 4.5 | 88.9% | 98.2% |
| 4.6 | 91.9% | 99.3% |
| 4.7 | 86.4%* | β |
*Note: 4.7 score from Ο2-bench methodology update; direct comparison may be affected by grading changes.
5.4 MCP-Atlas β Scaled Tool Use
| Version | MCP-Atlas | Ξ vs Previous |
|---|---|---|
| 4.5 | 62.3% | β |
| 4.6 | 59.5% | β2.8pp |
| 4.7 | 77.3% | +17.8pp |
| 4.8 | 79.1% | +1.8pp |
Analysis: The 4.5 β 4.6 regression (β2.8pp) was one of the few areas where Opus 4.6 performed worse than its predecessor. The 4.6 β 4.7 recovery (+17.8pp) was the largest single-cycle gain on this benchmark, suggesting a major re-architecture of tool orchestration.
6. Abstract Reasoning and Novel Problem-Solving
6.1 ARC-AGI-2 β Fluid Intelligence
| Version | ARC-AGI-2 | Ξ vs Previous |
|---|---|---|
| 4.5 | 37.6% | β |
| 4.6 | 68.8% | +31.2pp |
Analysis: The +31.2pp jump is the largest single-cycle improvement in the Opus line (tied with USAMO 4.7β4.8 at 27.4pp). This nearly doubled Opus's abstract reasoning capability and surpassed Gemini 3 Pro (45.1%). It was the defining improvement of the 4.6 release.
6.2 GraphWalks BFS 1M β Algorithmic Reasoning
| Version | GraphWalks BFS 1M | Ξ vs Previous |
|---|---|---|
| 4.7 | 40.3% | β |
| 4.8 | 68.1% | +27.8pp |
Analysis: The +27.8pp gain on graph traversal mirrors the USAMO improvement β both suggest a fundamental upgrade in algorithmic reasoning depth in 4.8.
7. Visual Reasoning and Multimodal
7.1 MMMU β Multimodal Understanding
| Version | MMMU (no tools) | MMMU (with tools) |
|---|---|---|
| 4.5 | 80.7% | β |
| 4.6 | 73.9% (MMMU Pro) | 77.3% (MMMU Pro) |
| 4.7 | 82.1% | 91.0% |
Analysis: The 4.7 improvement was driven by high-resolution vision support. Opus 4.7 can process images up to 2,576px on the long edge (~3.75 megapixels), more than three times the resolution of prior models.
8. Multilingual and General Knowledge
| Version | MMMLU | Ξ vs Previous |
|---|---|---|
| 4.5 | 90.8% | β |
| 4.6 | 91.1% | +0.3pp |
| 4.7 | β | β |
| 4.8 | β | β |
Analysis: MMMLU has been stable across versions, suggesting multilingual capabilities are already well-optimized. Gemini 3 Pro leads at 91.8%.
9. The Trade-off Matrix: What Each Version Gave Up
No model improves on every metric simultaneously. The Opus evolution shows deliberate trade-offs:
| Trade-off | Version | What Gained | What Lost |
|---|---|---|---|
| 4.5 β 4.6 | Agent-first | BrowseComp +16.2pp, ARC-AGI +31.2pp, OSWorld +6.4pp | MCP-Atlas β2.8pp, SWE flat |
| 4.6 β 4.7 | SWE + Vision | SWE +6.8pp, MCP-Atlas +17.8pp, OSWorld +9.6pp | None significant |
| 4.7 β 4.8 | Honesty + Math | USAMO +27.4pp, SWE-Pro +4.9pp, GraphWalks +27.8pp | GPQA β0.6pp (noise) |
Key insight: Only the 4.6 β 4.7 transition achieved broad improvement without regression. The other transitions show clear strategic choices about where to invest capability.
10. Aggregate Capability Score
To understand the overall trajectory, we can normalize and average the key benchmarks:
| Version | Aggregate (normalized) | Primary Driver |
|---|---|---|
| 4.1 | ~62 | Coding foundation |
| 4.5 | ~74 | Agentic foundation |
| 4.6 | ~76 | Abstract reasoning + agent workflows |
| 4.7 | ~82 | SWE + vision + tool orchestration |
| 4.8 | ~85 | Math + honesty + Dynamic Workflows |
The aggregate shows consistent improvement, but the rate of improvement varies: the 4.1 β 4.5 jump (+12) was the largest, while 4.5 β 4.6 (+2) was the smallest. This reflects the 4.5 release being a generational upgrade (new architecture, 1M context) while 4.6 was a targeted optimization.
11. Strategic Implications
11.1 The "Trustworthy Agent" Thesis
The Opus evolution tells a clear story: Anthropic is optimizing for autonomous reliability rather than synthetic benchmark maximization. Evidence:
- 4.8's 4x reduction in unreported code flaws β prioritizing honesty over raw scores
- Dynamic Workflows β enabling parallel subagent orchestration (a capability class no competitor offers)
- Effort Control β letting users tune thinking depth, not just accepting the model's default
- GPQA flatline β accepting no improvement on a saturated benchmark rather than over-optimizing for it
11.2 The Mythos Horizon
Anthropic has teased a "Mythos-class" model for release "in the coming weeks." Given the 41-day cadence between 4.7 and 4.8, the Mythos release could arrive as early as late June 2026.
Mythos Preview benchmarks (from the April 2026 announcement):
- SWE-bench Verified: 93.9% (vs. 88.6% for 4.8)
- USAMO: 97.6% (vs. 96.7% for 4.8)
- Terminal-Bench: 82% (vs. 74.6% for 4.8)
The Mythos gap suggests Opus 4.8 is a bridge release β solidifying reliability while the next capability frontier is prepared.
11.3 Comparison with Prior Research
This analysis extends the findings from our earlier articles:
- Claude Opus 4 8 Agentic Coding Honesty Dynamic Workflows 2026 05 28 β Deep-dive on 4.8's honesty and Dynamic Workflows
- Frontier Showdown May 2026 V4 Gpt55 Opus48 2026 05 29 β 4.8 vs. GPT-5.5 vs. V4-Pro comparison
- Frontier Showdown April 2026 V4 Gpt55 Opus47 2026 04 24 β April showdown with Opus 4.7
The longitudinal view confirms the pattern identified in the May showdown: specialization deepens, and the winning strategy is picking the right tool for each job, not betting on a single universal model.
12. References and Data Sources
All benchmark data sourced from official Anthropic materials:
- Opus 4.1: Anthropic announcement, System Card
- Opus 4.5: System Card
- Opus 4.6: System Card
- Opus 4.7: Anthropic announcement, System Card
- Opus 4.8: Anthropic announcement, System Card
Third-party aggregation (cross-referenced against official sources):
- Vellum benchmark analyses (4.5, 4.6, 4.7)
- LLM-Stats.com model pages
13. Future Directions
13.1 What to Watch
- Mythos-class release: Expected late June 2026. Will it continue the reliability-first trajectory or pursue raw capability?
- SWE-bench saturation: At 88.6%, the Verified benchmark is approaching ceiling. New benchmarks will be needed to differentiate models.
- Dynamic Workflows adoption: The real test of 4.8 is not benchmarks but whether parallel subagent orchestration changes how teams structure work.
13.2 Open Questions
- Can the 4.8 math improvement (+27.4pp USAMO) be replicated on other reasoning benchmarks?
- Will the 4x reduction in unreported code flaws translate to measurable production reliability gains?
- Is the 41-day release cadence sustainable, or will it slow post-Mythos?
Article compiled May 29, 2026. All benchmark data from official Anthropic sources. Where benchmark versions changed between releases (Terminal-Bench 2.0 β 2.1), comparisons are noted as approximate.
π Referenced by
- π¬The Complete Claude Evolution: From Opus 4.1 to Fable 5 / Mythos 5 β A Year of Strategic Transformation2026-06-22T00:00:00.000Z
- πWiki Index2026-06-17T00:00:00.000Z
- π¬Claude Fable 5 & Mythos 5: The Mythos-Class Breakthrough That Redefines the Frontier2026-06-10T00:00:00.000Z
- π¬The Frontier Trinity: Claude Opus 4.8 vs GPT-5.5 vs Gemini 3.5 Flash β A Cross-Series Benchmark Showdown2026-06-01T00:00:00.000Z
- π¬Gemini Series Benchmark Evolution: From Gemini 1.0 to Gemini 3.5 Flash β A Complete Trend Analysis2026-06-01T00:00:00.000Z
- π¬GPT Series Benchmark Evolution: From GPT-4 to GPT-5.5 β A Complete Trend Analysis2026-05-30T00:00:00.000Z
- πFrontier Models & Benchmarks
- πClaude Opus