Claude Sonnet 5: Anthropic's Most Agentic Mid-Tier Model
Claude Sonnet 5: Anthropic's Most Agentic Mid-Tier Model
Published: June 30, 2026 Source: Anthropic Official Announcement
Anthropic launched Claude Sonnet 5 on June 30, 2026, positioning it as the "most agentic Sonnet model yet." The model represents a substantial upgrade over Sonnet 4.6, narrowing the capability gap to Opus 4.8 while maintaining Sonnet-tier pricing.
Key Specifications
| Attribute | Detail |
|---|---|
| Context Window | 1,000,000 tokens |
| Maximum Output | 128,000 tokens |
| Introductory Pricing | $2/M input, $10/M output (through Aug 31, 2026) |
| Standard Pricing | $3/M input, $15/M output (from Sep 1, 2026) |
| API Identifier | claude-sonnet-5 |
| Tokenizer | Updated tokenizer (similar change introduced with Opus 4.7) |
| Default Model | Yes — for Free and Pro plans |
Benchmark Performance
All benchmark figures sourced from the Claude Sonnet 5 System Card published by Anthropic.
Coding: SWE-Bench Pro
| Model | Score |
|---|---|
| Opus 4.8 | 69.2% |
| Sonnet 5 | 63.2% |
| Sonnet 4.6 | 58.1% |
Sonnet 5 achieves a 5.1-point improvement over Sonnet 4.6, closing much of the gap to Opus 4.8.
Terminal Work: Terminal-Bench 2.1
| Model | Score |
|---|---|
| Sonnet 5 | 80.4% |
| Opus 4.8 | 74.6% |
| Sonnet 4.6 | 67.0% |
Sonnet 5 surpasses Opus 4.8 on terminal tasks — the first benchmark in the suite where the mid-tier model outperforms the flagship.
Hard Reasoning: Humanity's Last Exam (with tools)
| Model | Score |
|---|---|
| Opus 4.8 | 57.9% |
| Sonnet 5 | 57.4% |
| Sonnet 4.6 | 46.8% |
Sonnet 5 is essentially tied with Opus 4.8 on hard reasoning with tool use, representing a 10.6-point jump over Sonnet 4.6.
Computer Use: OSWorld-Verified
| Model | Score |
|---|---|
| Opus 4.8 | 83.4% |
| Sonnet 5 | 81.2% |
| Sonnet 4.6 | 78.5% |
Sonnet 5 narrows the gap to Opus 4.8 to just 2.2 points on computer use tasks.
Agentic Search: BrowseComp
| Model | Score |
|---|---|
| Sonnet 5 | 84.7% |
| Opus 4.8 | 83.4% |
| Sonnet 4.6 | 73.1% |
Sonnet 5 leads on agentic search, outperforming both Opus 4.8 and Sonnet 4.6.
Knowledge Work: GDPval-AA v2
| Model | Score |
|---|---|
| Sonnet 5 | 78.5% |
| Opus 4.8 | 77.2% |
| Sonnet 4.6 | 65.3% |
Sonnet 5 exceeds Opus 4.8 on knowledge work tasks, marking the second benchmark where it leads the flagship.
Safety Assessments
Anthropic's pre-deployment safety evaluations found Sonnet 5 to be an overall improvement over Sonnet 4.6:
- Lower rate of undesirable behaviors than Sonnet 4.6 across the automated behavioral audit
- Better at refusing malicious requests and resisting prompt injection hijack attempts
- Lower rates of hallucination and sycophancy than Sonnet 4.6
- Substantially poorer cybersecurity capabilities than Opus 4.8 and Mythos 5 — Sonnet 5 scored 0.0% on developing working exploits for Firefox vulnerabilities (same as Sonnet 4.6)
- Cyber safeguards enabled by default — same safeguards as Opus 4.7 and 4.8, less strict than those on Fable 5
However, Sonnet 5 shows a slightly higher rate of partial success on cybersecurity tasks than Sonnet 4.6, attributed to improvements in general intelligence rather than specific training.
Effort Levels
Sonnet 5 supports adjustable effort levels (low through extra-high), enabling users to trade cost for performance:
- At medium effort, Sonnet 5 provides substantially improved cost efficiency
- At extra-high effort, Sonnet 5's performance on OSWorld-Verified and BrowseComp approaches Opus 4.8's medium-to-high setting
- Users can select between Sonnet 5 and Opus 4.8 at different effort levels to find the optimal cost-performance balance
Availability
Claude Sonnet 5 is available across all Anthropic platforms:
- Claude Chat — default model for Free and Pro plans
- Claude Cowork — available to all plan tiers
- Claude Code — available for agentic coding workflows
- Claude Platform (API) — available via
claude-sonnet-5identifier - Max, Team, and Enterprise plans — available alongside other models
Significance
Sonnet 5 represents a structural shift in Anthropic's model lineup. For the first time, a mid-tier Sonnet model outperforms the flagship Opus on specific benchmarks (Terminal-Bench 2.1 and GDPval-AA v2), while matching or nearly matching Opus 4.8 across coding, reasoning, computer use, and agentic search.
The pricing strategy — introductory rates of $2/$10 per million tokens, rising to $3/$15 after August 31 — positions Sonnet 5 as the default choice for most agentic workloads, with Opus 4.8 reserved for tasks requiring the last few accuracy points.
The updated tokenizer (similar to the change introduced with Opus 4.7) increases token counts by roughly 1.0–1.35× depending on content type, which Anthropic offsets with the introductory pricing to make the transition roughly cost-neutral for users migrating from Sonnet 4.6.
References
- Anthropic. (2026, June 30). Introducing Claude Sonnet 5. https://www.anthropic.com/news/claude-sonnet-5
- Anthropic. (2026, June 30). Claude Sonnet 5 System Card. https://www.anthropic.com/claude-sonnet-5-system-card
- Anthropic. (2026, June 30). Release Notes: Claude Sonnet 5 launch. https://docs.anthropic.com/en/release-notes/claude-apps
- Anthropic. (2026). Claude Platform — Models Overview. https://platform.claude.com/docs/en/about-claude/models/overview
- Anthropic. (2026). Real-time Cyber Safeguards on Claude. https://support.claude.com/en/articles/14604842-real-time-cyber-safeguards-on-claude
🔗 Referenced by
- 🔬DeepSeek V4 Flash & Pro: API Migration Deadline, Hybrid Attention Architecture, and the $0.14/M Token Price Floor2026-07-10T00:00:00.000Z
- 🔬GPT-5.6 Public Launch: Sol, Terra, Luna Go Global with Ultra Mode, 750 TPS on Cerebras, and the Most Robust Cyber Safeguards Yet2026-07-09T00:00:00.000Z
- 🔬Meta's Muse Ecosystem: Muse Image Launch, Superintelligence Labs, and the Watermelon Model2026-07-08T00:00:00.000Z
- 🔬Claude Science: Anthropic's AI Workbench for Drug Discovery and Biomedical Research2026-07-07T00:00:00.000Z
- 🔬GPT-5.6 Sol, Terra, and Luna: OpenAI's Subagent Era, Ultra Mode, and the Most Robust Safety Stack Yet2026-07-06T00:00:00.000Z