Claude Sonnet 5 vs Opus 5 Coding Benchmarks (2026)
Anthropic’s current lineup is Opus 5 (flagship), Sonnet 5, Haiku 4.5, and Fable 5.1 for long-horizon agentic work. The most recently published agentic-coding benchmark scores, however, are from the previous generation: Sonnet 5 vs Opus 4.8. This post explains what those numbers still tell you, where the current lineup fits, and — because benchmark tables always lag the model list — how to measure the actual current models on your own agent in 48–72 hours.
TL;DR: Published scores say the Sonnet-5-vs-Opus tier split held last generation: Opus led the hardest coding tasks, Sonnet 5 won cost/performance and computer use. For the current lineup, default agents to Sonnet 5 ($2 / $10, permanent), route deep coding to Opus 5 (1M-token context), and let your own evals — not vendor tables — make the final call.
SWE-bench Pro — the last published agentic-coding scores
What it measures: SWE-bench Pro evaluates models on real-world software engineering tasks — reading a codebase, understanding an issue, and submitting a patch that passes hidden tests. It’s the closest widely-used public benchmark to a real coding agent’s day-to-day work.
Last published scores: Opus 4.8 at 69.2%, Sonnet 5 at 63.2%, Sonnet 4.6 at 58.1%.
What this still tells you:
- The tier split within a generation: the Opus tier completed more hard, multi-file tasks without giving up mid-run, while Sonnet handled the routine tail nearly as well
- Sonnet 5’s generational jump over Sonnet 4.6 (+5.1 pp) is why agents pinned to Opus “for safety margin” could mostly move down a tier
- The ~6-point gap concentrated on the hardest tasks — simple tasks showed little difference
What it can’t tell you: how Opus 5 compares. It’s a new generation, and published SWE-bench-style numbers for it aren’t yet on Anthropic’s model overview (checked September 2026). Directionally, the flagship tier exists for the same reason it always has — hardest coding tasks — and Anthropic positions Opus 5 for exactly that.
OSWorld-Verified — the computer-use number
What it measures: OSWorld-Verified evaluates a model’s ability to drive a real computer environment — clicking, typing, navigating applications and browsers to complete a task, not just writing code in a sandbox.
Last published score: Sonnet 5 at 81.2%. Anthropic didn’t publish an Opus-tier number on this benchmark, and Sonnet 5 was explicitly positioned as the more agentic, tool-driving model of its generation.
What this means in your agent:
- Browser-automation agents, form-filling agents, and any agent that operates a UI rather than an API have a proven default in Sonnet 5
- This benchmark is the clearest evidence that “bigger model = better agent” doesn’t hold uniformly — the mid-tier model was the stronger pick for this task shape
GDPval-AA v2 — the knowledge-work number
What it measures: GDPval-AA v2 evaluates general knowledge-work tasks — the kind of judgment-heavy, non-coding work agents do in support, research, and operations roles.
Last published scores: Sonnet 5 at 1,618, Opus 4.8 at 1,615 — essentially a tie.
What this means in your agent:
- For support, research-summary, and operations agents that aren’t doing deep coding, paying flagship-tier prices has bought no measurable quality advantage in the last two generations — and Sonnet 5’s price dropped to a permanent $2 / $10
- Assume the pattern holds until your own evals show otherwise
Where does Opus 5 fit?
Anthropic’s current positioning (from its model overview, verified September 2026):
| Model | Anthropic’s positioning | Price (in / out per 1M tokens) |
|---|---|---|
| Opus 5 | Complex agentic coding and enterprise work; 1M-token context | $5 / $25 |
| Sonnet 5 | Best combination of speed and intelligence; most workloads | $2 / $10 |
| Haiku 4.5 | Fastest model with near-frontier intelligence | $1 / $5 |
| Fable 5.1 | Demanding reasoning and long-horizon agentic work | see Anthropic’s docs |
Note what’s absent: published per-benchmark scores for the new generation. That’s normal — tables trail launches — and it’s precisely why the measurement loop below, not the tables, should drive your per-agent decisions.
What the benchmarks don’t measure
Reliability under production prompts. Benchmarks use clean, researcher-tuned prompts. Your production prompts may carry baggage compensating for older model quirks — worth re-testing simplified prompts now.
Tool ecosystem quality. A stronger model with weak tools (bad code search, flaky test runner) still produces mediocre agent results. The model score is a ceiling, not a guarantee.
Real-world repo scale. SWE-bench Pro tasks are real but curated. Your production repo may have far more history and unfamiliar conventions than any benchmark set — expect benchmark deltas to be directional, not exact.
How to measure the current lineup in your own agent
Don’t trust vendor benchmarks blindly — especially across a generation gap. Run your own comparison over 48–72 hours:
- Fix a frozen task set — 50-100 recent tasks covering your real difficulty distribution.
- Run the set on your current model (baseline) — record pass/fail, tokens, latency, tool calls.
- Run the set on Sonnet 5 and, separately, Opus 5 — same recording.
- Compute the delta on pass rate, tokens per successful task, and tool call count.
- Decide per agent. Coding-heavy agents may justify Opus 5’s $5 / $25; most others won’t.
On Paperclip, this loop is a dashboard action: clone the agent, swap the model, replay tasks, compare metrics — no infra work.
Practical recommendation
- Today: Default new agents to Sonnet 5 — permanent $2 / $10 pricing and Anthropic’s own starting recommendation for most workloads.
- This week: Re-run your golden task set against Sonnet 5 for any agent still on Opus 4.x, and against Opus 5 for agents where coding depth is the bottleneck.
- Route accordingly: Opus 5 for complex agentic coding and enterprise work, Haiku 4.5 ($1 / $5) for high-volume simple tasks, Fable 5.1 for long-horizon reasoning where Opus 5 at higher effort falls short.
FAQ
Do the Sonnet 5 vs Opus 4.8 benchmark scores still apply? As directional context, mostly — they’re the most recent published scores, and tiers rarely regress between generations. But Opus 5 is now the flagship; treat published scores as a hypothesis to re-verify on your own workload.
Which Claude model is best for coding agents today? Anthropic positions Opus 5 for complex agentic coding and enterprise work with a 1M-token context. Sonnet 5 remains the cost/performance default for most agent workloads — verify the Opus 5 lift on your own task set before paying flagship prices.
Where are the published Opus 5 benchmark scores? Not on Anthropic’s model overview yet (checked September 2026). Benchmark tables lag launches — run the measurement loop above instead of waiting.
How does the Claude lineup compare to GPT-5.6 and Gemini 3.1? GPT-5.6 Sol targets the hardest reasoning problems and Gemini 3.1 Pro leads on context volume; Sonnet 5 leads on cost/performance and computer-use tasks. See the full cross-provider comparison for benchmark-by-benchmark numbers.
Can I run Claude models self-hosted? No, they are API-only. If self-hosting matters, open-weight alternatives trail significantly on agent benchmarks but give you full data control.
Related: Deploy Claude Opus 5 & Sonnet 5 Agents on Paperclip → · Migrate to the current Claude lineup → · Claude vs GPT-5.6 vs Gemini →
HostAgentes Team
Engineering & product
The HostAgentes team is part of ZUI TECHNOLOGY, S.L. — we build managed hosting for AI agents and write about the infrastructure, models and patterns we use ourselves.
About us →Related articles
Deploy Claude Opus 5 on Paperclip: Model IDs & Pricing
Claude Opus 5 on Paperclip: model ID claude-opus-5, $5/$25 per 1M tokens, 1M context, no code changes. Sonnet 5 from $2/$10. Anthropic docs verified Sep 2026.
Claude vs GPT-5.6 vs Gemini for AI Agents
Head-to-head: Anthropic Claude Sonnet 5 / Opus 5, OpenAI GPT-5.6 Sol, and Google Gemini 3.1 Pro on coding, context, price, and agent workflows.
Migrate to Claude Opus 5 or Sonnet 5: Complete Guide (2026)
Step-by-step guide to migrating production AI agents to Claude Opus 5 or Sonnet 5 — plus how to switch back in 30 seconds if quality dips.