anthropic claude sonnet

Claude Sonnet 5 vs Opus 5 Coding Benchmarks (2026)

July 11, 2026 · Updated September 14, 2026 · HostAgentes Team · 7 min read

Anthropic’s current lineup is Opus 5 (flagship), Sonnet 5, Haiku 4.5, and Fable 5.1 for long-horizon agentic work. The most recently published agentic-coding benchmark scores, however, are from the previous generation: Sonnet 5 vs Opus 4.8. This post explains what those numbers still tell you, where the current lineup fits, and — because benchmark tables always lag the model list — how to measure the actual current models on your own agent in 48–72 hours.

TL;DR: Published scores say the Sonnet-5-vs-Opus tier split held last generation: Opus led the hardest coding tasks, Sonnet 5 won cost/performance and computer use. For the current lineup, default agents to Sonnet 5 ($2 / $10, permanent), route deep coding to Opus 5 (1M-token context), and let your own evals — not vendor tables — make the final call.

SWE-bench Pro — the last published agentic-coding scores

What it measures: SWE-bench Pro evaluates models on real-world software engineering tasks — reading a codebase, understanding an issue, and submitting a patch that passes hidden tests. It’s the closest widely-used public benchmark to a real coding agent’s day-to-day work.

Last published scores: Opus 4.8 at 69.2%, Sonnet 5 at 63.2%, Sonnet 4.6 at 58.1%.

What this still tells you:

  • The tier split within a generation: the Opus tier completed more hard, multi-file tasks without giving up mid-run, while Sonnet handled the routine tail nearly as well
  • Sonnet 5’s generational jump over Sonnet 4.6 (+5.1 pp) is why agents pinned to Opus “for safety margin” could mostly move down a tier
  • The ~6-point gap concentrated on the hardest tasks — simple tasks showed little difference

What it can’t tell you: how Opus 5 compares. It’s a new generation, and published SWE-bench-style numbers for it aren’t yet on Anthropic’s model overview (checked September 2026). Directionally, the flagship tier exists for the same reason it always has — hardest coding tasks — and Anthropic positions Opus 5 for exactly that.

OSWorld-Verified — the computer-use number

What it measures: OSWorld-Verified evaluates a model’s ability to drive a real computer environment — clicking, typing, navigating applications and browsers to complete a task, not just writing code in a sandbox.

Last published score: Sonnet 5 at 81.2%. Anthropic didn’t publish an Opus-tier number on this benchmark, and Sonnet 5 was explicitly positioned as the more agentic, tool-driving model of its generation.

What this means in your agent:

  • Browser-automation agents, form-filling agents, and any agent that operates a UI rather than an API have a proven default in Sonnet 5
  • This benchmark is the clearest evidence that “bigger model = better agent” doesn’t hold uniformly — the mid-tier model was the stronger pick for this task shape

GDPval-AA v2 — the knowledge-work number

What it measures: GDPval-AA v2 evaluates general knowledge-work tasks — the kind of judgment-heavy, non-coding work agents do in support, research, and operations roles.

Last published scores: Sonnet 5 at 1,618, Opus 4.8 at 1,615 — essentially a tie.

What this means in your agent:

  • For support, research-summary, and operations agents that aren’t doing deep coding, paying flagship-tier prices has bought no measurable quality advantage in the last two generations — and Sonnet 5’s price dropped to a permanent $2 / $10
  • Assume the pattern holds until your own evals show otherwise

Where does Opus 5 fit?

Anthropic’s current positioning (from its model overview, verified September 2026):

ModelAnthropic’s positioningPrice (in / out per 1M tokens)
Opus 5Complex agentic coding and enterprise work; 1M-token context$5 / $25
Sonnet 5Best combination of speed and intelligence; most workloads$2 / $10
Haiku 4.5Fastest model with near-frontier intelligence$1 / $5
Fable 5.1Demanding reasoning and long-horizon agentic worksee Anthropic’s docs

Note what’s absent: published per-benchmark scores for the new generation. That’s normal — tables trail launches — and it’s precisely why the measurement loop below, not the tables, should drive your per-agent decisions.

What the benchmarks don’t measure

Reliability under production prompts. Benchmarks use clean, researcher-tuned prompts. Your production prompts may carry baggage compensating for older model quirks — worth re-testing simplified prompts now.

Tool ecosystem quality. A stronger model with weak tools (bad code search, flaky test runner) still produces mediocre agent results. The model score is a ceiling, not a guarantee.

Real-world repo scale. SWE-bench Pro tasks are real but curated. Your production repo may have far more history and unfamiliar conventions than any benchmark set — expect benchmark deltas to be directional, not exact.

How to measure the current lineup in your own agent

Don’t trust vendor benchmarks blindly — especially across a generation gap. Run your own comparison over 48–72 hours:

  1. Fix a frozen task set — 50-100 recent tasks covering your real difficulty distribution.
  2. Run the set on your current model (baseline) — record pass/fail, tokens, latency, tool calls.
  3. Run the set on Sonnet 5 and, separately, Opus 5 — same recording.
  4. Compute the delta on pass rate, tokens per successful task, and tool call count.
  5. Decide per agent. Coding-heavy agents may justify Opus 5’s $5 / $25; most others won’t.

On Paperclip, this loop is a dashboard action: clone the agent, swap the model, replay tasks, compare metrics — no infra work.

Practical recommendation

  1. Today: Default new agents to Sonnet 5 — permanent $2 / $10 pricing and Anthropic’s own starting recommendation for most workloads.
  2. This week: Re-run your golden task set against Sonnet 5 for any agent still on Opus 4.x, and against Opus 5 for agents where coding depth is the bottleneck.
  3. Route accordingly: Opus 5 for complex agentic coding and enterprise work, Haiku 4.5 ($1 / $5) for high-volume simple tasks, Fable 5.1 for long-horizon reasoning where Opus 5 at higher effort falls short.

FAQ

Do the Sonnet 5 vs Opus 4.8 benchmark scores still apply? As directional context, mostly — they’re the most recent published scores, and tiers rarely regress between generations. But Opus 5 is now the flagship; treat published scores as a hypothesis to re-verify on your own workload.

Which Claude model is best for coding agents today? Anthropic positions Opus 5 for complex agentic coding and enterprise work with a 1M-token context. Sonnet 5 remains the cost/performance default for most agent workloads — verify the Opus 5 lift on your own task set before paying flagship prices.

Where are the published Opus 5 benchmark scores? Not on Anthropic’s model overview yet (checked September 2026). Benchmark tables lag launches — run the measurement loop above instead of waiting.

How does the Claude lineup compare to GPT-5.6 and Gemini 3.1? GPT-5.6 Sol targets the hardest reasoning problems and Gemini 3.1 Pro leads on context volume; Sonnet 5 leads on cost/performance and computer-use tasks. See the full cross-provider comparison for benchmark-by-benchmark numbers.

Can I run Claude models self-hosted? No, they are API-only. If self-hosting matters, open-weight alternatives trail significantly on agent benchmarks but give you full data control.


Related: Deploy Claude Opus 5 & Sonnet 5 Agents on Paperclip → · Migrate to the current Claude lineup → · Claude vs GPT-5.6 vs Gemini →

H

HostAgentes Team

Engineering & product

The HostAgentes team is part of ZUI TECHNOLOGY, S.L. — we build managed hosting for AI agents and write about the infrastructure, models and patterns we use ourselves.

About us →

Ready to deploy your agents?

Managed hosting from $3.99/mo. Zero headaches.

View plans