Claude Sonnet
AnthropicAnthropic's mid-tier Claude model and its production default for coding, positioned between Haiku (small/fast) and Opus (heavy reasoning). Tracked as a tier rather than a point release, so the generation currently shipping is scored and the measured release is named in the score note. Generations to date are Sonnet 4 (May 2025), Sonnet 4.5 (Sept 2025), Sonnet 4.6 (Feb 2026) and Sonnet 5 (June 2026). Closed weights, available through the Anthropic API and the Claude apps.
Tier product. Anthropic ships a new Sonnet roughly every few months, so a versioned entry goes stale faster than it can be reviewed; the slug is stable and the score carries the release it was measured against. Supersedes the separate claude-sonnet-4 and claude-sonnet-4-6 entries. The scores are still measured against Sonnet 4.6 while Sonnet 5 (2026-06-30) is shipping; capability is flagged for rescoring on it. The headline SWE-bench figure sits in a client-rendered chart on the 4.6 launch page and is not re-derivable from a fetch. Verified 2026-08-13 via the Sonnet 4.6 and Sonnet 5 launch posts.
Openness
1 high confidence- weights
- closed
- data
- closed
- code
- closed
- license
- Proprietary(API-only)
Closed by construction and stable across generations - no Sonnet release has ever shipped weights, so this axis does not move when the tier advances. Measured against Claude Sonnet 4.6 (released 2026-02-17), Anthropic's production default coding tier at the time of scoring. The tier has since advanced to Sonnet 5, sold the same way at $2 per million input tokens and $10 per million output tokens, so the axis reads closed on both releases.
- https://www.anthropic.com/news/claude-sonnet-4-6 recorded 2026-08-13
"Claude Sonnet 4.6 is available now on all Claude plans, Claude Cowork, Claude Code, our API, and all major cloud platforms"; access is by API model id claude-sonnet-4-6, with no weights, corpus or training code offered
- https://www.anthropic.com/news/claude-sonnet-5 recorded 2026-08-13
the tier's current release - "Introducing Claude Sonnet 5, Jun 30, 2026"; "priced at $2 per million input tokens and $10 per million output tokens. Developers can use claude-sonnet-5 via the Claude API"; no weights, corpus or training code offered
Adoption
5 medium confidenceThe highest-adoption tier of the three Claude model tiers - Sonnet is the production default for coding and the model most API traffic lands on. Directional rather than measured: Anthropic publishes no per-model usage figure, so this is a judgment from reported traction and no reach band is recorded. The current Sonnet 5 release is the default model on the Free and Pro plans, which makes the breadth argument stronger rather than weaker.
- https://www.anthropic.com/news/claude-sonnet-4-6 recorded 2026-08-13
"We've also upgraded our free tier to Sonnet 4.6 by default"; available on all Claude plans, Claude Cowork, Claude Code, the API and all major cloud platforms; no usage or user figure published
- https://www.anthropic.com/news/claude-sonnet-5 recorded 2026-08-13
"From today, Claude Sonnet 5 is available across all plans - it is the default model for Free and Pro plans, and is available to Max, Team, and Enterprise users"; no usage or user figure published
Capability
4 medium confidenceA tier is scored on its strongest release, and the governing release here is Claude Sonnet 5 (released 2026-06-30). The band is 4, and Anthropic's own framing is why. The launch post says "Sonnet 5 narrows the gap - its performance is close to that of Opus 4.8, but at lower prices", calls Opus 4.8 "a more generally capable model, for reference", and describes Sonnet 5 as "approaching Opus 4.8's capability levels" while outperforming it only "on some tasks". Close to, approaching, and better on some tasks is not level with: this tier sits below the Opus and Fable tiers, which sit at 5. The evidence for a 5 would have to be the vendor placing Sonnet at or above its own flagship, and the vendor does the opposite. Two limitations are worth keeping. Sonnet 4.6's headline 79.6% on SWE-bench Verified sits in a chart that renders in the browser and cannot be read back from the page, and the Sonnet 5 post publishes no SWE-bench figure in its readable text at all - only OSWorld-Verified 78.5%, 46.8% and 34.6% appear as plain text - so the benchmark named here is the one the judgment tracks rather than a number the page states. And SWE-bench Verified's public leaderboard has had no submissions since 2025-12-15, so recent figures are vendor-reported either way. Confidence stays medium for both reasons.
- https://www.anthropic.com/news/claude-sonnet-5 recorded 2026-08-14
the tier's governing release - "Sonnet 5 narrows the gap: its performance is close to that of Opus 4.8, but at lower prices. It's a substantial improvement over its predecessor, Sonnet 4.6, on important aspects of agentic performance like reasoning, tool use, coding, and knowledge work." Opus 4.8 is described as "a more generally capable model, for reference"; Sonnet 5 is "approaching Opus 4.8's capability levels" and outperforms it "on some tasks". OSWorld-Verified figures of 78.5%, 46.8% and 34.6% are the only benchmark numbers in the fetchable body; no SWE-bench figure appears.
- https://www.anthropic.com/news/claude-sonnet-4-6 recorded 2026-08-13
the superseded release the old score read. SWE-bench Verified footnote - "Our score was averaged over 10 trials. With a prompt modification, we saw a score of 80.2%". The headline 79.6% is in a client-rendered chart and is not in the fetched body.
Verified 2026-08-13