AI Potluck
Model components / Fine-tuned / chat models

Claude Opus 4.x

Anthropic

Anthropic's heavy-reasoning Opus tier, tuned for sustained agentic coding and long-horizon planning across large codebases. Opus 4 scored 72.5% on SWE-bench Verified at release (May 2025 era); the line has iterated through Opus 4.5 (80.9%), Opus 4.6 (80.8%), and Opus 4.7 (April 2026), which hits 87.6% SWE-bench Verified and 98.5% computer-use visual acuity. Opus 4.6 and 4.7 are tied at the top of LMArena (Elo ~1500). A leaked 'Claude Mythos Preview' tops SWE-bench at 93.9% and is presumed to be the next-generation Opus / Claude 5 candidate.

Anthropic Claude Opus 4.x flagship line: Opus 4 (May 22 2025) -> 4.5 (Nov 2025) -> 4.6 (Feb 2026) -> 4.7 (Apr 16 2026) -> 4.8 (May 28 2026, current frontier, model id claude-opus-4-8). Post-trained chat/agentic flagship, closed/API + Bedrock + Vertex + claude.ai. Note: Opus 4.7 is a frozen base_pretrained anchor; this record scores the Opus 4.x line in finetuned_chat. Verified via Anthropic Opus 4.8 + Claude 4 pages June 2026. Consolidated on 2026-07-29 from claude-opus-4-7; openness follows claude-opus-4-x, the release that currently governs. See sources/slug_aliases.yaml.

Openness

1 high confidence
1.0
weights
closed
data
closed
code
closed
license
Proprietary(API/Bedrock/Vertex/claude.ai)

No weights distributed in any form. Served only through the Anthropic API, Bedrock, Vertex and claude.ai, so 1/closed on every dimension. Opus 4.x governs as the current release; 4.7 scored identically, and nothing about the tier's openness turns on which release is read. Consolidated because Anthropic markets Claude Opus as the product and 4.7 and 4.8 are versions of it.

Adoption

4 medium confidence
4.0

Directional market-position estimate (not a measured usage figure). Opus is Anthropic's premium tier top-tier reach but lower volume than the Sonnet workhorse, so 4 rather than 5. Anthropic's launch post lists 28 enterprise testimonials (Stripe, Replit, Vercel, etc.); Opus flagships appear on public LLM marketplace leaderboards (e.g. OpenRouter). Overrides the prior frozen-anchor convention (held null), which systematically understated dominant closed incumbents and could mask the openness gap.

Capability

5 high confidence
5.0

Anthropic-reported: state-of-the-art on GDPval-AA and Finance Agent; CursorBench 70% (vs 58% for Opus 4.6); BigLaw Bench 90.9% at high effort; gains on SWE-bench Verified/Pro/Multilingual. Independent Gemini 3.5 Flash comparisons place Opus 4.7 in the frontier field. Vendor-reported.

Unchanged since 2026-07-29 (last edited, not re-checked)