AI Potluck
Model components / Fine-tuned / chat models

Claude Opus 4.x

Anthropic

Anthropic's heavy-reasoning Opus tier, tuned for sustained agentic coding and long-horizon planning across large codebases. The line has iterated from Opus 4 in May 2025 through 4.5, 4.6 and 4.7 to Opus 4.8, the current frontier release, served through the API, Bedrock, Vertex and claude.ai.

Opus 4.7 is a frozen base_pretrained anchor; this record scores the Opus 4.x line in finetuned_chat. An earlier capability note cited a Gemini comparison that names neither Claude nor Opus, and was withdrawn. Verified 2026-08-13 via the Opus 4.8, Opus 4.7 and Claude 4 announcements.

Openness

1 high confidence
1.0
weights
closed
data
closed
code
closed
license
Proprietary(API/Bedrock/Vertex/claude.ai)

No weights distributed in any form. Served only through the Anthropic API, Bedrock, Vertex and claude.ai, so 1/closed on every dimension. Opus 4.x governs as the current release; 4.7 scored identically, and nothing about the tier's openness turns on which release is read. Consolidated because Anthropic markets Claude Opus as the product and 4.7 and 4.8 are versions of it.

  • https://www.anthropic.com/news/claude-opus-4-8 recorded 2026-08-13

    "Claude Opus 4.8 is available everywhere today. Pricing for regular usage is unchanged from Opus 4.7 - $5 per million input tokens and $25 per million output tokens"; sold as metered API access, with no weights, corpus or training code offered

  • https://www.anthropic.com/news/claude-4 recorded 2026-08-13

    Opus 4 line origin (May 22 2025), proprietary, API/Bedrock/Vertex

Adoption

4 medium confidence
4.0

A directional read on market position rather than a measured usage figure. Opus is Anthropic's premium tier: top-tier reach, but lower volume than the Sonnet workhorse, so 4 rather than 5. Anthropic's launch post lists 28 enterprise testimonials (Stripe, Replit, Vercel and others), and Opus flagships appear on public LLM marketplace leaderboards such as OpenRouter - though those rankings render in the browser, so what the page confirms is that a live usage leaderboard exists, not Claude's place on it. No per-model usage or user figure is published, so no reach band is recorded.

  • https://www.anthropic.com/news/claude-opus-4-7 recorded 2026-08-13

    named enterprise testimonials (Cursor's Michael Truell on CursorBench, Harvey on BigLaw Bench, Replit's Michele Catasta); no quantitative per-model usage or user figure anywhere on the page

  • https://openrouter.ai/rankings recorded 2026-08-13

    "Live LLM rankings based on benchmarks and real data from millions of people using models through OpenRouter" - the leaderboard itself is client-rendered, so a plain fetch shows the page exists but not which models rank where

Capability

5 high confidence
5.0

Anthropic-reported: state-of-the-art on GDPval-AA and Finance Agent; CursorBench 70% (vs 58% for Opus 4.6); BigLaw Bench 90.9% at high effort; gains on SWE-bench Verified/Pro/Multilingual. Vendor-reported. No independent comparison stands behind these figures: the cited Google Gemini 3.5 post names neither Claude nor Opus, and the score does not rest on it.

Verified 2026-08-13