AI Potluck
Model components / Fine-tuned / chat models

GPT-5 line

OpenAI

OpenAI's current general-purpose flagship line spanning GPT-5 (August 2025), GPT-5.2, GPT-5.4, and GPT-5.5 (April 23, 2026), plus the GPT-5.5 Pro (API-only flagship), GPT-5.5 Instant (low-latency consumer tier), and GPT-5.3-Codex (the unified Codex+reasoning coder variant that succeeded o3/o4-mini). GPT-5.5 is marketed as OpenAI's 'smartest, most agentic' model; GPT-5.2 hits 80.0% on SWE-bench Verified and GPT-5.5 reportedly reaches 88.7% on a competing benchmark snapshot. GPT-5.5-high sits at #8 on the May 2026 LMArena text leaderboard (Elo 1481). Note: OpenAI has reportedly stopped reporting SWE-bench Verified scores, preferring SWE-bench Pro, citing test-case quality issues with Verified.

OpenAI GPT-5 line: GPT-5 (Aug 7 2025, became ChatGPT default replacing GPT-4o) through GPT-5.1/5.2/5.4 and GPT-5.5 / GPT-5.5 Instant (the current ChatGPT default as of May 2026). Unified reasoning+chat post-trained model family, closed/API + ChatGPT only. Consolidated on 2026-07-29 from gpt-5; openness follows gpt-5-line, the release that currently governs. See docs/reference/identity.md. OpenAI's own launch post answers HTTP 403 to a plain fetch, so the capability figures currently have no reachable source and that axis is unverified. Verified 2026-08-13 via the GPT-5 Wikipedia article and TechCrunch.

Openness

1 high confidence
1.0
weights
closed
data
closed
code
closed
license
Proprietary(API + ChatGPT only)

No downloadable weights, no published training data or code. Served through the OpenAI API and ChatGPT only, so 1/closed. Kept apart from GPT-4o and GPT-4.1 because OpenAI markets those as distinct products rather than versions of a single line, which is the same reason their slugs keep their version tokens. Adoption 5 reflects ChatGPT-scale reach.

Adoption

5 high confidence
5.0

GPT-5 was rolled out as the default model for logged-in ChatGPT users (Aug 2025). ChatGPT reached ~900M weekly active users (Feb 2026) and ~1B by May 2026, far exceeding the >10M threshold. Adoption attributed to the surface GPT-5 powers, not a standalone GPT-5 download count.

Capability

5 high confidence
5.0

OpenAI reported SWE-bench Verified 74.9%, MMMU 84.2%, AIME 2025 94.6% without tools and Aider Polyglot 88% at launch - frontier-tier coding and reasoning, though the figures are vendor-reported. Artificial Analysis rates plain GPT-5 (high) at 35, 84th of 188 on their recalibrated index, but a tier is scored on its strongest release and the current flagship there is GPT-5.6 Sol (max) at 61, so the 5 is not in doubt. OpenAI's system card describes the harness behind the headline figure - a fixed subset of 477 verified tasks at default verbosity - which the launch post does not.

  • https://cdn.openai.com/gpt-5-system-card.pdf recorded 2026-08-14

    OpenAI GPT-5 system card - "The SWE-Bench result in the GPT-5 launch blog post (74.9%), was run with the default verbosity setting in the API (verbosity = medium)"; "All SWE-bench evaluation runs use a fixed subset of n=477 verified tasks". Confirms the SWE-bench Verified basis figure from OpenAI directly. Carries no MMMU, AIME 2025 or Aider Polyglot headline.

  • https://openai.com/index/introducing-gpt-5/ recorded 2026-08-14

    OpenAI's GPT-5 launch post, the source of the four recorded figures, all four present in the body as fetched

Verified 2026-08-13