AI Potluck
Model components / Fine-tuned / chat models

GPT-5 line

OpenAI

OpenAI's current general-purpose flagship line spanning GPT-5 (August 2025), GPT-5.2, GPT-5.4, and GPT-5.5 (April 23, 2026), plus the GPT-5.5 Pro (API-only flagship), GPT-5.5 Instant (low-latency consumer tier), and GPT-5.3-Codex (the unified Codex+reasoning coder variant that succeeded o3/o4-mini). GPT-5.5 is marketed as OpenAI's 'smartest, most agentic' model; GPT-5.2 hits 80.0% on SWE-bench Verified and GPT-5.5 reportedly reaches 88.7% on a competing benchmark snapshot. GPT-5.5-high sits at #8 on the May 2026 LMArena text leaderboard (Elo 1481). Note: OpenAI has reportedly stopped reporting SWE-bench Verified scores, preferring SWE-bench Pro, citing test-case quality issues with Verified.

OpenAI GPT-5 line: GPT-5 (Aug 7 2025, became ChatGPT default replacing GPT-4o) through GPT-5.1/5.2/5.4 and GPT-5.5 / GPT-5.5 Instant (the current ChatGPT default as of May 2026). Unified reasoning+chat post-trained model family, closed/API + ChatGPT only. Verified via OpenAI pages + Wikipedia + TechCrunch June 2026. Consolidated on 2026-07-29 from gpt-5; openness follows gpt-5-line, the release that currently governs. See sources/slug_aliases.yaml.

Openness

1 high confidence
1.0
weights
closed
data
closed
code
closed
license
Proprietary(API + ChatGPT only)

No downloadable weights, no published training data or code. Served through the OpenAI API and ChatGPT only, so 1/closed. Kept apart from GPT-4o and GPT-4.1 because OpenAI markets those as distinct products rather than versions of a single line, which is the same reason their slugs keep their version tokens. Adoption 5 reflects ChatGPT-scale reach.

Adoption

5 high confidence
5.0

GPT-5 was rolled out as the default model for logged-in ChatGPT users (Aug 2025). ChatGPT reached ~900M weekly active users (Feb 2026) and ~1B by May 2026, far exceeding the >10M threshold. Adoption attributed to the surface GPT-5 powers, not a standalone GPT-5 download count.

Capability

5 high confidence
5.0

OpenAI-reported at launch: SWE-bench Verified 74.9%, MMMU 84.2%, AIME 2025 94.6% (no tools), Aider Polyglot 88%. Frontier-tier coding/reasoning. Vendor-reported figures.

Unchanged since 2026-07-29 (last edited, not re-checked)