AI Potluck
Model components / Fine-tuned / chat models

Grok 4.20

xAI

xAI Grok 4.20; beta Feb 17, 2026, API GA Mar 10, 2026 (Grok 4.20 + Multi-agent). Revision '4.20 0309 v2' dated Apr 7, 2026. Multi-agent architecture (4 parallel agents); 2M-token context.

xAI Grok 4.20; beta Feb 17, 2026, API GA Mar 10, 2026 (Grok 4.20 + Multi-agent). Revision '4.20 0309 v2' dated Apr 7, 2026. Multi-agent architecture (4 parallel agents); 2M-token context. Consolidated on 2026-07-29 from grok-3; openness follows grok-4-20, the release that currently governs. See docs/reference/identity.md. Two axes are flagged for rescoring - xAI's docs now headline grok-4.6 as the flagship, and Artificial Analysis has recalibrated its Intelligence Index, so the index score the capability record cites no longer exists on their scale. The vendor now brands itself SpaceXAI. Verified 2026-08-13 via the xAI release notes, the xAI model reference and the Artificial Analysis model page.

Openness

1 high confidence
1.0
weights
closed
data
closed
code
closed
license
Proprietary(API-only)

Proprietary, served through the xAI API, grok.com and X, with no downloadable weights, so fully closed. The bare grok slug names this model line - the consumer chat app is a separate product, recorded as grok-app. xAI's model reference now headlines grok-4.6 as the flagship rather than the grok-4.20 release this record is named for, and the vendor now brands itself SpaceXAI; neither moves a closed score.

  • https://docs.x.ai/developers/release-notes recorded 2026-08-13

    xAI release notes record "Grok 4.20 and Grok 4.20 Multi-agent are live" and point at the API docs; no weight release accompanies the entry

  • https://docs.x.ai/docs/models recorded 2026-08-13

    xAI's model reference now headlines "Meet grok-4.6 ... Our flagship model for code and everything else", with a 500k-token context, and documents pricing and model-selection guidance only, with no distribution or download path

Adoption

5 low confidence
5.0

A tier bands on its strongest release. xAI's model reference headlines "Meet grok-4.6 ... Our flagship model for code and everything else" and lists grok-4.3, grok-4.5 and grok-4.6 above the grok-4.20 the record is named for, so the tier's current release is grok-4.6, and it is the default across the X and grok.com surfaces. What is banded here is those surfaces, not a per-model count: xAI publishes no standalone per-model user or usage figure, so this is a reported-traction judgment, no reach band is recorded, and confidence stays low. Claude Sonnet and Gemini Pro are treated the same way, and Command R has the same shape. The consumer app itself is a separate product, grok-app; the bare grok slug is the model line.

  • https://docs.x.ai/docs/models recorded 2026-08-14

    xAI's model reference headlines "Meet grok-4.6 ... Our flagship model for code and everything else - agentic tool calling", and lists grok-4.3, grok-4.5 and grok-4.6 alongside the older grok-4.20. No user, usage or per-model traffic figure appears anywhere on it. The URL now redirects to /developers/models.

  • https://x.ai/news/grok-3 recorded 2026-08-14

    "Grok 3 Beta - The Age of Reasoning Agents"; the superseded release the old band was set on. No usage or user figure published.

Capability

5 high confidence
5.0

Artificial Analysis publishes "Grok 4.6 (high) scores 61 on the Artificial Analysis Intelligence Index, placing it well above average among comparable models (median - 34)", and its model record there is not marked deprecated, unlike the grok-4.20 page, which banners a replacement. Ranking the 456 model records the page embeds by their index score, the only distinct models above Grok 4.6 (high) at 60.92 are Claude Opus 5 at 63.05 and Claude Fable 5 at 62.07, with GPT-5.6 Sol level at 60.93 and Kimi K3 next at 59.70. The rank badge itself is drawn in the browser rather than served in the page body, so this is derived from the embedded data rather than read off the badge; it agrees with the "# 6 / 188" the page reports. grok-4.6 is the release the band is taken on, since a tier's adoption and capability are read on its strongest release, and xAI's own model reference headlines grok-4.6 above grok-4.3, grok-4.5 and the older grok-4.20. A model inside the top three of the index it is scored on is not a 4. Two things the band does not rest on: the older grok-4.20 page, which reads 38 and rank 62 of 184 on the same recalibrated scale and banners a replacement, and the roughly 1505-1535 LMArena figure on NextBigFuture, which is an explicit estimate rather than a measurement.

  • https://artificialanalysis.ai/models/grok-4-20 recorded 2026-08-13

    "Grok 4.20 0309 v2 (Reasoning) scores 38 on the Artificial Analysis Intelligence Index, placing it above average among comparable models (median: 34)"; Intelligence rank "# 62 / 184"; banner reads "This model is deprecated ... SpaceXAI has launched a newer model, Grok 4.3 (high). We suggest considering it instead."

  • https://www.nextbigfuture.com/2026/02/xai-launches-grok-4-20-and-it-has-4-ai-agents-collaborating.html recorded 2026-08-13

    "Estimated Arena (LMArena) Elo for Grok 4.20 - ~1505-1535 provisional ... Grok 4.1 Thinking is already at 1483". An estimate rather than a measured leaderboard position.

  • https://artificialanalysis.ai/models/grok-4-6 recorded 2026-08-14

    "Grok 4.6 (high) scores 61 on the Artificial Analysis Intelligence Index, placing it well above average among comparable models (median - 34)"; Intelligence rank "# 6 / 188"; no deprecation banner, unlike the grok-4-20 page. The current release for the tier, on the recalibrated scale that replaced the one the old 49 was recorded against.

  • https://artificialanalysis.ai/models/grok-4-6 recorded 2026-08-14

    The same page, fetched again for the 4 -> 5 ruling. Body carries "Grok 4.6 (high) scores 61 on the Artificial Analysis Intelligence Index, placing it well above average among comparable models (median - 34)" and the model record `"deprecated":false,"deprecatedTo":null`. Ranking the 456 embedded model records by `intelligenceIndex`, Grok 4.6 (high) is at 60.92 behind only Claude Opus 5 (63.05) and Claude Fable 5 (62.07), level with GPT-5.6 Sol (60.93) and ahead of Kimi K3 (59.70). The digest drifts between fetches because the page carries live pricing and throughput figures.

  • https://docs.x.ai/developers/models recorded 2026-08-14

    xAI's model reference, the page docs.x.ai/docs/models now redirects to. It still headlines "Meet grok-4.6" and lists grok-4.3, grok-4.5 and grok-4.6 above the grok-4.20 this record's openness axis is named for, which is what makes grok-4.6 the release the max-across-releases rule takes capability on. The served body carries a per-request token, so its digest drifts between fetches without the page changing.

Verified 2026-08-13