AI Potluck
Model components / Evaluation code

Chatbot Arena

LMArena

Hosted human-preference LLM evaluation platform; rebranded LMSYS Chatbot Arena -> LMArena -> Arena (arena.ai; lmarena.ai 301-redirects to arena.ai as of June 2026). The PRODUCT is the hosted leaderboard/battle platform (closed). Its historical OSS code is FastChat (Apache-2.0) but stale (v0.2.36, Feb 2024); auxiliary OSS eval tooling Arena-Hard-Auto (Apache-2.0) is maintained. Confirmed live June 2026.

Hosted human-preference LLM evaluation platform; rebranded LMSYS Chatbot Arena -> LMArena -> Arena (arena.ai; lmarena.ai 301-redirects to arena.ai as of June 2026). The PRODUCT is the hosted leaderboard/battle platform (closed). Its historical OSS code is FastChat (Apache-2.0) but stale (v0.2.36, Feb 2024); auxiliary OSS eval tooling Arena-Hard-Auto (Apache-2.0) is maintained. Confirmed live June 2026.

Openness

2 medium confidence
2.0
platform
closed(hosted arena.ai, proprietary live leaderboard + private vote data)
source
partial(FastChat is a real arena implementation but stale relative to the operating platform
reference-code
FastChat Apache-2.0 (powers Chatbot Arena but stale, last release Feb 2024, 39.5k stars)
aux-OSS
Arena-Hard-Auto Apache-2.0 (automatic LLM eval, highest correlation to LMArena)

Scored as the hosted platform: the live evaluation product, its vote stream and the current ranking pipeline are closed, while substantial OSS heritage (FastChat) and an OSS automatic-eval companion (Arena-Hard-Auto) exist. The OSS code is stale relative to the operating platform, which is what makes this source_available rather than closed: FastChat is a real arena implementation, just not the one running. Class corrected from open_core on 2026-07-30. In this ladder open_core means an OSI core with functionality withheld for a paid tier, and there is no open core here, only an open periphery. A repo that exists while the runtime does not ship is source:partial, which the ladder scores 2/source_available. The score itself was already 2.

  • https://github.com/lm-sys/FastChat recorded 2026-06-04

    Apache-2.0; 'FastChat powers Chatbot Arena (lmarena.ai), serving over 10M chat requests for 70+ LLMs'; latest release v0.2.36 Feb 2024

  • https://github.com/lmarena/arena-hard-auto recorded 2026-06-04

    Apache-2.0 automatic LLM eval tool, highest correlation to LMArena

  • https://arena.ai/ recorded 2026-06-04

    hosted 'Official AI Ranking & LLM Leaderboard' with Battle Mode; conversations disclosed to providers (proprietary platform)

Adoption

4 medium confidence
4.0

FastChat README (primary) states Chatbot Arena has served 'over 10 million chat requests for 70+ LLMs'; Arena Elo is the most-cited human-preference leaderboard in frontier model releases (LMArena Elo appears across vendor launch posts). Headline = the 10M+ chat-request figure (primary, FastChat repo). Banded 1M-10M conservatively on request volume; could be higher cumulatively.

Capability

4 medium confidence
4.0

Reference platform for human-preference evaluation and the de-facto Elo standard cited everywhere; not a 5 because as eval CODE it is largely closed/hosted and narrower than execution-based harnesses on agentic/task coverage. Capability assessed on its evaluation methodology breadth.

Unchanged since 2026-07-30 (last edited, not re-checked)