AI Potluck
Model components / Benchmark / eval datasets

MT-Bench

SGLang (LMSYS / UC Berkeley)

An 80-question, multi-turn, open-ended benchmark for LLM instruction-following and conversational ability, scored with a strong LLM (e.g. GPT-4) as judge. Created by LMSYS, it ships inside FastChat with both questions and GPT-4 reference answers public. Introduced in the NeurIPS 2023 'LLM-as-a-Judge' paper.

Questions + GPT-4 reference answers public in FastChat (Apache-2.0); companion human-judgments dataset is CC-BY-4.0. Small (80 Q) and partly saturated now.

Openness

5 high confidence
5.0
questions
public
answers
public(GPT-4 references in FastChat)
license
Apache-2.0

Both questions and reference answers are public and openly licensed.

Adoption

3 medium confidence
3.0

Ships in FastChat (~39.5k stars, shared across the platform) - stars_fallback (capped at 3).

Capability

4 medium confidence
4.0

Seminal multi-turn chat benchmark that pioneered LLM-as-judge; now small and partly saturated vs harder peers.

Unchanged since 2026-06-17 (last edited, not re-checked)