MT-Bench
SGLang (LMSYS / UC Berkeley)An 80-question, multi-turn, open-ended benchmark for LLM instruction-following and conversational ability, scored with a strong LLM (e.g. GPT-4) as judge. Created by LMSYS, it ships inside FastChat with both questions and GPT-4 reference answers public. Introduced in the NeurIPS 2023 'LLM-as-a-Judge' paper.
Questions + GPT-4 reference answers public in FastChat (Apache-2.0); companion human-judgments dataset is CC-BY-4.0. Small (80 Q) and partly saturated now.
Openness
5 high confidence- questions
- public
- answers
- public(GPT-4 references in FastChat)
- license
- Apache-2.0
Both questions and reference answers are public and openly licensed.
- https://github.com/lm-sys/FastChat/tree/main/fastchat/llm_judge/data/mt_bench recorded 2026-06-17
public questions + GPT-4 references
Adoption
3 medium confidenceShips in FastChat (~39.5k stars, shared across the platform) - stars_fallback (capped at 3).
- https://github.com/lm-sys/FastChat recorded 2026-06-17
Apache-2.0, ~39.5k stars, MT-Bench in llm_judge
Capability
4 medium confidenceSeminal multi-turn chat benchmark that pioneered LLM-as-judge; now small and partly saturated vs harder peers.
- https://arxiv.org/abs/2306.05685 recorded 2026-06-17
MT-Bench / LLM-as-a-Judge paper
Unchanged since 2026-06-17 (last edited, not re-checked)