Open LLM Leaderboard
Hugging FaceHugging Face's flagship public leaderboard for comparing open LLMs on a fixed, reproducible benchmark suite run on a shared cluster. v2 (2024) raised difficulty with IFEval, BBH, MATH-L5, GPQA, MuSR, and MMLU-Pro. It evaluated 13,000+ models and drew 2M+ visitors before being archived.
ARCHIVED/retired (announced 2025-03-13); HF redirected users to the OpenEvals hub. HF Space was Apache-2.0; backend used lm-evaluation-harness (v1) and lighteval (v2). Listed for historical significance.
Openness
5 high confidence- license
- Apache-2.0(HF-Space)
- backend
- lm-eval-harness/lighteval(open)
- source
- public
Open Apache-2.0 leaderboard built on open eval backends (now archived).
- https://huggingface.co/spaces/open-llm-leaderboard/open_llm_leaderboard recorded 2026-06-17
Apache-2.0 Space, ~14k likes
Adoption
4 high confidenceARCHIVED 2025-03-13. At peak: 13,000+ models evaluated, 2M+ unique visitors - historically the most influential open-model leaderboard.
- https://huggingface.co/spaces/open-llm-leaderboard/open_llm_leaderboard/discussions/1135 recorded 2026-06-17
retirement announcement, 2025-03-13
Capability
4 high confidenceHistorically the most visible open-model leaderboard; static rather than human-preference; now archived.
- https://huggingface.co/docs/leaderboards/en/open_llm_leaderboard/archive recorded 2026-06-17
v2 benchmark suite, 2M visitors, archive
Unchanged since 2026-06-17 (last edited, not re-checked)