AI Potluck
Model components / Evaluation code

Open LLM Leaderboard

Hugging Face

Hugging Face's flagship public leaderboard for comparing open LLMs on a fixed, reproducible benchmark suite run on a shared cluster. v2 (2024) raised difficulty with IFEval, BBH, MATH-L5, GPQA, MuSR, and MMLU-Pro. It evaluated 13,000+ models and drew 2M+ visitors before being archived.

ARCHIVED/retired (announced 2025-03-13); HF redirected users to the OpenEvals hub. HF Space was Apache-2.0; backend used lm-evaluation-harness (v1) and lighteval (v2). Listed for historical significance.

Openness

5 high confidence
5.0
license
Apache-2.0(HF-Space)
backend
lm-eval-harness/lighteval(open)
source
public

Open Apache-2.0 leaderboard built on open eval backends (now archived).

Adoption

4 high confidence
4.0

ARCHIVED 2025-03-13. At peak: 13,000+ models evaluated, 2M+ unique visitors - historically the most influential open-model leaderboard.

Capability

4 high confidence
4.0

Historically the most visible open-model leaderboard; static rather than human-preference; now archived.

Unchanged since 2026-06-17 (last edited, not re-checked)