AI Potluck
Model components / Evaluation code

Open LLM Leaderboard

Hugging Face

Hugging Face's public leaderboard for comparing open language models on a fixed, reproducible benchmark suite run on a shared cluster. Its second version raised difficulty with IFEval, BBH, MATH Level 5, GPQA, MuSR and MMLU-Pro, replacing the original six-benchmark suite.

Archived in March 2025 with users redirected to the OpenEvals hub, though the Space still runs. Listed for historical significance. Verified 2026-08-13 via the Space README, the Space API and the leaderboard archive page.

Openness

5 high confidence
5.0
license
Apache-2.0(HF-Space)
backend
lm-eval-harness/lighteval(open)
source
public
core-gated
ungated(Space source public and docker-compose deployable, no paid tier)

An Apache-2.0 leaderboard built on open eval backends, now archived. The Space is still public and still running, its card data declares apache-2.0 and a Docker SDK, and the README brings the React frontend and FastAPI backend up with `docker-compose up`; nothing has been withdrawn or gated since archival.

Adoption

4 high confidence
4.0

Archived on 2025-03-13, after evaluating 13,000+ models and drawing 2M+ unique visitors - historically the most influential open-model leaderboard. The two figures come from different pages: the retirement post gives the 13,000+ models evaluated, and the archive page reports 'more than 2 million unique people over the last 10 months' plus around 300,000 community members a month. The Space still shows 14,072 likes. These are peak lifetime numbers by nature, the product being retired.

Capability

4 high confidence
4.0

Historically the most visible open-model leaderboard: a standardized static benchmark suite rather than human-preference voting, and now archived. One caveat on the evidence: the archive page cited here documents the v1 leaderboard rather than v2, describing six benchmarks (ARC 25-shot, HellaSwag 10-shot, MMLU 5-shot, TruthfulQA, Winogrande 5-shot, GSM8k 5-shot) run on a single node of eight H100s through lm-evaluation-harness, with a reproduction command. It therefore establishes the shared-cluster reproducible runs and the 2M visitors, but not the v2 suite, which the Space's own tags corroborate instead; the 13,000+ model figure comes from the retirement post.

Verified 2026-08-13