Open LLM Leaderboard
Hugging FaceHugging Face's public leaderboard for comparing open language models on a fixed, reproducible benchmark suite run on a shared cluster. Its second version raised difficulty with IFEval, BBH, MATH Level 5, GPQA, MuSR and MMLU-Pro, replacing the original six-benchmark suite.
Archived in March 2025 with users redirected to the OpenEvals hub, though the Space still runs. Listed for historical significance. Verified 2026-08-13 via the Space README, the Space API and the leaderboard archive page.
Openness
5 high confidence- license
- Apache-2.0(HF-Space)
- backend
- lm-eval-harness/lighteval(open)
- source
- public
- core-gated
- ungated(Space source public and docker-compose deployable, no paid tier)
An Apache-2.0 leaderboard built on open eval backends, now archived. The Space is still public and still running, its card data declares apache-2.0 and a Docker SDK, and the README brings the React frontend and FastAPI backend up with `docker-compose up`; nothing has been withdrawn or gated since archival.
- https://huggingface.co/spaces/open-llm-leaderboard/open_llm_leaderboard recorded 2026-06-17
Apache-2.0 Space, ~14k likes
- https://huggingface.co/spaces/open-llm-leaderboard/open_llm_leaderboard/raw/main/README.md recorded 2026-08-11
Space README frontmatter declares 'license: apache-2.0' and 'sdk: docker'; the body documents bringing the React frontend and FastAPI backend up with `docker-compose up`. The Space's files are readable and downloadable over the raw endpoint, and no feature is described as paid, licensed, or restricted.
- https://huggingface.co/spaces/open-llm-leaderboard/open_llm_leaderboard/raw/main/README.md recorded 2026-08-13
Frontmatter declares 'license: apache-2.0' and 'sdk: docker', and the body documents bringing the React frontend and FastAPI backend up with `docker-compose up`. No feature is described as paid, licensed or restricted.
- https://huggingface.co/api/spaces/open-llm-leaderboard/open_llm_leaderboard recorded 2026-08-13
Space API reports private = false, cardData.license = apache-2.0, sdk = docker, runtime stage RUNNING, and 14,072 likes.
Adoption
4 high confidenceArchived on 2025-03-13, after evaluating 13,000+ models and drawing 2M+ unique visitors - historically the most influential open-model leaderboard. The two figures come from different pages: the retirement post gives the 13,000+ models evaluated, and the archive page reports 'more than 2 million unique people over the last 10 months' plus around 300,000 community members a month. The Space still shows 14,072 likes. These are peak lifetime numbers by nature, the product being retired.
- https://huggingface.co/spaces/open-llm-leaderboard/open_llm_leaderboard/discussions/1135 recorded 2026-06-17
retirement announcement, 2025-03-13
- https://huggingface.co/spaces/open-llm-leaderboard/open_llm_leaderboard/discussions/1135 recorded 2026-08-13
Retirement announcement dated 2025-03-13, stating 'For the last 2 years, we've evaluated over 13K models with the Open LLM Leaderboard' and noting that over 200 community-led leaderboards now exist on Hugging Face. It carries no visitor figure.
- https://huggingface.co/docs/leaderboards/en/open_llm_leaderboard/archive recorded 2026-08-13
archive page: 'visited by more than 2 million unique people over the last 10 months' and 'Around 300,000 community members used and collaborated on it monthly'.
- https://huggingface.co/api/spaces/open-llm-leaderboard/open_llm_leaderboard recorded 2026-08-13
likes = 14,072 on the archived Space
Capability
4 high confidenceHistorically the most visible open-model leaderboard: a standardized static benchmark suite rather than human-preference voting, and now archived. One caveat on the evidence: the archive page cited here documents the v1 leaderboard rather than v2, describing six benchmarks (ARC 25-shot, HellaSwag 10-shot, MMLU 5-shot, TruthfulQA, Winogrande 5-shot, GSM8k 5-shot) run on a single node of eight H100s through lm-evaluation-harness, with a reproduction command. It therefore establishes the shared-cluster reproducible runs and the 2M visitors, but not the v2 suite, which the Space's own tags corroborate instead; the 13,000+ model figure comes from the retirement post.
- https://huggingface.co/docs/leaderboards/en/open_llm_leaderboard/archive recorded 2026-06-17
v2 benchmark suite, 2M visitors, archive
- https://huggingface.co/docs/leaderboards/en/open_llm_leaderboard/archive recorded 2026-08-13
This page covers Open LLM Leaderboard v1 - six benchmarks (ARC, HellaSwag, MMLU, TruthfulQA, Winogrande, GSM8k) evaluated with the EleutherAI harness on a single node of 8 H100s, with the exact reproduction command published. Archived June 2024 and replaced by v2. It does NOT list the v2 suite this value names.
- https://huggingface.co/api/spaces/open-llm-leaderboard/open_llm_leaderboard recorded 2026-08-13
Space card data still tags the leaderboard 'eval:code' and 'eval:math' with automatic submission and a public test set, and describes it as 'Track, rank and evaluate open LLMs and chatbots'.
Verified 2026-08-13