lm-evaluation-harness
EleutherAIUnified framework for evaluating language models on 200+ academic NLP benchmarks (MMLU, ARC, GSM8K, HellaSwag, etc.) with a single config-driven CLI. Picked as the canonical reference harness; it powers the Hugging Face Open LLM Leaderboard and is the standard tool for reporting reproducible benchmark numbers in papers. 12.6k GitHub stars, 392 contributors, 58 commits in last 90 days.
lm-eval v0.4.12 (May 11 2026); confirmed live June 2026. Backend for HuggingFace Open LLM Leaderboard. The de-facto-standard model-eval harness.
Openness
5 high confidence- license
- MIT(OSI)
- source
- public(GitHub)
- core-gated
- ungated
MIT-licensed, fully public; OSI -> open_source.
- https://github.com/EleutherAI/lm-evaluation-harness recorded 2026-06-04
MIT license, public source, v0.4.12 (May 2026)
Adoption
4 high confidence~1.57M monthly PyPI downloads (lm-eval); backs HuggingFace Open LLM Leaderboard and is used in hundreds of papers and model release reports -- the strongest 'cited as standard' signal in the category. Headline = PyPI primary source; standard-status corroborates.
- https://pypistats.org/packages/lm-eval recorded 2026-06-04
~1,571,581 monthly PyPI downloads
- https://github.com/EleutherAI/lm-evaluation-harness recorded 2026-06-04
backend for HF Open LLM Leaderboard; used in hundreds of papers
Capability
5 high confidenceFrontier-definer within evaluation_code -- widest task + backend coverage and the standardization reference harness for LLM eval.
- https://github.com/EleutherAI/lm-evaluation-harness recorded 2026-06-04
60+ benchmarks, multi-backend support, Open LLM Leaderboard backend
Unchanged since 2026-07-30 (last edited, not re-checked)