AI Potluck
Model components / Evaluation code

lm-evaluation-harness

EleutherAI

Unified framework for evaluating language models across more than 200 academic NLP benchmarks, including MMLU, ARC, GSM8K and HellaSwag, driven by a single config-based CLI. It is the backend behind the Hugging Face Open LLM Leaderboard and is used to report reproducible benchmark numbers in papers.

Verified 2026-08-13 via the EleutherAI/lm-evaluation-harness repository, its license endpoint and its README.

Openness

5 high confidence
5.0
license
MIT(OSI)
source
public(GitHub)
core-gated
ungated

MIT-licensed and fully public, and MIT is an OSI license, so this is open source.

Adoption

4 high confidence
4.0

The lm-eval package draws 1,658,820 PyPI downloads in the trailing 30 days, which bands at level 4 (1M-10M). Downloads are the primary measure here, and standard status corroborates them: the harness backs the Hugging Face Open LLM Leaderboard - the README names the leaderboard task group - and is used in hundreds of papers and model release reports, the strongest 'cited as standard' signal in the category.

Capability

5 high confidence
5.0

The widest task and backend coverage in the category, and the standardization reference for LLM evaluation. The README states 'Over 60 standard academic benchmarks for LLMs, with hundreds of subtasks and variants implemented', and lists the spread of model backends behind them.

Verified 2026-08-13