AI Potluck
Model components / Evaluation code

lm-evaluation-harness

EleutherAI

Unified framework for evaluating language models on 200+ academic NLP benchmarks (MMLU, ARC, GSM8K, HellaSwag, etc.) with a single config-driven CLI. Picked as the canonical reference harness; it powers the Hugging Face Open LLM Leaderboard and is the standard tool for reporting reproducible benchmark numbers in papers. 12.6k GitHub stars, 392 contributors, 58 commits in last 90 days.

lm-eval v0.4.12 (May 11 2026); confirmed live June 2026. Backend for HuggingFace Open LLM Leaderboard. The de-facto-standard model-eval harness.

Openness

5 high confidence
5.0
license
MIT(OSI)
source
public(GitHub)
core-gated
ungated

MIT-licensed, fully public; OSI -> open_source.

Adoption

4 high confidence
4.0

~1.57M monthly PyPI downloads (lm-eval); backs HuggingFace Open LLM Leaderboard and is used in hundreds of papers and model release reports -- the strongest 'cited as standard' signal in the category. Headline = PyPI primary source; standard-status corroborates.

Capability

5 high confidence
5.0

Frontier-definer within evaluation_code -- widest task + backend coverage and the standardization reference harness for LLM eval.

Unchanged since 2026-07-30 (last edited, not re-checked)