AI Potluck
Model components / Evaluation code

lighteval

Hugging Face

Lightweight successor to the eval harness HF used internally; supports vLLM, transformers, TGI and OpenAI-compatible backends with a focus on speed and easy custom metric authoring. Picked when teams want a fast, hackable runner that integrates cleanly with the HF ecosystem (datasets, accelerate). 2.4k stars, 120 contributors, 17 commits in last 90 days; powers parts of the Open LLM Leaderboard v2.

lighteval v0.13.0 (24 Nov 2025), latest GitHub release; confirmed live on GitHub June 2026. All-in-one LLM eval toolkit from HF's Leaderboard & Evals team; powers the HF Open LLM Leaderboard. Columbia extension (model.code.evaluation) caveat applies.

Openness

5 high confidence
5.0
license
MIT(OSI)
source
public(github.com/huggingface/lighteval)
core-gated
ungated

Fully OSI-licensed (MIT), full source public, no managed/commercial tier gating the eval engine.

Adoption

2 medium confidence
2.0

~37,239 PyPI downloads in the last month (pypistats). Powers the HF Open LLM Leaderboard and is the reproducible eval suite behind it; meaningful in the eval/research community but a smaller install base than general-purpose harnesses. 2.4k GitHub stars corroborate only. Headline = PyPI volume (primary).

Capability

4 high confidence
4.0

Strong across the recipe's feature-matrix dimensions (coverage, backend breadth, reproducibility, some agentic support). Not a 5; lm-evaluation-harness remains the broader de-facto standard cited in more frontier release reports; lighteval is the HF-leaderboard-tied counterpart.

Unchanged since 2026-07-30 (last edited, not re-checked)