AI Potluck
Model components / Evaluation code

AlpacaEval

Tatsu Lab (Stanford)

Automatic LLM-judge evaluator for instruction-following models: an LLM annotator compares a candidate's outputs against a reference on a fixed instruction set to produce a win-rate and leaderboard. AlpacaEval 2.0 adds a length-controlled win-rate to debias for output length (raising correlation with Chatbot Arena to ~0.98).

Apache-2.0. ~2k stars. A full eval runs in ~5 min for under ~$10.

Openness

5 high confidence
5.0
license
Apache-2.0(OSI)
source
public
core-gated
ungated

Fully Apache-2.0.

Adoption

3 medium confidence
3.0

~2k stars - stars_fallback; widely used as the standard LLM-judge instruction-following eval.

Capability

4 medium confidence
4.0

De-facto standard for its niche with an influential LC metric; narrower than multi-task harnesses.

Unchanged since 2026-07-30 (last edited, not re-checked)