AlpacaEval
Tatsu Lab (Stanford)Automatic LLM-judge evaluator for instruction-following models: an LLM annotator compares a candidate's outputs against a reference on a fixed instruction set to produce a win-rate and leaderboard. AlpacaEval 2.0 adds a length-controlled win-rate to debias for output length (raising correlation with Chatbot Arena to ~0.98).
Apache-2.0. ~2k stars. A full eval runs in ~5 min for under ~$10.
Openness
5 high confidence5.0
- license
- Apache-2.0(OSI)
- source
- public
- core-gated
- ungated
Fully Apache-2.0.
- https://github.com/tatsu-lab/alpaca_eval recorded 2026-06-17
Apache-2.0, ~2.0k stars
Adoption
3 medium confidence3.0
~2k stars - stars_fallback; widely used as the standard LLM-judge instruction-following eval.
- https://github.com/tatsu-lab/alpaca_eval recorded 2026-06-17
~1,996 stars
Capability
4 medium confidence4.0
De-facto standard for its niche with an influential LC metric; narrower than multi-task harnesses.
- https://github.com/tatsu-lab/alpaca_eval recorded 2026-06-17
AlpacaEval 2.0 LC win-rate
Unchanged since 2026-07-30 (last edited, not re-checked)