AI Potluck
Model components / Evaluation code

AlpacaEval

Tatsu Lab (Stanford)

Automatic judge-based evaluator for instruction-following models: a model annotator compares a candidate's outputs against a reference on a fixed instruction set to produce a win rate and a leaderboard. AlpacaEval 2.0 adds a length-controlled win rate to remove the bias toward longer outputs, which raises correlation with Chatbot Arena.

The README's data badge points at a sibling repository rather than anything in this one, which the openness axis resolves. Verified 2026-08-13 via the tatsu-lab/alpaca_eval repository, its license endpoint and its README.

Openness

5 high confidence
5.0
license
Apache-2.0(OSI)
source
public
core-gated
ungated

Fully Apache-2.0. The README header carries a 'Data License CC By NC 4.0' badge, but both badges link to the sibling tatsu-lab/alpaca_farm repository rather than to anything in this one, so the NC term is not a license this product makes you accept.

Adoption

2 medium confidence
2.0

tatsu-lab/alpaca_eval carries 2,012 GitHub stars, which bands at 1K-10K stars, level 2. Stars cap a level at 3 in any case, because a star is not a use. No download, install or customer figure is published for this product, so stars are the only honest signal available and the level is directional.

Capability

4 medium confidence
4.0

The de-facto standard for its niche, with an influential length-controlled win-rate metric, but narrower in scope than the multi-task harnesses it sits beside.

Verified 2026-08-12