AI Potluck
Model components / Benchmark / eval datasets

LiveBench

LiveBench

Contamination-limited language-model benchmark with objective ground-truth answers and automatic scoring rather than judge models, spanning math, coding, reasoning, language, instruction following and data analysis. New questions are released monthly and older ones retired to limit contamination. Built by Abacus.AI with NYU, Nvidia, UMD and USC, and published as an ICLR 2025 spotlight.

The monthly rotation is freshness management rather than gating; questions ship with their ground-truth answers. Verified 2026-08-12 via the repository DATASHEET.md, the repository LICENSE and the livebench/reasoning dataset card.

Openness

5 high confidence
5.0
questions
public
answers
public(ground_truth released)
license
Apache-2.0/MIT(repo LICENSE is Apache-2.0, carried over from FastChat
datasheet
present(docs/DATASHEET.md, a full datasheets-for-datasets questionnaire covering motivation, composition, collection and distribution)
access
ungated(HF datasets download without a gate)

All questions, answers, and code are released openly; rotation is for freshness, not access control. The documentation is a full datasheets-for-datasets questionnaire in the repository rather than a Hugging Face card, and it states the distribution license outright.

Adoption

3 medium confidence
3.0

36,032 Hugging Face downloads in the trailing 30 days, summed across the ten LiveBench category corpora: math 6,923; reasoning 5,633; model_judgment 5,365; coding 4,761; data_analysis 4,270; instruction_following 4,250; language 4,165; liveswebench 370; model_answer 150; and liveswebench-patches 145. That puts it in the 10K-100K band, level 3 on the dataset adoption scale.

Capability

5 high confidence
5.0

Among the most influential contamination-resistant benchmarks: broader than the code-only LiveCodeBench, and judge-free, since every question is auto-scored against objective ground truth.

  • https://arxiv.org/abs/2406.19314 recorded 2026-08-13

    "LiveBench - A Challenging, Contamination-Limited LLM Benchmark"; the abstract still lists the three release properties, names the six task domains, reports top models below 70% accuracy, and says "We release all questions, code, and model answers. Questions are added and updated on a monthly basis"

Verified 2026-08-12