AI Potluck
Model components / Benchmark / eval datasets

LiveBench

LiveBench

An open, contamination-limited LLM benchmark with objective ground-truth answers and automatic scoring (no LLM judges), spanning six categories (math, coding, reasoning, language, instruction following, data analysis). New questions are released monthly and older ones retired to limit contamination. Built by Abacus.AI, NYU, Nvidia, UMD, USC; ICLR 2025 Spotlight.

Questions and ground-truth answers public (Apache-2.0/MIT). ~1.2k stars; multi-category HF dataset downloads. Monthly rotation is freshness management, not gating.

Openness

5 high confidence
5.0
questions
public
answers
public(ground_truth released)
license
Apache-2.0/MIT

All questions, answers, and code are released openly; rotation is for freshness, not access control.

Adoption

4 medium confidence
4.0

~1.2k GitHub stars plus multi-category HF dataset downloads (genuine multi-source adoption).

Capability

5 high confidence
5.0

Among the most influential contamination-resistant benchmarks; broader than code-only LiveCodeBench and judge-free.

Unchanged since 2026-06-17 (last edited, not re-checked)