LiveBench
LiveBenchAn open, contamination-limited LLM benchmark with objective ground-truth answers and automatic scoring (no LLM judges), spanning six categories (math, coding, reasoning, language, instruction following, data analysis). New questions are released monthly and older ones retired to limit contamination. Built by Abacus.AI, NYU, Nvidia, UMD, USC; ICLR 2025 Spotlight.
Questions and ground-truth answers public (Apache-2.0/MIT). ~1.2k stars; multi-category HF dataset downloads. Monthly rotation is freshness management, not gating.
Openness
5 high confidence- questions
- public
- answers
- public(ground_truth released)
- license
- Apache-2.0/MIT
All questions, answers, and code are released openly; rotation is for freshness, not access control.
- https://huggingface.co/datasets/livebench/reasoning recorded 2026-06-17
public ground_truth column, release/removal dates
Adoption
4 medium confidence~1.2k GitHub stars plus multi-category HF dataset downloads (genuine multi-source adoption).
- https://github.com/livebench/livebench recorded 2026-06-17
~1.2k stars, six categories, monthly releases
Capability
5 high confidenceAmong the most influential contamination-resistant benchmarks; broader than code-only LiveCodeBench and judge-free.
- https://arxiv.org/abs/2406.19314 recorded 2026-06-17
ICLR 2025 Spotlight, releases all questions/code/answers
Unchanged since 2026-06-17 (last edited, not re-checked)