LiveCodeBench
LiveCodeBench TeamContinuously updated benchmark that scrapes fresh LeetCode/Codeforces/AtCoder problems after a model's training cutoff, ensuring contamination-free evaluation across code generation, self-repair, execution, and test-output prediction. v6 includes 1,000+ problems collected May 2023–2025. Per Artificial Analysis (May 2026), Gemini 3 Pro Preview leads at 91.7%, followed by Gemini 3 Flash Reasoning at 90.8% and DeepSeek V3.2 Speciale at 89.6%; DeepSeek V4-Pro-Max reached 93.5% on a later snapshot. ~870 GitHub stars but disproportionate influence as the trusted dynamic complement to HumanEval/MBPP.
LiveCodeBench code_generation_lite, release_v5 = 880 problems (May 2023-Jan 2025); a 'live', continuously-updated contamination-free coding benchmark (problems from LeetCode/AtCoder/Codeforces) covering code generation, self-repair, test-output prediction, execution. HF dataset verified live June 2026 (71,393 monthly downloads).
Openness
4 medium confidence- data
- downloadable(HF parquet)
- license
- cc(generic CC tag, specific variant unspecified on dataset card)
- datasheet
- present(README documents versions/sources/structure)
- redistributable
- yes
- splits
- public problems with hidden test cases bundled
Open and downloadable with a datasheet, but the HF card declares only a generic 'cc' license (no specific CC variant), so the exact reuse terms are ambiguous. Scored open but not a clean CC-BY; capped at 4.
- https://huggingface.co/datasets/livecodebench/code_generation_lite recorded 2026-06-04
license:cc, 880 problems (release_v5), 71,393 monthly downloads, datasheet/README
- https://huggingface.co/datasets/livecodebench/code_generation_lite/blob/main/README.md recorded 2026-06-04
YAML metadata license: cc
Adoption
4 medium confidence71,393 monthly HF downloads (primary). De facto standard contamination-free coding benchmark; LiveCodeBench appears as a headline coding metric in frontier model release reports (e.g. DeepSeek-V4-Pro reports LiveCodeBench Pass@1 93.5 in the frozen flagship set). 'Usage' here is downloads + benchmark citation, not end-users; placed at 4 on download volume plus release-report standard status, which is the top adoption signal for a benchmark.
- https://huggingface.co/datasets/livecodebench/code_generation_lite recorded 2026-06-04
71,393 monthly downloads
- https://livecodebench.github.io/ recorded 2026-06-04
contamination-free coding benchmark across generation/self-repair/execution
Capability
not assessedA dataset is not 'capable'. Contamination-resistance (the 'live' continuously-updated design) is a quality strength but not scored on the 1-5 capability axis per the recipe; openness + adoption carry this category.
Unchanged since 2026-06-24 (last edited, not re-checked)