AI Potluck
Model components / Benchmark / eval datasets

LiveCodeBench

LiveCodeBench Team

Continuously updated coding benchmark that collects fresh problems from LeetCode, Codeforces and AtCoder after a model's training cutoff, so evaluation stays contamination-free. It covers code generation, self-repair, execution and test-output prediction, and ships release-versioned snapshots with hidden test cases bundled.

The dataset card declares a bare Creative Commons tag naming no variant, and the harness repository's terms make no statement about the problems, which the openness axis resolves. Verified 2026-08-13 via the livecodebench/code_generation_lite card, its raw README and the harness repository.

Openness

3 medium confidence
3.0
data
downloadable(HF parquet)
license
cc(the card YAML names the Creative Commons family and no variant
datasheet
present(README documents versions/sources/structure)
redistributable
yes
splits
public problems with hidden test cases bundled

Downloadable with a datasheet, but the only license statement the project makes about the corpus is a bare "cc" naming no variant, and CC-BY and CC-BY-NC sit two tiers apart. No variant is stated anywhere the project publishes: not in the card front matter, the Hub API, the sibling code_generation and test_generation corpora, the repository LICENSE (MIT, which covers the lcb_runner harness rather than the problems), the README, ERRATA.md, the project site, the leaderboard or the paper. Files that download freely without a grant to use them are not open, and not closed either, so this sits at gated - the barrier is the missing permission rather than a login. The problems are collected from LeetCode, AtCoder and Codeforces, which is a plausible reason no variant has been committed to.

Adoption

3 medium confidence
3.0

91,719 Hugging Face downloads in the trailing 30 days for livecodebench/code_generation_lite, the single declared artifact, which puts it in the 10K-100K band, level 3 on the dataset adoption scale. That scale tops out at >1M.

Capability

not assessed

A dataset is not 'capable', so this axis is left unscored. Its contamination resistance - the "live", continuously updated design - is a real quality strength, but openness and adoption are what carry this category.

Verified 2026-08-13