AI Potluck
Model components / Evaluation code

BigCodeBench

BigCode Project

Evaluation harness for code generation models on 1,140 practical programming tasks requiring use of multiple libraries and tools, scored with branch-coverage tests. Picked over HumanEval because it tests real-world library use rather than self-contained algorithmic puzzles. Hosts the BigCode Models Leaderboard; 500+ GitHub stars on the harness, plus 1k+ stars on the sister bigcode-evaluation-harness for code-completion models.

BigCodeBench v0.2.4 (Mar 2 2025); 1140 software-engineering programming tasks (Complete + Instruct splits); active HF leaderboard, 163 models evaluated. Confirmed live June 2026. Borderline: it is primarily a single benchmark+its harness; the dataset itself maps to benchmark_eval_data, but the runnable harness fits here.

Openness

5 high confidence
5.0
license
Apache-2.0(OSI)
source
public(GitHub)

Apache-2.0, fully public; OSI -> open_source.

Adoption

2 medium confidence
2.0

Narrow, specialist code-gen benchmark: 503 GitHub stars, active HF leaderboard with 163 models evaluated. No clean PyPI/download figure for the harness; no published install volume. stars_fallback -> capped at level 3, set to 2 on the modest star + leaderboard footprint.

Capability

3 medium confidence
3.0

Solid, rigorous code-gen eval but single-benchmark scope -- below the broad multi-task harnesses (lm-eval, OpenCompass) on coverage breadth.

Unchanged since 2026-06-09 (last edited, not re-checked)