BigCodeBench
BigCode ProjectEvaluation harness for code generation models on 1,140 practical programming tasks requiring use of multiple libraries and tools, scored with branch-coverage tests. Picked over HumanEval because it tests real-world library use rather than self-contained algorithmic puzzles. Hosts the BigCode Models Leaderboard; 500+ GitHub stars on the harness, plus 1k+ stars on the sister bigcode-evaluation-harness for code-completion models.
BigCodeBench v0.2.4 (Mar 2 2025); 1140 software-engineering programming tasks (Complete + Instruct splits); active HF leaderboard, 163 models evaluated. Confirmed live June 2026. Borderline: it is primarily a single benchmark+its harness; the dataset itself maps to benchmark_eval_data, but the runnable harness fits here.
Openness
5 high confidence- license
- Apache-2.0(OSI)
- source
- public(GitHub)
Apache-2.0, fully public; OSI -> open_source.
- https://github.com/bigcode-project/bigcodebench recorded 2026-06-04
Apache-2.0 license, public harness, 1140 tasks
Adoption
2 medium confidenceNarrow, specialist code-gen benchmark: 503 GitHub stars, active HF leaderboard with 163 models evaluated. No clean PyPI/download figure for the harness; no published install volume. stars_fallback -> capped at level 3, set to 2 on the modest star + leaderboard footprint.
- https://github.com/bigcode-project/bigcodebench recorded 2026-06-04
503 stars; leaderboard with 163 models
- https://huggingface.co/spaces/bigcode/bigcodebench-leaderboard recorded 2026-06-04
active leaderboard ranking evaluated models
Capability
3 medium confidenceSolid, rigorous code-gen eval but single-benchmark scope -- below the broad multi-task harnesses (lm-eval, OpenCompass) on coverage breadth.
- https://github.com/bigcode-project/bigcodebench recorded 2026-06-04
1140 tasks, Complete/Instruct splits, execution-based eval
Unchanged since 2026-06-09 (last edited, not re-checked)