AI Potluck
Model components / Benchmark / eval datasets

MultiPL-E

Northeastern University / Brown / Wellesley

Multi-programming language translation of HumanEval and MBPP into 18+ languages (JavaScript, TypeScript, C++, Java, Rust, Go, etc.). Enables comparing code generation performance across languages rather than just Python. Published at TOCS 2023. The standard benchmark for evaluating multilingual code models, used in the BigCode and Open LLM Leaderboard evaluations.

MultiPL-E: translations of HumanEval (161 problems) and MBPP (397 problems) into 22-23 programming languages (Ada, Clojure, C++, C#, D, Dart, Elixir, Go, Haskell, Java, Julia, JS, Lua, OCaml, PHP, Perl, R, Ruby, Racket, Rust, Scala, Shell, Swift, TS). HF dataset (nuprl/MultiPL-E, 47 subsets) verified live June 2026; integrated into the BigCode evaluation harness. arXiv 2208.08227 (IEEE TSE).

Openness

4 medium confidence
4.0
data
downloadable(HF parquet, 47 subsets)
license
MIT(on HF dataset card)
datasheet
present(card documents prompts/doctests/tests/stop-tokens)
redistributable
yes(HF)
caveat
GitHub repo LICENSE is a non-OSI 'BSD-3-Clause + no-ML-training' modified license, conflicting with the MIT on the HF data artifact

The scored artifact (HF dataset) is MIT and openly downloadable with a datasheet, open class. But the source repo carries a modified BSD-3 with an explicit 'may not be used as training data for ML' clause (non-OSI), a conflict worth flagging; scored on the data artifact's MIT terms, capped at 4 for the licensing inconsistency.

Adoption

3 low confidence
3.0

The standard multilingual code-generation benchmark: integrated into the BigCode Code Generation LM Harness and HF's multilingual code-model evaluation, cited in multilingual code-model release reports (e.g. StarCoder2/CodeQwen). HF page did not surface a clean monthly-download count this run (showed likes only), and GitHub stars (~308) are last-resort and cannot exceed 3. Level 3 on harness-adoption / release-report traction; no verified usage_volume number, hence low confidence.

Capability

not assessed

Dataset/translation framework, not capability-scored.

Unchanged since 2026-06-24 (last edited, not re-checked)