MultiPL-E
Northeastern University / Brown / WellesleyMulti-programming language translation of HumanEval and MBPP into 18+ languages (JavaScript, TypeScript, C++, Java, Rust, Go, etc.). Enables comparing code generation performance across languages rather than just Python. Published at TOCS 2023. The standard benchmark for evaluating multilingual code models, used in the BigCode and Open LLM Leaderboard evaluations.
MultiPL-E: translations of HumanEval (161 problems) and MBPP (397 problems) into 22-23 programming languages (Ada, Clojure, C++, C#, D, Dart, Elixir, Go, Haskell, Java, Julia, JS, Lua, OCaml, PHP, Perl, R, Ruby, Racket, Rust, Scala, Shell, Swift, TS). HF dataset (nuprl/MultiPL-E, 47 subsets) verified live June 2026; integrated into the BigCode evaluation harness. arXiv 2208.08227 (IEEE TSE).
Openness
4 medium confidence- data
- downloadable(HF parquet, 47 subsets)
- license
- MIT(on HF dataset card)
- datasheet
- present(card documents prompts/doctests/tests/stop-tokens)
- redistributable
- yes(HF)
- caveat
- GitHub repo LICENSE is a non-OSI 'BSD-3-Clause + no-ML-training' modified license, conflicting with the MIT on the HF data artifact
The scored artifact (HF dataset) is MIT and openly downloadable with a datasheet, open class. But the source repo carries a modified BSD-3 with an explicit 'may not be used as training data for ML' clause (non-OSI), a conflict worth flagging; scored on the data artifact's MIT terms, capped at 4 for the licensing inconsistency.
- https://huggingface.co/datasets/nuprl/MultiPL-E recorded 2026-06-04
MIT license, 47 subsets across 22-23 languages, downloadable parquet, documented schema
- https://github.com/nuprl/MultiPL-E/blob/main/LICENSE recorded 2026-06-04
modified BSD-3-Clause with no-ML-training-data restriction (non-OSI)
Adoption
3 low confidenceThe standard multilingual code-generation benchmark: integrated into the BigCode Code Generation LM Harness and HF's multilingual code-model evaluation, cited in multilingual code-model release reports (e.g. StarCoder2/CodeQwen). HF page did not surface a clean monthly-download count this run (showed likes only), and GitHub stars (~308) are last-resort and cannot exceed 3. Level 3 on harness-adoption / release-report traction; no verified usage_volume number, hence low confidence.
- https://huggingface.co/datasets/nuprl/MultiPL-E recorded 2026-06-04
BigCode harness integration; multilingual code eval standard
- https://github.com/nuprl/MultiPL-E recorded 2026-06-04
BigCode harness integration; ~308 stars (last-resort signal)
Capability
not assessedDataset/translation framework, not capability-scored.
Unchanged since 2026-06-24 (last edited, not re-checked)