AI Potluck
Model components / Training & synthetic datasets

MADLAD-400

Allen Institute for AI

Document-level multilingual web corpus derived from Common Crawl and covering 419 languages, published by the Allen Institute for AI. It ships a noisy minimally-filtered variant and a clean quality-filtered variant totalling roughly 38TB, designed to improve coverage of low-resource languages, and underpins the MADLAD-400 machine-translation models.

Verified 2026-08-13 via the allenai/MADLAD-400 dataset metadata on Hugging Face.

Openness

5 high confidence
5.0
license
odc-by(API tag
gated
false
card
present

Ungated with a permissive open-data license and a complete dataset card.

Adoption

3 high confidence
3.0

15,089 downloads in the trailing 30 days for allenai/MADLAD-400, which bands at 10K-100K, level 3 on the dataset adoption scale.

Capability

4 high confidence
4.0

Underpins Google's MADLAD-400 machine-translation models (10.7B/8B, 2023), which are competitive with NLLB-54B: the 10.7B averages 31.8 BLEU against NLLB-54B's 32.7 on WMT across 400+ languages. That is a strong documented downstream training result rather than a current lead.

Verified 2026-08-13