AI Potluck
Model components / Training & synthetic datasets

MADLAD-400

Allen Institute for AI

MADLAD-400 is a document-level multilingual web corpus derived from Common Crawl, covering 419 languages, released by the Allen Institute for AI (AI2). It ships in a noisy (minimally filtered) variant and a clean (quality-filtered) variant, totaling roughly 38TB, and was designed to improve coverage of low-resource languages. It underpins the MADLAD-400 machine-translation models.

Verified live 2026-06-22 via primary sources. Ungated with a permissive open-data license and a complete dataset card.

Openness

5 high confidence
5.0
license
odc-by(API tag
gated
false
card
present

Ungated with a permissive open-data license and a complete dataset card.

Adoption

3 high confidence
3.0

32,957 monthly downloads on Hugging Face, graded on the training-corpus bands.

Capability

4 high confidence
4.0

Underpins Google's MADLAD-400 MT models (10.7B/8B, 2023) competitive with NLLB-54B; a strong documented downstream training result, not the current leader.

Unchanged since 2026-07-04 (last edited, not re-checked)