MADLAD-400
Allen Institute for AIMADLAD-400 is a document-level multilingual web corpus derived from Common Crawl, covering 419 languages, released by the Allen Institute for AI (AI2). It ships in a noisy (minimally filtered) variant and a clean (quality-filtered) variant, totaling roughly 38TB, and was designed to improve coverage of low-resource languages. It underpins the MADLAD-400 machine-translation models.
Verified live 2026-06-22 via primary sources. Ungated with a permissive open-data license and a complete dataset card.
Openness
5 high confidence- license
- odc-by(API tag
- gated
- false
- card
- present
Ungated with a permissive open-data license and a complete dataset card.
- https://huggingface.co/api/datasets/allenai/MADLAD-400 recorded 2026-06-22
gated = false, license = odc-by
Adoption
3 high confidence32,957 monthly downloads on Hugging Face, graded on the training-corpus bands.
- https://huggingface.co/api/datasets/allenai/MADLAD-400 recorded 2026-07-04
downloads field = 32,957 (30-day)
Capability
4 high confidenceUnderpins Google's MADLAD-400 MT models (10.7B/8B, 2023) competitive with NLLB-54B; a strong documented downstream training result, not the current leader.
- https://proceedings.neurips.cc/paper_files/paper/2023/file/d49042a5d49818711c401d34172f9900-Paper-Datasets_and_Benchmarks.pdf recorded 2026-07-04
MADLAD-400 10.7B MT model averages 31.8 BLEU vs NLLB-54B's 32.7 across 400+ languages
Unchanged since 2026-07-04 (last edited, not re-checked)