MADLAD-400
Allen Institute for AIDocument-level multilingual web corpus derived from Common Crawl and covering 419 languages, published by the Allen Institute for AI. It ships a noisy minimally-filtered variant and a clean quality-filtered variant totalling roughly 38TB, designed to improve coverage of low-resource languages, and underpins the MADLAD-400 machine-translation models.
Verified 2026-08-13 via the allenai/MADLAD-400 dataset metadata on Hugging Face.
Openness
5 high confidence- license
- odc-by(API tag
- gated
- false
- card
- present
Ungated with a permissive open-data license and a complete dataset card.
- https://huggingface.co/api/datasets/allenai/MADLAD-400 recorded 2026-08-13
gated = false, license = odc-by
Adoption
3 high confidence15,089 downloads in the trailing 30 days for allenai/MADLAD-400, which bands at 10K-100K, level 3 on the dataset adoption scale.
- https://huggingface.co/api/datasets/allenai/MADLAD-400 recorded 2026-08-13
15,089 downloads in the trailing 30 days for allenai/MADLAD-400
Capability
4 high confidenceUnderpins Google's MADLAD-400 machine-translation models (10.7B/8B, 2023), which are competitive with NLLB-54B: the 10.7B averages 31.8 BLEU against NLLB-54B's 32.7 on WMT across 400+ languages. That is a strong documented downstream training result rather than a current lead.
- https://proceedings.neurips.cc/paper_files/paper/2023/file/d49042a5d49818711c401d34172f9900-Paper-Datasets_and_Benchmarks.pdf recorded 2026-08-13
MADLAD-400 10.7B MT model averages 31.8 BLEU vs NLLB-54B's 32.7 across 400+ languages
Verified 2026-08-13