OLDI Seed
Open Language Data InitiativeOLDI Seed is a parallel corpus of 6,193 English Wikipedia sentences translated into several dozen low-resource languages, meant to kick-start machine translation where no large training set exists. It extends the NLLB-Seed data that Meta commissioned from professional translators for No Language Left Behind. The Open Language Data Initiative now maintains it and adds community translations.
Openness
5 high confidence- license
- cc-by-sa-4.0
- access
- public
- dataset_card
- present
The translations carry a share-alike license and download from Hugging Face without a gate.
- https://huggingface.co/api/datasets/openlanguagedata/oldi_seed recorded 2026-09-24
gated: false; license tag cc-by-sa-4.0.
- https://huggingface.co/datasets/openlanguagedata/oldi_seed recorded 2026-09-24
"The data, which is licensed under CC BY-SA 4.0, is currently being managed by OLDI".
Adoption
1 high confidenceHugging Face downloads over the trailing month for the OLDI Seed repository. Earlier NLLB-Seed downloads from Meta's own servers are not counted.
- https://huggingface.co/api/datasets/openlanguagedata/oldi_seed recorded 2026-09-24
977 downloads in the trailing 30 days for openlanguagedata/oldi_seed
Capability
1 medium confidenceOLDI Seed gives forty-four languages English Wikipedia sentences translated by people, is documented in the NLLB paper and a detailed card, and was part of the training data for Meta's NLLB translation model. With a few thousand sentences per language it is a seed set rather than a training corpus, far smaller than a single-pair collection such as the IIT Bombay English-Hindi corpus and nowhere near the billions of pairs in the largest parallel collections.
- https://arxiv.org/pdf/2207.04672 recorded 2026-09-24
"NLLB-Seed: Seed training data in 39 languages"; section "Bootstrapping models with NLLB-Seed"; figure: "NLLB-200 Model ... Incorporating NLLB-Seed".
- https://huggingface.co/datasets/openlanguagedata/oldi_seed recorded 2026-09-24
"a parallel corpus which consists of 6,193 sentences sampled from English Wikipedia and translated into 44 languages".
Verified 2026-09-24