AI Potluck
Back to Gap Map Model components / Language-specific datasets

OLDI Seed

Open Language Data Initiative
open / Overall score: 1.0

OLDI Seed is a parallel corpus of 6,193 English Wikipedia sentences translated into several dozen low-resource languages, meant to kick-start machine translation where no large training set exists. It extends the NLLB-Seed data that Meta commissioned from professional translators for No Language Left Behind. The Open Language Data Initiative now maintains it and adds community translations.

Openness

5 high confidence
5.0
license
cc-by-sa-4.0
access
public
dataset_card
present

The translations carry a share-alike license and download from Hugging Face without a gate.

Adoption

1 high confidence
1.0

Hugging Face downloads over the trailing month for the OLDI Seed repository. Earlier NLLB-Seed downloads from Meta's own servers are not counted.

Capability

1 medium confidence
1.0

OLDI Seed gives forty-four languages English Wikipedia sentences translated by people, is documented in the NLLB paper and a detailed card, and was part of the training data for Meta's NLLB translation model. With a few thousand sentences per language it is a seed set rather than a training corpus, far smaller than a single-pair collection such as the IIT Bombay English-Hindi corpus and nowhere near the billions of pairs in the largest parallel collections.

Verified 2026-09-24