AI Potluck
Back to Gap Map Model components / Language-specific datasets

Glot500-c

CIS, LMU Munich
open / Overall score: 3.0

Glot500-c is a multilingual pretraining text corpus covering 511 mostly low-resource languages, assembled from more than 150 existing monolingual and parallel datasets plus crawls of multilingual websites. The public release leaves out sources that cannot be redistributed, and each sentence records its source dataset. LMU Munich's Center for Information and Language Processing built it to train the Glot500-m model.

Openness

4 high confidence
4.0
license
mixed-per-subset(Each sentence keeps its source dataset's license
access
public(Ungated for redistributable sources
dataset_card
present

The redistributable part downloads without a gate, and every sentence keeps the license of the dataset it came from, which ranges from permissive to non-commercial. Sources that forbid redistribution are withheld and offered through a request form.

Adoption

3 high confidence
3.0

Hugging Face downloads over the trailing month for the Glot500 repository. The withheld parts shared by request are not counted.

Capability

3 medium confidence
3.0

The Glot corpus reaches more than five hundred mostly low-resource languages, more than any other open pretraining corpus here, records the source, collection method and license of each part, and trained its builders' Glot encoder, though the public release leaves out sources that cannot be redistributed. It holds a few tens of billions of tokens spread over all those languages, a sliver beside English pretraining sets such as FineWeb that run to trillions.

Verified 2026-09-24