Glot500-c
CIS, LMU MunichGlot500-c is a multilingual pretraining text corpus covering 511 mostly low-resource languages, assembled from more than 150 existing monolingual and parallel datasets plus crawls of multilingual websites. The public release leaves out sources that cannot be redistributed, and each sentence records its source dataset. LMU Munich's Center for Information and Language Processing built it to train the Glot500-m model.
Openness
4 high confidence- license
- mixed-per-subset(Each sentence keeps its source dataset's license
- access
- public(Ungated for redistributable sources
- dataset_card
- present
The redistributable part downloads without a gate, and every sentence keeps the license of the dataset it came from, which ranges from permissive to non-commercial. Sources that forbid redistribution are withheld and offered through a request form.
- https://huggingface.co/api/datasets/cis-lmu/Glot500 recorded 2026-09-24
gated: false; license tag other.
- https://huggingface.co/datasets/cis-lmu/Glot500/raw/main/README.md recorded 2026-09-24
"We don't own any part of the data. The original source of each sentence of the data is indicated in dataset field." "We license the actual packaging, the metadata and the annotations of these data under the cc0-1.0."
- https://raw.githubusercontent.com/cisnlp/Glot500/main/LICENSE recorded 2026-09-24
Apache License, Version 2.0, copyright Ayyoob Imani, Peiqin Lin.
- https://raw.githubusercontent.com/cisnlp/Glot500/main/README.md recorded 2026-09-24
"certain sources prohibit the redistribution of data. As such, data from these sources is omitted from the published version of Glot500-c." Overview table lists each source's license.
Adoption
3 high confidenceHugging Face downloads over the trailing month for the Glot500 repository. The withheld parts shared by request are not counted.
- https://huggingface.co/api/datasets/cis-lmu/Glot500 recorded 2026-09-24
41846 downloads in the trailing 30 days for cis-lmu/Glot500
Capability
3 medium confidenceThe Glot corpus reaches more than five hundred mostly low-resource languages, more than any other open pretraining corpus here, records the source, collection method and license of each part, and trained its builders' Glot encoder, though the public release leaves out sources that cannot be redistributed. It holds a few tens of billions of tokens spread over all those languages, a sliver beside English pretraining sets such as FineWeb that run to trillions.
- https://arxiv.org/abs/2305.12182 recorded 2026-09-24
"collect and clean Glot500-c, a corpus that covers these 511 languages and allows us to train Glot500-m".
- https://raw.githubusercontent.com/cisnlp/Glot500/main/README.md recorded 2026-09-24
"Glot500-c is a subset of Glot2000-c for over 500 languages, including languages with more than 30,000 sentences."
Verified 2026-09-24