SMUGRI data
TartuNLP, University of TartuSMUGRI data is a news-based monolingual text collection for seven low-resource Finno-Ugric languages: Komi, Udmurt, Moksha, Erzya, Mansi, Khanty and Livvi Karelian, about 130,000 rows in all. It accompanies a paper on machine translation for these languages presented at the 24th Nordic Conference on Computational Linguistics. TartuNLP at the University of Tartu publishes it.
The card names the languages and the news origin only; the outlets and collection method were not found on any page read.
Openness
4 high confidence- license
- cc-by-4.0
- access
- public(Hugging Face gated: false.)
- dataset_card
- partial(One-line summary and citation
The data is released under the Creative Commons Attribution license, with no gate on the download. The card says only that the texts are news, with no word on which outlets were collected or how.
- https://huggingface.co/api/datasets/tartuNLP/smugri-data recorded 2026-09-24
"gated":false, license cc-by-4.0
- https://huggingface.co/datasets/tartuNLP/smugri-data/raw/main/README.md recorded 2026-09-24
license: cc-by-4.0; "News-based datasets for Komi, Udmurt, Moksha, Erzya, Mansi, Khanty, and Livvi Karelian."
Adoption
1 high confidenceHugging Face downloads of the dataset repository over the trailing 30 days. A download is a file fetch rather than a trained model, and copies mirrored elsewhere are not counted.
- https://huggingface.co/api/datasets/tartuNLP/smugri-data recorded 2026-09-24
133 downloads in the trailing 30 days for tartuNLP/smugri-data
Capability
1 medium confidenceSMUGRI data gives seven Finno-Ugric languages a news collection with a one-line card, documented only through the machine translation paper it supports, and no model on Hugging Face declares it as training data. At a few tens of thousands of news texts it is tiny beside English pretraining corpora of trillions of tokens, and the Northern Sámi web corpus, a comparable small Uralic crawl, at least explains its collection method.
- https://datasets-server.huggingface.co/size?dataset=tartuNLP/smugri-data recorded 2026-09-24
num_rows 129863 over configs kca, kv, mdf, mns, myv, olo, udm.
- https://huggingface.co/api/models?filter=dataset:tartuNLP/smugri-data recorded 2026-09-24
Empty list: no models tagged with the dataset.
Verified 2026-09-24