AI Potluck
Back to Gap Map Model components / Language-specific datasets

SMUGRI data

TartuNLP, University of Tartu
open / Overall score: 1.0

SMUGRI data is a news-based monolingual text collection for seven low-resource Finno-Ugric languages: Komi, Udmurt, Moksha, Erzya, Mansi, Khanty and Livvi Karelian, about 130,000 rows in all. It accompanies a paper on machine translation for these languages presented at the 24th Nordic Conference on Computational Linguistics. TartuNLP at the University of Tartu publishes it.

The card names the languages and the news origin only; the outlets and collection method were not found on any page read.

Openness

4 high confidence
4.0
license
cc-by-4.0
access
public(Hugging Face gated: false.)
dataset_card
partial(One-line summary and citation

The data is released under the Creative Commons Attribution license, with no gate on the download. The card says only that the texts are news, with no word on which outlets were collected or how.

Adoption

1 high confidence
1.0

Hugging Face downloads of the dataset repository over the trailing 30 days. A download is a file fetch rather than a trained model, and copies mirrored elsewhere are not counted.

Capability

1 medium confidence
1.0

SMUGRI data gives seven Finno-Ugric languages a news collection with a one-line card, documented only through the machine translation paper it supports, and no model on Hugging Face declares it as training data. At a few tens of thousands of news texts it is tiny beside English pretraining corpora of trillions of tokens, and the Northern Sámi web corpus, a comparable small Uralic crawl, at least explains its collection method.

Verified 2026-09-24