TartuNLP, University of Tartu
lab · EstoniaScores
1 product on the map — 1 open.
Openness
4 high confidence- license
- cc-by-4.0
- access
- public(Hugging Face gated: false.)
- dataset_card
- partial(One-line summary and citation
The data is released under the Creative Commons Attribution license, with no gate on the download. The card says only that the texts are news, with no word on which outlets were collected or how.
- https://huggingface.co/api/datasets/tartuNLP/smugri-data recorded 2026-09-24
"gated":false, license cc-by-4.0
- https://huggingface.co/datasets/tartuNLP/smugri-data/raw/main/README.md recorded 2026-09-24
license: cc-by-4.0; "News-based datasets for Komi, Udmurt, Moksha, Erzya, Mansi, Khanty, and Livvi Karelian."
Adoption
1 high confidenceHugging Face downloads of the dataset repository over the trailing 30 days. A download is a file fetch rather than a trained model, and copies mirrored elsewhere are not counted.
- https://huggingface.co/api/datasets/tartuNLP/smugri-data recorded 2026-09-24
133 downloads in the trailing 30 days for tartuNLP/smugri-data
Capability
1 medium confidenceSMUGRI data gives seven Finno-Ugric languages a news collection with a one-line card, documented only through the machine translation paper it supports, and no model on Hugging Face declares it as training data. At a few tens of thousands of news texts it is tiny beside English pretraining corpora of trillions of tokens, and the Northern Sámi web corpus, a comparable small Uralic crawl, at least explains its collection method.
- https://datasets-server.huggingface.co/size?dataset=tartuNLP/smugri-data recorded 2026-09-24
num_rows 129863 over configs kca, kv, mdf, mns, myv, olo, udm.
- https://huggingface.co/api/models?filter=dataset:tartuNLP/smugri-data recorded 2026-09-24
Empty list: no models tagged with the dataset.