AI Potluck
Back to Gap Map Model components / Language-specific datasets

Northern Sámi web corpus

Language Technology Group, University of Oslo
open / Overall score: 1.0

The Northern Sámi web corpus is a crawl of about 33,000 web documents in Northern Sámi, seeded from the external links of the Sámi Wikipedia and continued breadth-first through pages that GlotLID identified as Northern Sámi and whose robots.txt allowed crawling. Text was extracted with Trafilatura and fuzzy-deduplicated document by document. The Language Technology Group at the University of Oslo publishes it.

Openness

5 high confidence
5.0
license
cc0-1.0
access
public(Hugging Face gated: false.)
dataset_card
present(Short card: crawl seeding, language filter, extraction and deduplication

The corpus is dedicated to the public domain under CC0 and downloads without a gate. The crawl kept only pages whose robots.txt allowed it, and the card says the license adds no constraints on the crawled content.

Adoption

1 high confidence
1.0

Hugging Face downloads of the dataset repository over the trailing 30 days. A download is a file fetch rather than a trained model, and copies mirrored elsewhere are not counted.

Capability

1 medium confidence
1.0

The corpus gives Northern Sámi a crawled web collection whose short card explains how the crawl was seeded, filtered and deduplicated, and TurkuNLP lists it among the training data of its Finnish ModernBERT models. At a few tens of millions of tokens at most it is tiny beside English pretraining corpora of trillions of tokens, and far from a documented multi-source corpus such as Danish Dynaword.

Verified 2026-09-24