Language Technology Group, University of Oslo
lab · NorwayScores
1 product on the map — 1 open.
Openness
5 high confidence- license
- cc0-1.0
- access
- public(Hugging Face gated: false.)
- dataset_card
- present(Short card: crawl seeding, language filter, extraction and deduplication
The corpus is dedicated to the public domain under CC0 and downloads without a gate. The crawl kept only pages whose robots.txt allowed it, and the card says the license adds no constraints on the crawled content.
- https://huggingface.co/api/datasets/ltg/saami-web recorded 2026-09-24
"gated":false, license cc0-1.0
- https://huggingface.co/datasets/ltg/saami-web/raw/main/README.md recorded 2026-09-24
license: cc0-1.0; "Our license does not impose any additional constraints on the contents of this corpus."
Adoption
1 high confidenceHugging Face downloads of the dataset repository over the trailing 30 days. A download is a file fetch rather than a trained model, and copies mirrored elsewhere are not counted.
- https://huggingface.co/api/datasets/ltg/saami-web recorded 2026-09-24
26 downloads in the trailing 30 days for ltg/saami-web
Capability
1 medium confidenceThe corpus gives Northern Sámi a crawled web collection whose short card explains how the crawl was seeded, filtered and deduplicated, and TurkuNLP lists it among the training data of its Finnish ModernBERT models. At a few tens of millions of tokens at most it is tiny beside English pretraining corpora of trillions of tokens, and far from a documented multi-source corpus such as Danish Dynaword.
- https://datasets-server.huggingface.co/size?dataset=ltg/saami-web recorded 2026-09-24
num_rows: 33064; num_bytes_original_files: 113497001
- https://huggingface.co/api/models?filter=dataset:ltg/saami-web recorded 2026-09-24
Models tagged with the dataset include TurkuNLP/finnish-modernbert-large, -base and -tiny.