AI Potluck
Back to Gap Map Model components / Language-specific datasets

Icelandic Dynaword

Danish Foundation Models
open / Overall score: 1.7

Icelandic Dynaword is a continually updated Icelandic pretraining corpus of more than 2.6 billion tokens, gathered only from openly licensed sources: the open subcorpora of the Icelandic Gigaword Corpus, Wikipedia and its sister projects, the official state gazette and out-of-copyright books. Each document carries model-generated quality and content annotations. Danish Foundation Models maintains it, applying the approach of Danish Dynaword.

Openness

4 high confidence
4.0
license
mixed-per-subset(Text keeps each source's license: CC-BY 4.0 (IGC subcorpora and others, 2.52B tokens), CC-BY-SA 4.0 (Wikimedia), Icelandic Copyright Law (state gazette, out-of-copyright books), Gutenberg terms. CC0-1.0, the Hugging Face tag, covers the collection and metadata.)
access
public(Hugging Face gated: false.)
dataset_card
present(Per-source datasheets, domain and license tables.)

Most of the text comes from the Gigaword subcorpora under CC BY 4.0, the rest from Wikimedia projects, out-of-copyright books and the state gazette, and the repository downloads without a gate. There is no single license for the text; each source's datasheet names its own.

Adoption

1 high confidence
1.0

Hugging Face downloads of the dataset repository over the trailing 30 days. A download is a file fetch rather than a trained model, and copies mirrored elsewhere are not counted.

Capability

2 medium confidence
2.0

Icelandic Dynaword documents each source of its openly licensed Icelandic text in the manner of Danish Dynaword, though most of it is the open part of the Icelandic Gigaword Corpus in a new format, and its card states no models have been trained on it. At under three billion tokens it is far below the trillions of tokens in English pretraining corpora, on a par with Danish Dynaword.

Verified 2026-09-24