Icelandic Dynaword
Danish Foundation ModelsIcelandic Dynaword is a continually updated Icelandic pretraining corpus of more than 2.6 billion tokens, gathered only from openly licensed sources: the open subcorpora of the Icelandic Gigaword Corpus, Wikipedia and its sister projects, the official state gazette and out-of-copyright books. Each document carries model-generated quality and content annotations. Danish Foundation Models maintains it, applying the approach of Danish Dynaword.
Openness
4 high confidence- license
- mixed-per-subset(Text keeps each source's license: CC-BY 4.0 (IGC subcorpora and others, 2.52B tokens), CC-BY-SA 4.0 (Wikimedia), Icelandic Copyright Law (state gazette, out-of-copyright books), Gutenberg terms. CC0-1.0, the Hugging Face tag, covers the collection and metadata.)
- access
- public(Hugging Face gated: false.)
- dataset_card
- present(Per-source datasheets, domain and license tables.)
Most of the text comes from the Gigaword subcorpora under CC BY 4.0, the rest from Wikimedia projects, out-of-copyright books and the state gazette, and the repository downloads without a gate. There is no single license for the text; each source's datasheet names its own.
- https://huggingface.co/api/datasets/danish-foundation-models/icelandic-dynaword recorded 2026-09-24
"gated":false, license cc0-1.0
- https://huggingface.co/datasets/danish-foundation-models/icelandic-dynaword/raw/main/README.md recorded 2026-09-24
License table: CC-BY 4.0 2.52B tokens; Icelandic Copyright Law 102.80M; CC-BY-SA 4.0 47.70M; "The collection of datasets (metadata, quality control, etc.) is licensed under CC-0."
Adoption
1 high confidenceHugging Face downloads of the dataset repository over the trailing 30 days. A download is a file fetch rather than a trained model, and copies mirrored elsewhere are not counted.
- https://huggingface.co/api/datasets/danish-foundation-models/icelandic-dynaword recorded 2026-09-24
799 downloads in the trailing 30 days for danish-foundation-models/icelandic-dynaword
Capability
2 medium confidenceIcelandic Dynaword documents each source of its openly licensed Icelandic text in the manner of Danish Dynaword, though most of it is the open part of the Icelandic Gigaword Corpus in a new format, and its card states no models have been trained on it. At under three billion tokens it is far below the trillions of tokens in English pretraining corpora, on a par with Danish Dynaword.
- https://huggingface.co/api/models?filter=dataset:danish-foundation-models/icelandic-dynaword recorded 2026-09-24
Empty list: no models tagged with the dataset.
- https://huggingface.co/datasets/danish-foundation-models/icelandic-dynaword/raw/main/README.md recorded 2026-09-24
"Models: Currently there is no models trained on this dataset"; "Number of tokens (Llama 3): 2.67B"
Verified 2026-09-24