Danish Dynaword
Danish Foundation ModelsDanish Dynaword is a continually updated Danish pretraining corpus of more than 9 billion tokens, assembled only from openly licensed sources across legal text, parliament records, news, books, conversation, social media and the web, including the Bornholm and South Jutland dialects. Each source has its own datasheet, and contributors add new ones through a tested submission process. Danish Foundation Models maintains it.
Openness
4 high confidence- license
- mixed-per-subset(Text keeps each source's open license (CC0, CC BY 4.0, CC BY-SA 4.0, attribution licenses, public domain, MIT). CC0-1.0, the Hugging Face license tag, covers the collection and metadata only.)
- access
- public(Hugging Face gated: false.)
- dataset_card
- present(Per-source datasheets, domain and license tables, annotation description.)
Every source in the corpus is openly licensed, mostly under CC0, CC BY or CC BY-SA, and the repository downloads without a gate. The text carries no single license, so users check each source's datasheet for its terms.
- https://huggingface.co/api/datasets/danish-foundation-models/danish-dynaword recorded 2026-09-24
"gated":false, license cc0-1.0
- https://huggingface.co/datasets/danish-foundation-models/danish-dynaword/raw/main/README.md recorded 2026-09-24
"These license is applied to the constituent data, i.e., the text. The collection of datasets (metadata, quality control, etc.) is licensed under CC-0."
Adoption
3 high confidenceHugging Face downloads of the dataset repository over the trailing 30 days. A download is a file fetch rather than a trained model, and copies mirrored elsewhere are not counted.
- https://huggingface.co/api/datasets/danish-foundation-models/danish-dynaword recorded 2026-09-24
11157 downloads in the trailing 30 days for danish-foundation-models/danish-dynaword
Capability
2 medium confidenceDanish Dynaword gathers openly licensed Danish text with a paper and a datasheet for each source, takes new sources through a tested submission process, and trained Danish Foundation Models' open seven-billion-parameter decoder and a set of Gemma comparison models. At just under ten billion tokens it holds over four times as much as comparable openly licensed Danish releases, but remains small beside English pretraining corpora that run to trillions.
- https://arxiv.org/abs/2508.02271 recorded 2026-09-24
"Danish Dynaword contains over four times as many tokens as comparable releases, is exclusively openly licensed, and has received multiple contributions across industry and research."
- https://huggingface.co/api/models?filter=dataset:danish-foundation-models/danish-dynaword recorded 2026-09-24
Models tagged with the dataset include danish-foundation-models/dfm-decoder-open-v0-7b-pt and gemma-3-1b-cpt-dynaword-full-v1.
- https://huggingface.co/datasets/danish-foundation-models/danish-dynaword/raw/main/README.md recorded 2026-09-24
"Number of tokens (Llama 3): 9.81B"; "Number of samples: 7.40M"
Verified 2026-09-24