AI Potluck
Back to Gap Map Model components / Language-specific datasets

Danish Dynaword

Danish Foundation Models
open / Overall score: 2.3

Danish Dynaword is a continually updated Danish pretraining corpus of more than 9 billion tokens, assembled only from openly licensed sources across legal text, parliament records, news, books, conversation, social media and the web, including the Bornholm and South Jutland dialects. Each source has its own datasheet, and contributors add new ones through a tested submission process. Danish Foundation Models maintains it.

Openness

4 high confidence
4.0
license
mixed-per-subset(Text keeps each source's open license (CC0, CC BY 4.0, CC BY-SA 4.0, attribution licenses, public domain, MIT). CC0-1.0, the Hugging Face license tag, covers the collection and metadata only.)
access
public(Hugging Face gated: false.)
dataset_card
present(Per-source datasheets, domain and license tables, annotation description.)

Every source in the corpus is openly licensed, mostly under CC0, CC BY or CC BY-SA, and the repository downloads without a gate. The text carries no single license, so users check each source's datasheet for its terms.

Adoption

3 high confidence
3.0

Hugging Face downloads of the dataset repository over the trailing 30 days. A download is a file fetch rather than a trained model, and copies mirrored elsewhere are not counted.

Capability

2 medium confidence
2.0

Danish Dynaword gathers openly licensed Danish text with a paper and a datasheet for each source, takes new sources through a tested submission process, and trained Danish Foundation Models' open seven-billion-parameter decoder and a set of Gemma comparison models. At just under ten billion tokens it holds over four times as much as comparable openly licensed Danish releases, but remains small beside English pretraining corpora that run to trillions.

Verified 2026-09-24