Icelandic Gigaword Corpus
Árni Magnússon Institute for Icelandic StudiesThe Icelandic Gigaword Corpus (IGC) is a tagged and lemmatized corpus of Icelandic drawn from news media, parliamentary speeches back to 1911, laws, court judgments, journals, social media, books and Wikipedia. Its 2022 version holds about 2.4 billion running words, extended by yearly releases. A JSONL release of the openly licensed subcorpora, about 1.46 billion words, serves model training. The Árni Magnússon Institute for Icelandic Studies builds it.
Openness
5 high confidence- license
- cc-by-4.0(Applies to the Hugging Face release. IGC-Books and IGC-News2 (913.6M words in IGC-2022) are under the restricted IGC (MIM) license, which forbids republishing the texts, and are not in the repository
- access
- public(Hugging Face gated: false. Restricted subcorpora require accepting a user license with email registration.)
- dataset_card
- present(Subcorpus sizes, format, domains and quality grading
The Hugging Face release holds only the subcorpora published under the Creative Commons Attribution license, with no gate on the download. Published books and part of the news, over a third of the full corpus, sit under a custom license that allows building models but forbids republishing the text, so they are left out.
- https://huggingface.co/api/datasets/arnastofnun/IGC-2024 recorded 2026-09-24
"gated":false, license cc-by-4.0
- https://huggingface.co/datasets/arnastofnun/IGC-2024/raw/main/README.md recorded 2026-09-24
"This package contains those subcorpora ... that have been published with an open licence (CC-BY)"; restricted licence table: IGC-Books 13.8, IGC-News2 899.8 (millions of words).
- https://igc.arnastofnun.is/ recorded 2026-09-24
"The main difference between the licenses is that the texts published under a custom license can not be republished. Both licenses allow building language models"
Adoption
1 high confidenceHugging Face downloads of the dataset repository over the trailing 30 days. A download is a file fetch rather than a trained model, and copies mirrored elsewhere are not counted. The restricted subcorpora and the tagged TEI release on CLARIN-IS are not counted.
- https://huggingface.co/api/datasets/arnastofnun/IGC-2024 recorded 2026-09-24
190 downloads in the trailing 30 days for arnastofnun/IGC-2024
Capability
2 medium confidenceThe IGC is the largest documented corpus of Icelandic, built from news, parliament, law, books and more and released in yearly versions, and Icelandic Dynaword repackages its open parts, though only a community model lists it as training data. With about a billion and a half words in its open release it is far below the trillions of tokens in English pretraining corpora, much like Danish Dynaword for Danish.
- https://huggingface.co/api/models?filter=dataset:arnastofnun/IGC-2024 recorded 2026-09-24
One model tagged with the dataset: Magnuss85/Norn-v3-Icelandic-9B.
- https://huggingface.co/datasets/arnastofnun/IGC-2024/raw/main/README.md recorded 2026-09-24
"It contains around 1,457 million running words."
- https://igc.arnastofnun.is/ recorded 2026-09-24
"IGC 2022 contains texts from until the end of the year 2021, in total 2429 million running words."
Verified 2026-09-24