AI Potluck
Back to Gap Map Model components / Language-specific datasets

Icelandic Gigaword Corpus

Árni Magnússon Institute for Icelandic Studies
open / Overall score: 1.7

The Icelandic Gigaword Corpus (IGC) is a tagged and lemmatized corpus of Icelandic drawn from news media, parliamentary speeches back to 1911, laws, court judgments, journals, social media, books and Wikipedia. Its 2022 version holds about 2.4 billion running words, extended by yearly releases. A JSONL release of the openly licensed subcorpora, about 1.46 billion words, serves model training. The Árni Magnússon Institute for Icelandic Studies builds it.

Openness

5 high confidence
5.0
license
cc-by-4.0(Applies to the Hugging Face release. IGC-Books and IGC-News2 (913.6M words in IGC-2022) are under the restricted IGC (MIM) license, which forbids republishing the texts, and are not in the repository
access
public(Hugging Face gated: false. Restricted subcorpora require accepting a user license with email registration.)
dataset_card
present(Subcorpus sizes, format, domains and quality grading

The Hugging Face release holds only the subcorpora published under the Creative Commons Attribution license, with no gate on the download. Published books and part of the news, over a third of the full corpus, sit under a custom license that allows building models but forbids republishing the text, so they are left out.

Adoption

1 high confidence
1.0

Hugging Face downloads of the dataset repository over the trailing 30 days. A download is a file fetch rather than a trained model, and copies mirrored elsewhere are not counted. The restricted subcorpora and the tagged TEI release on CLARIN-IS are not counted.

Capability

2 medium confidence
2.0

The IGC is the largest documented corpus of Icelandic, built from news, parliament, law, books and more and released in yearly versions, and Icelandic Dynaword repackages its open parts, though only a community model lists it as training data. With about a billion and a half words in its open release it is far below the trillions of tokens in English pretraining corpora, much like Danish Dynaword for Danish.

Verified 2026-09-24