AI Potluck
Back to Gap Map Model components / Language-specific datasets

Norwegian Colossal Corpus

National Library of Norway AI Lab
open / Overall score: 1.7

The Norwegian Colossal Corpus (NCC) is a Norwegian pretraining corpus, mainly Bokmål and Nynorsk, drawn from books and newspapers digitized by the National Library, government reports, parliament collections, public-sector websites and Wikipedia, with each document tagged by source type and year. Newspapers distributed under the Språkbank agreement were withdrawn at the request of media houses, leaving about 4.6 billion words. The National Library of Norway AI Lab builds it.

Openness

4 high confidence
4.0
license
mixed-per-subset(License per doc_type: NLOD 2.0 (government, parliament, Målfrid), CC0 1.0 (books, library newspapers), CC BY-NC 2.0 (online newspapers), CC BY-SA 3.0 (Wikipedia, OpenSubtitles). Hugging Face tag: cc. The notram Apache-2.0 LICENSE covers the code.)
access
public(Hugging Face gated: false. Språkbank-agreement newspapers no longer distributed since December 2024.)
dataset_card
present(Per-source word and document counts, languages, decades, license table

The corpus downloads without a gate, and each document's source type maps to a license, from CC0 and the Norwegian public-data license to a non-commercial license for online newspapers. Users who cannot accept one of them filter that source type out.

Adoption

1 high confidence
1.0

Hugging Face downloads of the dataset repository over the trailing 30 days. A download is a file fetch rather than a trained model, and copies mirrored elsewhere are not counted.

Capability

2 medium confidence
2.0

The Norwegian Colossal Corpus, built largely from OCR of National Library books and newspapers, is behind the library's NB-BERT and NB-GPT-J models and is a training source for NorMistral and NorBLOOM, with an LREC paper and a detailed card. Withdrawing the agreement-licensed newspapers left it under five billion words, down from over seven billion and a small fraction of the trillions of tokens in English pretraining corpora.

Verified 2026-09-24