Norwegian Colossal Corpus
National Library of Norway AI LabThe Norwegian Colossal Corpus (NCC) is a Norwegian pretraining corpus, mainly Bokmål and Nynorsk, drawn from books and newspapers digitized by the National Library, government reports, parliament collections, public-sector websites and Wikipedia, with each document tagged by source type and year. Newspapers distributed under the Språkbank agreement were withdrawn at the request of media houses, leaving about 4.6 billion words. The National Library of Norway AI Lab builds it.
Openness
4 high confidence- license
- mixed-per-subset(License per doc_type: NLOD 2.0 (government, parliament, Målfrid), CC0 1.0 (books, library newspapers), CC BY-NC 2.0 (online newspapers), CC BY-SA 3.0 (Wikipedia, OpenSubtitles). Hugging Face tag: cc. The notram Apache-2.0 LICENSE covers the code.)
- access
- public(Hugging Face gated: false. Språkbank-agreement newspapers no longer distributed since December 2024.)
- dataset_card
- present(Per-source word and document counts, languages, decades, license table
The corpus downloads without a gate, and each document's source type maps to a license, from CC0 and the Norwegian public-data license to a non-commercial license for online newspapers. Users who cannot accept one of them filter that source type out.
- https://huggingface.co/api/datasets/NbAiLab/NCC recorded 2026-09-24
"gated":false, license cc
- https://huggingface.co/datasets/NbAiLab/NCC/raw/main/README.md recorded 2026-09-24
"Various licences applies to different parts of the corpus. Every document in the corpus has a tag telling what doc_type it belongs to. If you are unable to accept any of the licenses, you should filter out the doc_type"
- https://huggingface.co/datasets/NbAiLab/NCC/raw/main/README.md recorded 2026-09-24
"As of December 2024, at the request of media houses, we have ceased distributing newspapers under this agreement, including the Norsk Aviskorpus."
- https://raw.githubusercontent.com/NbAiLab/notram/master/LICENSE recorded 2026-09-24
Apache License, Version 2.0 (repository of corpus scripts and model code).
Adoption
1 high confidenceHugging Face downloads of the dataset repository over the trailing 30 days. A download is a file fetch rather than a trained model, and copies mirrored elsewhere are not counted.
- https://huggingface.co/api/datasets/NbAiLab/NCC recorded 2026-09-24
543 downloads in the trailing 30 days for NbAiLab/NCC
Capability
2 medium confidenceThe Norwegian Colossal Corpus, built largely from OCR of National Library books and newspapers, is behind the library's NB-BERT and NB-GPT-J models and is a training source for NorMistral and NorBLOOM, with an LREC paper and a detailed card. Withdrawing the agreement-licensed newspapers left it under five billion words, down from over seven billion and a small fraction of the trillions of tokens in English pretraining corpora.
- https://arxiv.org/abs/2104.09617 recorded 2026-09-24
arXiv abstract page for "Operationalizing a National Digital Library: The Case for a Norwegian Transformer Model".
- https://huggingface.co/api/models?filter=dataset:NbAiLab/NCC recorded 2026-09-24
Models tagged with the dataset include NbAiLab/nb-gpt-j-6B, ltg/normistral-7b-scratch, ltg/norbloom-7b-scratch and vesteinn/ScandiBERT.
- https://huggingface.co/datasets/NbAiLab/NCC/raw/main/README.md recorded 2026-09-24
Summary table: 4,629,683,886 words, 8,176,399 documents; LREC abstract: "comprises 49GB of clean Norwegian textual data containing over 7B words"
- https://raw.githubusercontent.com/NbAiLab/notram/master/README.md recorded 2026-09-24
Pretrained Models table lists nb-bert-base and other NB models alongside the Norwegian Colossal Corpus.
Verified 2026-09-24