National Library of Norway AI Lab
government · NorwayScores
1 product on the map — 1 open.
Openness
4 high confidence- license
- mixed-per-subset(License per doc_type: NLOD 2.0 (government, parliament, Målfrid), CC0 1.0 (books, library newspapers), CC BY-NC 2.0 (online newspapers), CC BY-SA 3.0 (Wikipedia, OpenSubtitles). Hugging Face tag: cc. The notram Apache-2.0 LICENSE covers the code.)
- access
- public(Hugging Face gated: false. Språkbank-agreement newspapers no longer distributed since December 2024.)
- dataset_card
- present(Per-source word and document counts, languages, decades, license table
The corpus downloads without a gate, and each document's source type maps to a license, from CC0 and the Norwegian public-data license to a non-commercial license for online newspapers. Users who cannot accept one of them filter that source type out.
- https://huggingface.co/api/datasets/NbAiLab/NCC recorded 2026-09-24
"gated":false, license cc
- https://huggingface.co/datasets/NbAiLab/NCC/raw/main/README.md recorded 2026-09-24
"Various licences applies to different parts of the corpus. Every document in the corpus has a tag telling what doc_type it belongs to. If you are unable to accept any of the licenses, you should filter out the doc_type"
- https://huggingface.co/datasets/NbAiLab/NCC/raw/main/README.md recorded 2026-09-24
"As of December 2024, at the request of media houses, we have ceased distributing newspapers under this agreement, including the Norsk Aviskorpus."
- https://raw.githubusercontent.com/NbAiLab/notram/master/LICENSE recorded 2026-09-24
Apache License, Version 2.0 (repository of corpus scripts and model code).
Adoption
1 high confidenceHugging Face downloads of the dataset repository over the trailing 30 days. A download is a file fetch rather than a trained model, and copies mirrored elsewhere are not counted.
- https://huggingface.co/api/datasets/NbAiLab/NCC recorded 2026-09-24
543 downloads in the trailing 30 days for NbAiLab/NCC
Capability
2 medium confidenceThe Norwegian Colossal Corpus, built largely from OCR of National Library books and newspapers, is behind the library's NB-BERT and NB-GPT-J models and is a training source for NorMistral and NorBLOOM, with an LREC paper and a detailed card. Withdrawing the agreement-licensed newspapers left it under five billion words, down from over seven billion and a small fraction of the trillions of tokens in English pretraining corpora.
- https://arxiv.org/abs/2104.09617 recorded 2026-09-24
arXiv abstract page for "Operationalizing a National Digital Library: The Case for a Norwegian Transformer Model".
- https://huggingface.co/api/models?filter=dataset:NbAiLab/NCC recorded 2026-09-24
Models tagged with the dataset include NbAiLab/nb-gpt-j-6B, ltg/normistral-7b-scratch, ltg/norbloom-7b-scratch and vesteinn/ScandiBERT.
- https://huggingface.co/datasets/NbAiLab/NCC/raw/main/README.md recorded 2026-09-24
Summary table: 4,629,683,886 words, 8,176,399 documents; LREC abstract: "comprises 49GB of clean Norwegian textual data containing over 7B words"
- https://raw.githubusercontent.com/NbAiLab/notram/master/README.md recorded 2026-09-24
Pretrained Models table lists nb-bert-base and other NB models alongside the Norwegian Colossal Corpus.