AI Potluck
Back to Gap Map Organization

National Library of Norway AI Lab

government · Norway

Scores

1 product on the map — 1 open.

Norwegian Colossal Corpus

Openness

4 high confidence
4.0
license
mixed-per-subset(License per doc_type: NLOD 2.0 (government, parliament, Målfrid), CC0 1.0 (books, library newspapers), CC BY-NC 2.0 (online newspapers), CC BY-SA 3.0 (Wikipedia, OpenSubtitles). Hugging Face tag: cc. The notram Apache-2.0 LICENSE covers the code.)
access
public(Hugging Face gated: false. Språkbank-agreement newspapers no longer distributed since December 2024.)
dataset_card
present(Per-source word and document counts, languages, decades, license table

The corpus downloads without a gate, and each document's source type maps to a license, from CC0 and the Norwegian public-data license to a non-commercial license for online newspapers. Users who cannot accept one of them filter that source type out.

Adoption

1 high confidence
1.0

Hugging Face downloads of the dataset repository over the trailing 30 days. A download is a file fetch rather than a trained model, and copies mirrored elsewhere are not counted.

Capability

2 medium confidence
2.0

The Norwegian Colossal Corpus, built largely from OCR of National Library books and newspapers, is behind the library's NB-BERT and NB-GPT-J models and is a training source for NorMistral and NorBLOOM, with an LREC paper and a detailed card. Withdrawing the agreement-licensed newspapers left it under five billion words, down from over seven billion and a small fraction of the trillions of tokens in English pretraining corpora.