AI Potluck
Back to Gap Map Model components / Language-specific datasets

CorpusNós

Proxecto Nós
open / Overall score: 1.7

CorpusNós is a Galician pretraining corpus of about 1.9 billion tokens in 7.9 million documents, covering books, research articles, press and blogs, government texts, encyclopedic data, web crawls and translation corpora. It is split into public data and material obtained through transfer agreements with rights holders. Proxecto Nós builds it within the Spanish ILENIA program.

Openness

4 high confidence
4.0
license
mixed-per-subset(Card: TED2020 keeps CC BY-NC-ND 4.0, mC4 Apache 2.0, OSCAR CC0
"All other subcorpora that do not have a previously established original license are released under CC BY-SA 4.0." Hugging Face metadata
other.)
access
public(Hugging Face gated: false, including the transfer-agreement subcorpus.)
dataset_card
present(Per-genre token and document counts for both subcorpora

The whole corpus downloads without a gate, the transfer-agreement material included, and most of it is released under CC BY-SA 4.0. A few imported subcorpora keep their original terms, one of them non-commercial with no derivatives allowed.

Adoption

1 high confidence
1.0

Hugging Face downloads of the dataset repository over the trailing 30 days. A download is a file fetch rather than a trained model, and copies mirrored elsewhere are not counted.

Capability

2 medium confidence
2.0

CorpusNós is the Galician pretraining corpus behind the Carballo and Carvalho models from Proxecto Nós, documented in a PROPOR paper and a card with per-genre counts, and it adds books, press and research writing from rights holders under transfer agreements to its web crawls. Its roughly two billion tokens are a small fraction of the trillions in English pretraining corpora.

Verified 2026-09-24