CorpusNós
Proxecto NósCorpusNós is a Galician pretraining corpus of about 1.9 billion tokens in 7.9 million documents, covering books, research articles, press and blogs, government texts, encyclopedic data, web crawls and translation corpora. It is split into public data and material obtained through transfer agreements with rights holders. Proxecto Nós builds it within the Spanish ILENIA program.
Openness
4 high confidence- license
- mixed-per-subset(Card: TED2020 keeps CC BY-NC-ND 4.0, mC4 Apache 2.0, OSCAR CC0
- "All other subcorpora that do not have a previously established original license are released under CC BY-SA 4.0." Hugging Face metadata
- other.)
- access
- public(Hugging Face gated: false, including the transfer-agreement subcorpus.)
- dataset_card
- present(Per-genre token and document counts for both subcorpora
The whole corpus downloads without a gate, the transfer-agreement material included, and most of it is released under CC BY-SA 4.0. A few imported subcorpora keep their original terms, one of them non-commercial with no derivatives allowed.
- https://huggingface.co/api/datasets/proxectonos/corpusnos recorded 2026-09-24
"gated":false, license: other
- https://huggingface.co/datasets/proxectonos/corpusnos/raw/main/README.md recorded 2026-09-24
"the following subcorpora retain their original licenses: TED2020: CC BY-NC-ND 4.0; mC4: Apache License 2.0; OSCAR: CC0. All other subcorpora ... are released under CC BY-SA 4.0."
Adoption
1 high confidenceHugging Face downloads of the dataset repository over the trailing 30 days. A download is a file fetch rather than a trained model, and copies mirrored elsewhere are not counted.
- https://huggingface.co/api/datasets/proxectonos/corpusnos recorded 2026-09-24
131 downloads in the trailing 30 days for proxectonos/corpusnos
Capability
2 medium confidenceCorpusNós is the Galician pretraining corpus behind the Carballo and Carvalho models from Proxecto Nós, documented in a PROPOR paper and a card with per-genre counts, and it adds books, press and research writing from rights holders under transfer agreements to its web crawls. Its roughly two billion tokens are a small fraction of the trillions in English pretraining corpora.
- https://aclanthology.org/2024.propor-1.66/ recorded 2026-09-24
ACL Anthology entry: "CorpusNÓS: A massive Galician corpus for training large language models" (PROPOR 2024).
- https://huggingface.co/api/models?filter=dataset:proxectonos/corpusnos recorded 2026-09-24
Models tagged with the dataset include proxectonos/Carballo-bloom-1.3B, proxectonos/Llama-3.1-Carballo and Nos-PT/Llama-Carvalho-GL.
- https://huggingface.co/datasets/proxectonos/corpusnos/raw/main/README.md recorded 2026-09-24
Total: 1,917,196,788 tokens, 7,922,447 documents; "a massive Galician corpus primarily devised for training large language models"
Verified 2026-09-24