Proxecto Nós
lab · SpainScores
1 product on the map — 1 open.
Openness
4 high confidence- license
- mixed-per-subset(Card: TED2020 keeps CC BY-NC-ND 4.0, mC4 Apache 2.0, OSCAR CC0
- "All other subcorpora that do not have a previously established original license are released under CC BY-SA 4.0." Hugging Face metadata
- other.)
- access
- public(Hugging Face gated: false, including the transfer-agreement subcorpus.)
- dataset_card
- present(Per-genre token and document counts for both subcorpora
The whole corpus downloads without a gate, the transfer-agreement material included, and most of it is released under CC BY-SA 4.0. A few imported subcorpora keep their original terms, one of them non-commercial with no derivatives allowed.
- https://huggingface.co/api/datasets/proxectonos/corpusnos recorded 2026-09-24
"gated":false, license: other
- https://huggingface.co/datasets/proxectonos/corpusnos/raw/main/README.md recorded 2026-09-24
"the following subcorpora retain their original licenses: TED2020: CC BY-NC-ND 4.0; mC4: Apache License 2.0; OSCAR: CC0. All other subcorpora ... are released under CC BY-SA 4.0."
Adoption
1 high confidenceHugging Face downloads of the dataset repository over the trailing 30 days. A download is a file fetch rather than a trained model, and copies mirrored elsewhere are not counted.
- https://huggingface.co/api/datasets/proxectonos/corpusnos recorded 2026-09-24
131 downloads in the trailing 30 days for proxectonos/corpusnos
Capability
2 medium confidenceCorpusNós is the Galician pretraining corpus behind the Carballo and Carvalho models from Proxecto Nós, documented in a PROPOR paper and a card with per-genre counts, and it adds books, press and research writing from rights holders under transfer agreements to its web crawls. Its roughly two billion tokens are a small fraction of the trillions in English pretraining corpora.
- https://aclanthology.org/2024.propor-1.66/ recorded 2026-09-24
ACL Anthology entry: "CorpusNÓS: A massive Galician corpus for training large language models" (PROPOR 2024).
- https://huggingface.co/api/models?filter=dataset:proxectonos/corpusnos recorded 2026-09-24
Models tagged with the dataset include proxectonos/Carballo-bloom-1.3B, proxectonos/Llama-3.1-Carballo and Nos-PT/Llama-Carvalho-GL.
- https://huggingface.co/datasets/proxectonos/corpusnos/raw/main/README.md recorded 2026-09-24
Total: 1,917,196,788 tokens, 7,922,447 documents; "a massive Galician corpus primarily devised for training large language models"