CATalog
Projecte AinaCATalog is a Catalan pretraining corpus of about 17.45 billion words drawn from 26 sources, including filtered Common Crawl snapshots, news media, forums, digital libraries and public institutions. Every document carries a heuristic quality score from the CURATE pipeline, so users can filter at their own threshold instead of a fixed cut. The Language Technologies Unit at the Barcelona Supercomputing Center builds it for Projecte AINA.
Openness
4 high confidence- license
- mixed-per-subset(Card: "The dataset as a whole does not carry a unified license"
- access
- public(Hugging Face gated: false.)
- dataset_card
- present(Per-source table with license and word count
The corpus downloads without a gate, but it has no license of its own: each of its 26 sources keeps its original terms, which run from CC0 to CC BY-NC-ND and private data sharing agreements. Anyone redistributing it must ship the per-source license list with the data.
- https://huggingface.co/api/datasets/projecte-aina/CATalog recorded 2026-09-24
"gated":false, no license in cardData
- https://huggingface.co/datasets/projecte-aina/CATalog/raw/main/README.md recorded 2026-09-24
"The dataset as a whole does not carry a unified license. The use, reuse, and redistribution of each part of this dataset is governed by the license of its original source."
Adoption
2 high confidenceHugging Face downloads of the dataset repository over the trailing 30 days. A download is a file fetch rather than a trained model, and copies mirrored elsewhere are not counted.
- https://huggingface.co/api/datasets/projecte-aina/CATalog recorded 2026-09-24
3947 downloads in the trailing 30 days for projecte-aina/CATalog
Capability
3 medium confidenceCATalog is the Catalan pretraining corpus behind the Barcelona Supercomputing Center's Salamandra and ALIA models and the earlier FLOR model, its card documents every source with its license and word count, and each document carries a quality score so users can filter at their own threshold. At under twenty billion words it is large for Catalan yet far below the trillions of tokens in English web corpora such as FineWeb.
- https://huggingface.co/api/models?filter=dataset:projecte-aina/CATalog recorded 2026-09-24
Models tagged with the dataset include projecte-aina/FLOR-6.3B, BSC-LT/salamandra-7b, BSC-LT/salamandra-2b and BSC-LT/ALIA-40b.
- https://huggingface.co/datasets/projecte-aina/CATalog/raw/main/README.md recorded 2026-09-24
"It consists of text documents from 26 different sources ... totaling in 17.45 billion words"; num_examples: 34314510.
Verified 2026-09-24