AI Potluck
Back to Gap Map Model components / Language-specific datasets

CATalog

Projecte Aina
open / Overall score: 2.7

CATalog is a Catalan pretraining corpus of about 17.45 billion words drawn from 26 sources, including filtered Common Crawl snapshots, news media, forums, digital libraries and public institutions. Every document carries a heuristic quality score from the CURATE pipeline, so users can filter at their own threshold instead of a fixed cut. The Language Technologies Unit at the Barcelona Supercomputing Center builds it for Projecte AINA.

Openness

4 high confidence
4.0
license
mixed-per-subset(Card: "The dataset as a whole does not carry a unified license"
access
public(Hugging Face gated: false.)
dataset_card
present(Per-source table with license and word count

The corpus downloads without a gate, but it has no license of its own: each of its 26 sources keeps its original terms, which run from CC0 to CC BY-NC-ND and private data sharing agreements. Anyone redistributing it must ship the per-source license list with the data.

Adoption

2 high confidence
2.0

Hugging Face downloads of the dataset repository over the trailing 30 days. A download is a file fetch rather than a trained model, and copies mirrored elsewhere are not counted.

Capability

3 medium confidence
3.0

CATalog is the Catalan pretraining corpus behind the Barcelona Supercomputing Center's Salamandra and ALIA models and the earlier FLOR model, its card documents every source with its license and word count, and each document carries a quality score so users can filter at their own threshold. At under twenty billion words it is large for Catalan yet far below the trillions of tokens in English web corpora such as FineWeb.

Verified 2026-09-24