AI Potluck
Back to Gap Map Organization

Projecte Aina

government · Spain

Scores

1 product on the map — 1 open.

CATalog

Openness

4 high confidence
4.0
license
mixed-per-subset(Card: "The dataset as a whole does not carry a unified license"
access
public(Hugging Face gated: false.)
dataset_card
present(Per-source table with license and word count

The corpus downloads without a gate, but it has no license of its own: each of its 26 sources keeps its original terms, which run from CC0 to CC BY-NC-ND and private data sharing agreements. Anyone redistributing it must ship the per-source license list with the data.

Adoption

2 high confidence
2.0

Hugging Face downloads of the dataset repository over the trailing 30 days. A download is a file fetch rather than a trained model, and copies mirrored elsewhere are not counted.

Capability

3 medium confidence
3.0

CATalog is the Catalan pretraining corpus behind the Barcelona Supercomputing Center's Salamandra and ALIA models and the earlier FLOR model, its card documents every source with its license and word count, and each document carries a quality score so users can filter at their own threshold. At under twenty billion words it is large for Catalan yet far below the trillions of tokens in English web corpora such as FineWeb.