AI Potluck
Back to Gap Map Organization

HiTZ Center, University of the Basque Country

lab · Spain

Scores

1 product on the map — 1 open.

Latxa Corpus

Openness

4 high confidence
4.0
license
mixed-per-subset(No dataset-level license on either card or in the Hugging Face metadata
access
public(Hugging Face gated: false on both repositories.)
dataset_card
present(Sources, per-source document counts and funding on both cards.)

Both repositories download without a gate, but the corpus has no license of its own: each document keeps the terms of the crawl or public dataset it came from. The maintainers removed documents they lacked permission to redistribute, so the public copy is a curated subset of what the models saw.

Adoption

1 high confidence
1.0

Hugging Face downloads over the trailing 30 days, summed across the v1.1 and v2 repositories. A download is a file fetch rather than a trained model, and copies mirrored elsewhere are not counted.

Capability

2 medium confidence
2.0

Latxa Corpus is the pretraining set behind the whole Latxa family of Basque models, from the first Llama-based releases to the Qwen versions, joining a new Basque web crawl and institutional text with the Basque parts of large multilingual web corpora, and an ACL paper and both dataset cards document its sources and sizes. Its roughly four billion tokens are ample for Basque but a tiny share of the trillions in English pretraining corpora.