Latxa Corpus
HiTZ Center, University of the Basque CountryLatxa Corpus is a deduplicated monolingual Basque collection for language model pretraining, combining the EusCrawl web crawl, official gazettes, Basque Parliament transcripts, academic journals, Wikipedia and the Basque portions of CulturaX, HPLT, FineWeb2 and Colossal OSCAR. The version behind the first Latxa models held 4.2 billion tokens in 4.3 million documents. HiTZ and the IXA group at the University of the Basque Country build it.
The entry covers both published versions, v1.1 and v2, as one product line.
Openness
4 high confidence- license
- mixed-per-subset(No dataset-level license on either card or in the Hugging Face metadata
- access
- public(Hugging Face gated: false on both repositories.)
- dataset_card
- present(Sources, per-source document counts and funding on both cards.)
Both repositories download without a gate, but the corpus has no license of its own: each document keeps the terms of the crawl or public dataset it came from. The maintainers removed documents they lacked permission to redistribute, so the public copy is a curated subset of what the models saw.
- https://huggingface.co/api/datasets/HiTZ/latxa-corpus-v1.1 recorded 2026-09-24
"gated":false, no license in cardData
- https://huggingface.co/api/datasets/HiTZ/latxa-corpus-v2 recorded 2026-09-24
"gated":false, no license in cardData
- https://huggingface.co/datasets/HiTZ/latxa-corpus-v1.1/raw/main/README.md recorded 2026-09-24
v1.1 card lists five sources and says to consult each corpus's references for licenses; same permission-compliance notice.
- https://huggingface.co/datasets/HiTZ/latxa-corpus-v2/raw/main/README.md recorded 2026-09-24
"For detailed information regarding the licenses associated with each individual corpus and document ... please refer to the 'license' field within 'meta'"; notice that "Some data points have been removed to ensure permission compliance."
- https://raw.githubusercontent.com/hitz-zentroa/latxa/main/LICENSE recorded 2026-09-24
MIT License, Copyright (c) 2024 HiTZ zentroa; the README describes the repository as training code.
Adoption
1 high confidenceHugging Face downloads over the trailing 30 days, summed across the v1.1 and v2 repositories. A download is a file fetch rather than a trained model, and copies mirrored elsewhere are not counted.
- https://huggingface.co/api/datasets/HiTZ/latxa-corpus-v1.1 recorded 2026-09-24
565 downloads in the trailing 30 days for HiTZ/latxa-corpus-v1.1
- https://huggingface.co/api/datasets/HiTZ/latxa-corpus-v2 recorded 2026-09-24
328 downloads in the trailing 30 days for HiTZ/latxa-corpus-v2
Capability
2 medium confidenceLatxa Corpus is the pretraining set behind the whole Latxa family of Basque models, from the first Llama-based releases to the Qwen versions, joining a new Basque web crawl and institutional text with the Basque parts of large multilingual web corpora, and an ACL paper and both dataset cards document its sources and sizes. Its roughly four billion tokens are ample for Basque but a tiny share of the trillions in English pretraining corpora.
- https://arxiv.org/abs/2403.20266 recorded 2026-09-24
"we continue pretraining on a new Basque corpus comprising 4.3M documents and 4.2B tokens"
- https://huggingface.co/api/models?filter=dataset:HiTZ/latxa-corpus-v1.1 recorded 2026-09-24
Models tagged with the dataset include HiTZ/latxa-7b-v1.1, latxa-13b-v1.1, latxa-70b-v1.1 and the v1.2 releases.
- https://huggingface.co/api/models?filter=dataset:HiTZ/latxa-corpus-v2 recorded 2026-09-24
Models tagged with the dataset include HiTZ/Latxa-Llama-3.1-70B-Instruct-v2, HiTZ/Latxa-Qwen3.5-4B and HiTZ/Latxa-Qwen3.5-2B.
- https://huggingface.co/datasets/HiTZ/latxa-corpus-v1.1/raw/main/README.md recorded 2026-09-24
"This is the training corpus of Latxa v1.1, a family of large language models for Basque based on Llama 2."
- https://huggingface.co/datasets/HiTZ/latxa-corpus-v2/raw/main/README.md recorded 2026-09-24
Data Statistics table: per-source train/valid/test document counts for 14 sources, EusCrawl v2 largest at 1,636,010 train documents.
Verified 2026-09-24