ALIA-40b
Barcelona Supercomputing CenterSpain's national open LLM: a 40.4B-parameter decoder-only transformer pretrained from scratch by the Barcelona Supercomputing Center on the MareNostrum 5 supercomputer, under Spain's National Language Technology Plan. It was trained on 9.37 trillion tokens across 35 European languages and 92 programming languages, focusing on Spanish and the co-official Catalan, Valencian, Basque and Galician, with 48 layers, a 32,768-token context and a 256k vocabulary.
The corpus is documented across more than sixty named sources but will not be distributed, which is the one dimension keeping this below the Apertus and OLMo tier. An instruct variant and the earlier Salamandra family are not separately scored. Verified 2026-08-13 via the model card and the langtech-bsc/alia repository.
Openness
4 high confidence- weights
- open(Apache-2.0, safetensors on HF BSC-LT/ALIA-40b)
- data
- documented-not-released(60+ named sources described in the model card, but the pretraining corpus "will not be released or distributed to third parties")
- code
- open(all training scripts + configuration files at github.com/langtech-bsc/alia)
- checkpoints
- not-released
- license
- Apache-2.0(OSI)
A strong open-weights release that stops one notch short of full openness: permissive Apache-2.0 weights plus the complete training code and recipe published on GitHub, but a pretraining corpus that is only documented, through 60-odd named sources, rather than released or reconstructable. That missing open-data component is the bright line keeping it at open weights (4) rather than the open-source tier (5) that OLMo, Pythia and Apertus reach - Apertus gets there by shipping data-reconstruction scripts. No intermediate checkpoints are published either.
- https://huggingface.co/BSC-LT/ALIA-40b recorded 2026-08-13
Apache-2.0 license; downloadable safetensors weights; 40.4B params; 9.37T training tokens; corpus documented across 60+ sources but "will not be released or distributed"; links training code at github.com/langtech-bsc/alia
- https://api.github.com/repos/langtech-bsc/alia recorded 2026-08-13
`private: false`, `archived: false`, license `Apache-2.0`. The training scripts and configuration files the model card points at are published.
- https://github.com/langtech-bsc/alia recorded 2026-06-26
open training scripts and configuration files for the ALIA models
Adoption
1 medium confidenceA publicly funded sovereign model - Spain's national LLM, built by BSC on MareNostrum 5 - with strong institutional and policy interest but modest raw volume. The single declared artifact BSC-LT/ALIA-40b returns 475 Hugging Face downloads in the trailing 30 days and 88 likes, which bands at level 1 (<10K) on the model adoption scale. The niche multilingual focus - Spanish, Spain's co-official languages and 35 European languages - keeps adoption below Apertus's level. Counted across the wider ALIA Kit (instruct, GGUF, Salamandra) the total would be higher, but the base model alone sits in the low band.
- https://huggingface.co/api/models/BSC-LT/ALIA-40b recorded 2026-08-13
475 downloads in the trailing 30 days for BSC-LT/ALIA-40b
- https://huggingface.co/BSC-LT/ALIA-40b recorded 2026-06-26
~59 downloads/month, 88 likes (June 2026)
Capability
2 low confidenceA serious 40B multilingual base model, but smaller than Apertus-70B (capability 3) and below the 2026 frontier on reasoning/knowledge. Positioned for Iberian and European-language coverage and digital sovereignty rather than SOTA capability. Score kept conservative pending published benchmarks.
- https://huggingface.co/BSC-LT/ALIA-40b recorded 2026-08-13
40.43B safetensors parameters, 9.37T training tokens, 35 European languages + 92 programming languages; no benchmark table, with results deferred to a forthcoming technical report
- https://arxiv.org/abs/2502.08489 recorded 2026-08-13
Salamandra technical report (BSC-LT) describing the shared training methodology and benchmarks
Verified 2026-08-13