ALIA-40b
Barcelona Supercomputing CenterSpain's national open LLM: a 40.4B-parameter transformer decoder-only model pretrained from scratch by the Barcelona Supercomputing Center (BSC-LT) on the MareNostrum 5 supercomputer, as part of the ALIA project under Spain's National Language Technology Plan. Trained on 9.37 trillion tokens across 35 European languages plus 92 programming languages, with a focus on Spanish and the co-official languages (Catalan, Valencian, Basque, Galician). Released under Apache 2.0 with downloadable weights and fully published training code, though the pretraining corpus is documented rather than released. Distributed through the broader ALIA Kit (models, corpora, integration tools, docs). Capability is mid-tier and below the 2026 frontier; the value proposition is permissive licensing, multilingual breadth, and digital sovereignty for the ~600M Spanish speakers and Iberian-language communities.
ALIA-40b base model from BSC-LT (Barcelona Supercomputing Center, Language Technologies Lab), Apache 2.0. 40.4B params, 48 layers, 32,768 context, 256k vocab; 9.37T pretraining tokens (+6.3B for context extension). Trained on MareNostrum 5 under Spain's National Language Technology Plan with EU NextGenerationEU funds. Docs portal: langtech-bsc.gitbook.io/alia-kit. Instruct variant (ALIA-40b-instruct) and the earlier Salamandra family exist but are not separately scored. Verified live on Hugging Face, June 2026 (59 downloads/month, 88 likes).
Openness
4 high confidence- weights
- open(Apache-2.0, safetensors on HF BSC-LT/ALIA-40b)
- data
- documented-not-released(60+ named sources described in the model card, but the pretraining corpus "will not be released or distributed to third parties")
- code
- open(all training scripts + configuration files at github.com/langtech-bsc/alia)
- checkpoints
- not-released
- license
- Apache-2.0(OSI)
Strong open-weights release that stops one notch short of full openness: permissive Apache-2.0 weights plus the complete training code/recipe published on GitHub, but the pretraining corpus is only documented (60+ named sources) rather than released or reconstructable. That missing open-data component is the bright line that keeps it at open_weights (4) rather than the OLMo/Pythia/Apertus open_source (5) tier, where Apertus reaches 5 by shipping data-reconstruction scripts. No intermediate checkpoints.
- https://huggingface.co/BSC-LT/ALIA-40b recorded 2026-06-26
Apache-2.0 license; downloadable safetensors weights; 40.4B params; 9.37T training tokens; corpus documented across 60+ sources but "will not be released or distributed"; links training code at github.com/langtech-bsc/alia
- https://github.com/langtech-bsc/alia recorded 2026-06-26
open training scripts and configuration files for the ALIA models
Adoption
1 medium confidencePublic-funded sovereign model (Spain's national LLM, BSC / MareNostrum 5) with strong institutional and policy interest but modest raw volume. HF BSC-LT/ALIA-40b shows ~59 downloads/month and 88 likes (June 2026); the niche multilingual focus (Spanish + co-official + 35 European languages) keeps adoption below Apertus's tier. Cumulative across the wider ALIA Kit (instruct, GGUF, Salamandra) would be higher, but the base model alone sits in the low band.
- https://huggingface.co/BSC-LT/ALIA-40b recorded 2026-06-26
~59 downloads/month, 88 likes (June 2026)
Capability
2 low confidenceA serious 40B multilingual base model, but smaller than Apertus-70B (capability 3) and below the 2026 frontier on reasoning/knowledge. Positioned for Iberian and European-language coverage and digital sovereignty rather than SOTA capability. Score kept conservative pending published benchmarks.
- https://huggingface.co/BSC-LT/ALIA-40b recorded 2026-06-26
40.4B params, 48 layers, 9.37T tokens, 35 European languages + 92 programming languages
- https://arxiv.org/abs/2502.08489 recorded 2026-06-26
Salamandra technical report (BSC-LT) describing the shared training methodology and benchmarks
Unchanged since 2026-07-30 (last edited, not re-checked)