AI Potluck
Model components / Base / pretrained models

ALIA-40b

Barcelona Supercomputing Center

Spain's national open LLM: a 40.4B-parameter decoder-only transformer pretrained from scratch by the Barcelona Supercomputing Center on the MareNostrum 5 supercomputer, under Spain's National Language Technology Plan. It was trained on 9.37 trillion tokens across 35 European languages and 92 programming languages, focusing on Spanish and the co-official Catalan, Valencian, Basque and Galician, with 48 layers, a 32,768-token context and a 256k vocabulary.

The corpus is documented across more than sixty named sources but will not be distributed, which is the one dimension keeping this below the Apertus and OLMo tier. An instruct variant and the earlier Salamandra family are not separately scored. Verified 2026-08-13 via the model card and the langtech-bsc/alia repository.

Openness

4 high confidence
4.0
weights
open(Apache-2.0, safetensors on HF BSC-LT/ALIA-40b)
data
documented-not-released(60+ named sources described in the model card, but the pretraining corpus "will not be released or distributed to third parties")
code
open(all training scripts + configuration files at github.com/langtech-bsc/alia)
checkpoints
not-released
license
Apache-2.0(OSI)

A strong open-weights release that stops one notch short of full openness: permissive Apache-2.0 weights plus the complete training code and recipe published on GitHub, but a pretraining corpus that is only documented, through 60-odd named sources, rather than released or reconstructable. That missing open-data component is the bright line keeping it at open weights (4) rather than the open-source tier (5) that OLMo, Pythia and Apertus reach - Apertus gets there by shipping data-reconstruction scripts. No intermediate checkpoints are published either.

  • https://huggingface.co/BSC-LT/ALIA-40b recorded 2026-08-13

    Apache-2.0 license; downloadable safetensors weights; 40.4B params; 9.37T training tokens; corpus documented across 60+ sources but "will not be released or distributed"; links training code at github.com/langtech-bsc/alia

  • https://api.github.com/repos/langtech-bsc/alia recorded 2026-08-13

    `private: false`, `archived: false`, license `Apache-2.0`. The training scripts and configuration files the model card points at are published.

  • https://github.com/langtech-bsc/alia recorded 2026-06-26

    open training scripts and configuration files for the ALIA models

Adoption

1 medium confidence
1.0

A publicly funded sovereign model - Spain's national LLM, built by BSC on MareNostrum 5 - with strong institutional and policy interest but modest raw volume. The single declared artifact BSC-LT/ALIA-40b returns 475 Hugging Face downloads in the trailing 30 days and 88 likes, which bands at level 1 (<10K) on the model adoption scale. The niche multilingual focus - Spanish, Spain's co-official languages and 35 European languages - keeps adoption below Apertus's level. Counted across the wider ALIA Kit (instruct, GGUF, Salamandra) the total would be higher, but the base model alone sits in the low band.

Capability

2 low confidence
2.0

A serious 40B multilingual base model, but smaller than Apertus-70B (capability 3) and below the 2026 frontier on reasoning/knowledge. Positioned for Iberian and European-language coverage and digital sovereignty rather than SOTA capability. Score kept conservative pending published benchmarks.

  • https://huggingface.co/BSC-LT/ALIA-40b recorded 2026-08-13

    40.43B safetensors parameters, 9.37T training tokens, 35 European languages + 92 programming languages; no benchmark table, with results deferred to a forthcoming technical report

  • https://arxiv.org/abs/2502.08489 recorded 2026-08-13

    Salamandra technical report (BSC-LT) describing the shared training methodology and benchmarks

Verified 2026-08-13