AI Potluck
Back to Gap Map Model components / Base / pretrained models

EuroLLM

EuroLLM
open weights / Overall score: 3.0

An open-weights multilingual LLM family from the EuroLLM project, a European consortium led by Unbabel and the DeepSPIN group at Instituto Superior Técnico, built under the UTTER Horizon Europe project. The current 22B flagship, alongside 9B and 1.7B siblings, is trained from scratch on ~4 trillion tokens covering all 24 official EU languages plus 11 more (35 in total), on the MareNostrum5 supercomputer under a EuroHPC grant. The Apache-2.0 weights ship with base and instruction-tuned variants, intermediate per-phase Megatron-Core checkpoints, the EuroWeb web pretraining subset, the EuroBlocks instruction datasets, and an evaluation harness; the full pretraining corpus and training recipe are documented in the technical report but not released.

EuroLLM-22B (models 2512, technical report arXiv:2602.05879, published 2026-02-05) is an open-weights release: Apache-2.0 weights and public per-phase Megatron-Core checkpoints, but the pretraining corpus is only partly released (the EuroWeb-2512 web subset; parallel/curated/synthetic components are documented, not shipped) and no EuroLLM pretraining recipe is published (the cited deep-spin/Megatron-LM-pretrain is a generic year-stale Megatron-LM fork; only the deep-spin/eurollm-eval harness is EuroLLM-specific). All ten shipped SKUs are declared so the warehouse sums adoption over the same set. Verified 2026-09-04 via the EuroLLM-22B-2512 model card, the EuroWeb-2512 dataset card, the technical report arXiv:2602.05879, and the deep-spin/Megatron-LM-pretrain repository.

Openness

3 high confidence
3.0
weights
open(Apache-2.0, BF16 safetensors, ungated)
data
documented-not-released(EuroWeb-2512 web subset released, remaining parallel/curated/synthetic and long-context components documented in the report but withheld)
code
closed(only deep-spin/eurollm-eval evaluation harness is EuroLLM-specific, deep-spin/Megatron-LM-pretrain a generic year-stale Megatron-LM fork with no EuroLLM configs)
checkpoints
open(intermediate Megatron-Core checkpoints per pretraining phase, eurollm-22b/9b/1.7b-mcore-phase1/2/3)
license
Apache-2.0(OSI)

Open weights under an OSI license, but not a full open-source release, so the ladder lands at 3 rather than 5. The weights ship as ungated Apache-2.0 BF16 safetensors and the family publishes intermediate Megatron-Core checkpoints for every pretraining phase (eurollm-22b/9b/1.7b-mcore-phase1/2/3), which is the checkpoints:open leg both OLMo and Apertus record. But the two legs that would lift the score above the category's open-weights centre both fall short. Data: only the EuroWeb-2512 multilingual web subset is released; the technical report itemizes parallel, code, math, 1.7M synthetic-math, Wikipedia, ArXiv, Books, Apollo and Cosmopedia sources plus ~60B long-context tokens that are described but not shipped, so the corpus as a whole is documented-not-released rather than open (the peers at the 4/5 rungs released the whole corpus or full reconstruction scripts). Code: the only EuroLLM-specific code published is the deep-spin/eurollm-eval evaluation harness, and the rubric classes an evaluation harness alone as closed; the report also cites deep-spin/Megatron-LM-pretrain as the pre-training codebase, but that repository is a generic fork of NVIDIA/Megatron-LM whose last commit predates the 2512 release by about a year and carries no EuroLLM training configs, so it contributes nothing to this dimension and no pretraining recipe actually ships. Weights open, data documented, code closed and an OSI license fall through the recipe to 3, open_weights. Apache-2.0 covers the weights only: the released datasets carry no license tag, eurollm-eval has no LICENSE file, and the Megatron fork is NVIDIA-licensed.

  • https://huggingface.co/api/models/utter-project/EuroLLM-22B-2512 recorded 2026-09-04

    `private: false`, `gated: false`, `license: apache-2.0`, `safetensors.total` 22,637,328,384 BF16 parameters, LlamaForCausalLM. Establishes that the weights are downloadable, ungated and OSI-licensed.

  • https://huggingface.co/utter-project/EuroLLM-22B-2512 recorded 2026-09-04

    Model card "Model Card for EuroLLM-22B": "License: Apache License 2.0"; 22B parameter multilingual transformer trained from scratch on ~4T tokens; describes the three pretraining phases (initial 3.6T, annealing 400B, annealing-to-zero 100B) and the data sources (web, parallel, Wikipedia, Arxiv, books, math, code, Apollo) without releasing the non-web components.

  • https://arxiv.org/pdf/2602.05879 recorded 2026-09-04

    Technical report lists the pretraining corpus components (web, parallel, code, math, synthetic math, Wikipedia, ArXiv, Books, Apollo, Cosmopedia, ~60B long-context) and cites github.com/deep-spin/Megatron-LM-pretrain (pre-training) and github.com/deep-spin/eurollm-eval (evaluation) as codebases. Only the web subset (EuroWeb-2512) and the eval harness are actually released; the pretraining repo is a generic Megatron-LM fork (independently verified: last commit 2025-01-06, no EuroLLM configs).

  • https://huggingface.co/api/datasets/utter-project/EuroWeb-2512 recorded 2026-09-04

    `private: false`, `gated: false`, size category 1B<n<10B, no license tag; the released multilingual web pretraining subset for the 2512 release. Establishes that a documented subset of the corpus is released while the remainder is not.

  • https://huggingface.co/api/models/utter-project/eurollm-22b-mcore-phase1 recorded 2026-09-04

    `private: false`, `gated: false`; a public Megatron-Core intermediate checkpoint for the 22B model's first pretraining phase, one of the eurollm-22b/9b/1.7b-mcore-phase1/2/3 series. Establishes checkpoints:open.

Adoption

3 medium confidence
3.0

127,672 downloads in the trailing 30 days summed across the ten declared family SKUs on the utter-project Hugging Face org: EuroLLM-22B-Instruct-2512 31,111; 9B-Instruct-2512 28,779; 1.7B-Instruct 27,659; 9B-Instruct 18,632; 1.7B 11,974; 9B-2512 2,866; 9B 2,832; 22B-2512 2,603; 22B-Instruct-Preview 1,111; 22B-Preview 105. Bands at level 3 (100K-1M) on the model adoption scale; confidence medium because the total sits not far above the 100K floor and download figures move month to month. The product declares exactly these ten SKUs so the warehouse sums the same set.

Capability

3 high confidence
3.0

Mid-tier capability, the same band as Apertus though edging ahead of it on benchmarks: competitive with the prior open generation and unusually strong multilingually across the 35 supported languages, but below the 2026 frontier on reasoning and knowledge. The value proposition is openness and European-language breadth rather than raw SOTA. The EuroLLM-22B technical report's Table 3 measures EuroLLM-22B and the Apertus baselines on the same instruction-tuned English suite, where EuroLLM-22B leads on eight of ten benchmarks and the paper names it "the strongest model among the fully open European systems considered" while "the EuroLLM family still trails the very best open-weights models overall". Kept at 3 rather than lifted to 4: beating Apertus by single-digit margins on most benchmarks places it at the top of the same band, not in a higher capability tier, since it still trails the frontier open-weights.

  • https://arxiv.org/pdf/2602.05879 recorded 2026-09-04

    Table 3: EuroLLM-22B-Instruct English scores (IFEval 67.2, Hellaswag 69.7, MMLU 69.8, MMLU-Pro 50.8, BBH 55.3, ARC-C 89.8, GPQA 26.8, GSM8K 85.5, MATH-500 54.5, HumanEval 53.9) against Apertus, OLMo-3, Llama-3, Gemma-3 and Qwen-3 baselines; prose "EuroLLM-22B is the strongest model among the fully open European systems" and "still trails the very best open-weights models overall".

Verified 2026-09-04