Lucie 7B
LinagoraLinagora and OpenLLM-France's fully open multilingual 7B foundation model on the Llama 3.1 architecture, trained on roughly three trillion tokens with about equal French and English shares plus German, Spanish, Italian and code. The weights ship with the Lucie Training Dataset, the Megatron-DeepSpeed training code and intermediate checkpoints as repository branches.
Instruct variants are folded into this record. The corpus is published and ungated but non-commercial; the data dimension asks whether a corpus is released rather than on what terms, which is why the score holds while the project's own compliance claim does not. Successor: Luciole. Verified 2026-08-13 via the model card, the dataset card and the Lucie-Training repository.
Openness
5 high confidence- weights
- open(Apache-2.0)
- data
- open(Lucie-Training-Dataset on HF, redistributed in training form)
- code
- open(Lucie-Training Megatron-DeepSpeed fork)
- checkpoints
- open(intermediate HF revisions)
- license
- Apache-2.0(OSI)
Full open-source tier alongside OLMo and Apertus: open weights, open training data, open training code and intermediate checkpoints. The 7B card reads apache-2.0 with step0005000-style checkpoint branches, and Lucie-Training is the published pretraining code. The authors position Lucie as compliant with the OSI definition of open-source AI. One caveat: Lucie-Training-Dataset is licensed cc-by-nc-sa-4.0, so the corpus is published and ungated but not commercially reusable. This axis asks whether the corpus is released, not on what terms, so the score is unaffected - but the OSI-compliance claim does not survive a strict reading, and Lucie is a weaker full-openness exemplar than OLMo, whose Dolma corpus is odc-by.
- https://huggingface.co/api/datasets/OpenLLM-France/Lucie-Training-Dataset recorded 2026-08-13
`gated: false`, `private: false`, 2,823 downloads, 38 likes. Card license is `cc-by-nc-sa-4.0` - the corpus is published and ungated, but non-commercial. See the note.
- https://huggingface.co/OpenLLM-France/Lucie-7B recorded 2026-08-13
License apache-2.0, safetensors weights, links to Lucie-Training-Dataset, and intermediate checkpoints as repo branches (`step0005000`, `step0010000`, ...); `isGated: false`.
- https://huggingface.co/datasets/OpenLLM-France/Lucie-Training-Dataset recorded 2026-08-13
Dataset page for the multilingual pretraining corpus: `isGated: false`, `isPrivate: false`, Parquet format present. The URL needs its `/datasets/` path segment - without it the Hub resolves the path as a model repo and answers 401.
- https://github.com/OpenLLM-France/Lucie-Training recorded 2026-08-13
Repository described as "Code for continual pretraining of LUCIE" - the Megatron-DeepSpeed fork used for the run.
- https://arxiv.org/abs/2503.12294 recorded 2026-08-13
Paper claiming OSI-compliant open resources for multilingual generation. The claim is the authors', which is why it establishes no dimension by itself.
Adoption
1 medium confidence1,945 downloads in the trailing 30 days across the two declared artifacts (OpenLLM-France/Lucie-7B 1,315; OpenLLM-France/Lucie-7B-Instruct-v1.1 630), which bands at level 1 (<10K) on the model adoption scale.
- https://huggingface.co/api/models/OpenLLM-France/Lucie-7B recorded 2026-08-13
1,315 downloads in the trailing 30 days for OpenLLM-France/Lucie-7B
- https://huggingface.co/api/models/OpenLLM-France/Lucie-7B-Instruct-v1.1 recorded 2026-08-13
630 downloads in the trailing 30 days for OpenLLM-France/Lucie-7B-Instruct-v1.1
Capability
2 medium confidence7B multilingual base prioritizing French cultural representation and data-rights constraints — competitive for its size/language mix on French/English evals in the paper, but well below the 2026 open-weight frontier (trillion-param MoEs). Capability 2 reflects small-scale fully-open sovereign model, not frontier performance. Value is openness + FR/EU language coverage.
- https://arxiv.org/abs/2503.12294 recorded 2026-08-13
Lucie-7B / Instruct eval narrative vs contemporary open baselines
- https://huggingface.co/OpenLLM-France/Lucie-7B recorded 2026-08-13
training/eval learning curves and multilingual benchmark plots on the model card
Verified 2026-08-13