AI Potluck
Model components / Base / pretrained models

Lucie 7B

Linagora

Linagora and OpenLLM-France's fully open multilingual 7B foundation model on the Llama 3.1 architecture, trained on roughly three trillion tokens with about equal French and English shares plus German, Spanish, Italian and code. The weights ship with the Lucie Training Dataset, the Megatron-DeepSpeed training code and intermediate checkpoints as repository branches.

Instruct variants are folded into this record. The corpus is published and ungated but non-commercial; the data dimension asks whether a corpus is released rather than on what terms, which is why the score holds while the project's own compliance claim does not. Successor: Luciole. Verified 2026-08-13 via the model card, the dataset card and the Lucie-Training repository.

Openness

5 high confidence
5.0
weights
open(Apache-2.0)
data
open(Lucie-Training-Dataset on HF, redistributed in training form)
code
open(Lucie-Training Megatron-DeepSpeed fork)
checkpoints
open(intermediate HF revisions)
license
Apache-2.0(OSI)

Full open-source tier alongside OLMo and Apertus: open weights, open training data, open training code and intermediate checkpoints. The 7B card reads apache-2.0 with step0005000-style checkpoint branches, and Lucie-Training is the published pretraining code. The authors position Lucie as compliant with the OSI definition of open-source AI. One caveat: Lucie-Training-Dataset is licensed cc-by-nc-sa-4.0, so the corpus is published and ungated but not commercially reusable. This axis asks whether the corpus is released, not on what terms, so the score is unaffected - but the OSI-compliance claim does not survive a strict reading, and Lucie is a weaker full-openness exemplar than OLMo, whose Dolma corpus is odc-by.

Adoption

1 medium confidence
1.0

1,945 downloads in the trailing 30 days across the two declared artifacts (OpenLLM-France/Lucie-7B 1,315; OpenLLM-France/Lucie-7B-Instruct-v1.1 630), which bands at level 1 (<10K) on the model adoption scale.

Capability

2 medium confidence
2.0

7B multilingual base prioritizing French cultural representation and data-rights constraints — competitive for its size/language mix on French/English evals in the paper, but well below the 2026 open-weight frontier (trillion-param MoEs). Capability 2 reflects small-scale fully-open sovereign model, not frontier performance. Value is openness + FR/EU language coverage.

Verified 2026-08-13