AI Potluck
Model components / Base / pretrained models

Lucie 7B

Linagora

Linagora / OpenLLM-France's fully open multilingual 7B foundation model (Llama 3.1 architecture), trained on ~3T tokens with roughly equal French and English shares plus German, Spanish, Italian, and code. Apache-2.0 weights, redistributable Lucie Training Dataset, Megatron-DeepSpeed training code, and intermediate checkpoints — one of the few European releases that clears the full open-source bar. Instruct variants (v1.1 and human-data) are folded into this entry.

Lucie-7B base + Instruct-v1.1 / Instruct-human-data. Apache-2.0 weights; training data at OpenLLM-France/Lucie-Training-Dataset (not a separate map product for now); training code AGPL/community stack on GitHub. Paper arXiv:2503.12294. Collection https://huggingface.co/collections/OpenLLM-France/lucie-llm.

Openness

5 high confidence
5.0
weights
open(Apache-2.0)
data
open(Lucie-Training-Dataset on HF, redistributed in training form)
code
open(Lucie-Training Megatron-DeepSpeed fork)
checkpoints
open(intermediate HF revisions)
license
Apache-2.0(OSI)

Full open-source tier alongside OLMo / Apertus: open weights, open training data, open training code, intermediate checkpoints. Authors position Lucie as OSI-definition-compliant for open-source AI. Training-dataset product deferred — artifacts cited here for openness only. Caveat found on the 2026-07-30 re-read: Lucie-Training-Dataset carries `cc-by-nc-sa-4.0`, so the corpus is published and ungated but NOT commercially reusable. The `data` dimension asks whether the corpus is released, not under what terms, so the score is unchanged — but the authors' OSI-compliance claim does not survive a strict reading, and Lucie is a weaker full-openness exemplar than OLMo, whose Dolma is `odc-by`.

Adoption

2 medium confidence
2.0

Niche sovereign/French open-LLM audience. HF Lucie-7B ~780 downloads/month and 31 likes (Jul 2026); collection upvotes modest. Comparable adoption tier to Apertus / Falcon-class fully-open European releases rather than Llama/Qwen scale.

Capability

2 medium confidence
2.0

7B multilingual base prioritizing French cultural representation and data-rights constraints — competitive for its size/language mix on French/English evals in the paper, but well below the 2026 open-weight frontier (trillion-param MoEs). Capability 2 reflects small-scale fully-open sovereign model, not frontier performance. Value is openness + FR/EU language coverage.

Verified 2026-07-30