AI Potluck
Model components / Base / pretrained models

Luciole

Linagora

Linagora and OpenLLM-France's fully open French-English model family and the successor to Lucie, released at 1B, 8B and 23B in base and instruct forms. The weights ship with a 4.65-trillion-token training dataset drawn only from openly licensed corpora, the training pipeline, and intermediate and final checkpoints. Base models are pretrained on roughly 53 percent English and 16 percent French alongside German, Spanish, Italian, Portuguese, Dutch and Arabic; the instruct models are aligned with SFT and DPO.

The training corpus carries commercially reusable terms, unlike Lucie's. An earlier read worked from search snippets because direct fetches were blocked; this one fetched the cards. Verified 2026-08-13 via the 23B base and instruct cards, the dataset card and the Luciole-Training repository.

Openness

5 high confidence
5.0
weights
open(Apache-2.0, safetensors on HF for 1B/8B/23B base + Instruct-1.1)
data
open(Luciole-Training-Dataset on HF, ~4.65T tokens, "only corpora under open licenses")
code
open(Luciole-Training pretraining + data-processing pipeline on GitHub, GPL-3.0)
checkpoints
open(intermediate + final pretraining checkpoints released)
license
Apache-2.0(OSI)

Full open-source tier alongside Lucie, OLMo and Apertus: open weights, open training data, open training code and intermediate checkpoints. It is a cleaner full-openness exemplar than its Lucie predecessor - the Luciole-Training-Dataset is described as containing only openly licensed corpora, and its card license is cc-by-sa-4.0, share-alike but commercially reusable, where Lucie-Training-Dataset carried a non-commercial cc-by-nc-sa license. The weights are Apache-2.0 across every distributed size, so open data plus open code plus an OSI license lands at 5. The intermediate checkpoints are concrete: the 23B-Base card documents revision tags every 1,000 steps to 5,000, every 5,000 to 30,000 and every 10,000 beyond, plus one at the end of each training phase, mirrored at dl.labs.linagora.com. Luciole-Training is public under GPL-3.0.

Adoption

1 medium confidence
1.0

7,631 downloads in the trailing 30 days across the six declared artifacts (Luciole-23B-Base 746; Luciole-23B-Instruct-1.1 1,266; Luciole-8B-Base 965; Luciole-8B-Instruct-1.1 1,821; Luciole-1B-Base 693; Luciole-1B-Instruct-1.1 2,140), which bands at level 1 (<10K) on the model adoption scale - a niche sovereign French open-LLM audience. The whole family together is still under the 10K floor for level 2, though it is growing fast enough to cross it within a month or two.

Capability

2 low confidence
2.0

A multilingual French-English family up to 23B that prioritizes French cultural representation and open data rights over frontier performance. Larger than Lucie-7B, which also scores 2, but the same sovereign-model tier: mid-scale and well below the 2026 open-weight frontier of trillion-parameter MoEs, and positioned much like Alia-40b. The 23B-Base card publishes no benchmark table against external models, only training-convergence and evaluation curves, so the conservative 2 rests on scale and positioning rather than on a measured comparison. The value here is openness and French and European language coverage.

Verified 2026-08-13