Luciole
LinagoraLinagora and OpenLLM-France's fully open French-English model family and the successor to Lucie, released at 1B, 8B and 23B in base and instruct forms. The weights ship with a 4.65-trillion-token training dataset drawn only from openly licensed corpora, the training pipeline, and intermediate and final checkpoints. Base models are pretrained on roughly 53 percent English and 16 percent French alongside German, Spanish, Italian, Portuguese, Dutch and Arabic; the instruct models are aligned with SFT and DPO.
The training corpus carries commercially reusable terms, unlike Lucie's. An earlier read worked from search snippets because direct fetches were blocked; this one fetched the cards. Verified 2026-08-13 via the 23B base and instruct cards, the dataset card and the Luciole-Training repository.
Openness
5 high confidence- weights
- open(Apache-2.0, safetensors on HF for 1B/8B/23B base + Instruct-1.1)
- data
- open(Luciole-Training-Dataset on HF, ~4.65T tokens, "only corpora under open licenses")
- code
- open(Luciole-Training pretraining + data-processing pipeline on GitHub, GPL-3.0)
- checkpoints
- open(intermediate + final pretraining checkpoints released)
- license
- Apache-2.0(OSI)
Full open-source tier alongside Lucie, OLMo and Apertus: open weights, open training data, open training code and intermediate checkpoints. It is a cleaner full-openness exemplar than its Lucie predecessor - the Luciole-Training-Dataset is described as containing only openly licensed corpora, and its card license is cc-by-sa-4.0, share-alike but commercially reusable, where Lucie-Training-Dataset carried a non-commercial cc-by-nc-sa license. The weights are Apache-2.0 across every distributed size, so open data plus open code plus an OSI license lands at 5. The intermediate checkpoints are concrete: the 23B-Base card documents revision tags every 1,000 steps to 5,000, every 5,000 to 30,000 and every 10,000 beyond, plus one at the end of each training phase, mirrored at dl.labs.linagora.com. Luciole-Training is public under GPL-3.0.
- https://huggingface.co/OpenLLM-France/Luciole-23B-Instruct-1.1 recorded 2026-08-13
Apache-2.0 license; downloadable safetensors weights; fine-tuned + DPO-aligned version of Luciole-23B-Base; developed by LINAGORA and the OpenLLM-France consortium (France 2030 / BPI France)
- https://huggingface.co/api/datasets/OpenLLM-France/Luciole-Training-Dataset recorded 2026-08-13
`gated: false`, `private: false`, 13,136 downloads, card license `cc-by-sa-4.0` - the corpus is published, ungated and commercially reusable, unlike Lucie's cc-by-nc-sa corpus.
- https://huggingface.co/datasets/OpenLLM-France/Luciole-Training-Dataset recorded 2026-08-13
released multilingual pretraining corpus (~4.65T tokens; EN 53.4% / FR 16.3% / DE / ES / IT / PT / NL / AR + regional), "contains only corpora under open licenses"
- https://api.github.com/repos/OpenLLM-France/Luciole-Training recorded 2026-08-13
`private: false`, `archived: false`, license `GPL-3.0`, last pushed 2026-07-31 - the pretraining and data-processing pipeline for the Luciole series.
- https://huggingface.co/OpenLLM-France/Luciole-23B-Base recorded 2026-08-13
"Intermediate checkpoints are released under dedicated revision tags at regular intervals throughout training" - every 1,000 steps to 5,000, every 5,000 to 30,000, every 10,000 beyond, plus one per training phase, all mirrored at dl.labs.linagora.com.
- https://github.com/OpenLLM-France/Luciole-Training recorded 2026-08-13
training pipeline for the Luciole series — data prep, training recipe, and configs (GPL-3.0)
- https://huggingface.co/collections/OpenLLM-France/luciole-llm recorded 2026-08-13
official Luciole LLM collection (1B/8B/23B base + Instruct-1.1, training dataset)
Adoption
1 medium confidence7,631 downloads in the trailing 30 days across the six declared artifacts (Luciole-23B-Base 746; Luciole-23B-Instruct-1.1 1,266; Luciole-8B-Base 965; Luciole-8B-Instruct-1.1 1,821; Luciole-1B-Base 693; Luciole-1B-Instruct-1.1 2,140), which bands at level 1 (<10K) on the model adoption scale - a niche sovereign French open-LLM audience. The whole family together is still under the 10K floor for level 2, though it is growing fast enough to cross it within a month or two.
- https://huggingface.co/api/models/OpenLLM-France/Luciole-23B-Base recorded 2026-08-13
746 downloads in the trailing 30 days for OpenLLM-France/Luciole-23B-Base
- https://huggingface.co/api/models/OpenLLM-France/Luciole-23B-Instruct-1.1 recorded 2026-08-13
1,266 downloads in the trailing 30 days for OpenLLM-France/Luciole-23B-Instruct-1.1
- https://huggingface.co/api/models/OpenLLM-France/Luciole-8B-Base recorded 2026-08-13
965 downloads in the trailing 30 days for OpenLLM-France/Luciole-8B-Base
- https://huggingface.co/api/models/OpenLLM-France/Luciole-8B-Instruct-1.1 recorded 2026-08-13
1,821 downloads in the trailing 30 days for OpenLLM-France/Luciole-8B-Instruct-1.1
- https://huggingface.co/api/models/OpenLLM-France/Luciole-1B-Base recorded 2026-08-13
693 downloads in the trailing 30 days for OpenLLM-France/Luciole-1B-Base
- https://huggingface.co/api/models/OpenLLM-France/Luciole-1B-Instruct-1.1 recorded 2026-08-13
2,140 downloads in the trailing 30 days for OpenLLM-France/Luciole-1B-Instruct-1.1
Capability
2 low confidenceA multilingual French-English family up to 23B that prioritizes French cultural representation and open data rights over frontier performance. Larger than Lucie-7B, which also scores 2, but the same sovereign-model tier: mid-scale and well below the 2026 open-weight frontier of trillion-parameter MoEs, and positioned much like Alia-40b. The 23B-Base card publishes no benchmark table against external models, only training-convergence and evaluation curves, so the conservative 2 rests on scale and positioning rather than on a measured comparison. The value here is openness and French and European language coverage.
- https://huggingface.co/OpenLLM-France/Luciole-23B-Base recorded 2026-08-13
23B multilingual base model, ~30% French pretraining data, French/English focus; the card carries training-convergence and evaluation sections but no comparison table against external models
- https://huggingface.co/collections/OpenLLM-France/luciole-llm recorded 2026-08-13
1B / 8B / 23B base + Instruct-1.1 lineup; Instruct trained via SFT + DPO on math, science, coding, chat, RAG, and translation
Verified 2026-08-13