Luciole
LinagoraLinagora / OpenLLM-France's fully open multilingual (French–English) LLM family and the successor to Lucie, released in 1B, 8B, and 23B sizes (base + Instruct-1.1). Apache-2.0 weights, the redistributable Luciole-Training-Dataset (~4.65T tokens drawn only from openly licensed corpora), the Luciole-Training pipeline on GitHub, and intermediate + final pretraining checkpoints — clearing the full open-source bar. Base models are pretrained on roughly 30% French data (English 53%, French 16%, plus German/Spanish/Italian/Portuguese/Dutch/Arabic and regional languages); the Instruct-1.1 models are SFT + DPO aligned on math, science, coding, chat, RAG, and translation. Built by LINAGORA and the OpenLLM-France consortium, funded by BPI France through the France 2030 program.
Luciole family (1B / 8B / 23B, base + Instruct-1.1) from LINAGORA / OpenLLM-France, successor to Lucie-7B. Apache-2.0 weights; open Luciole-Training-Dataset (~4.65T tokens, open-licensed corpora only — cleaner terms than Lucie's cc-by-nc-sa corpus); Luciole-Training pipeline on GitHub (GPL-3.0); intermediate + final checkpoints published. Funded by BPI France / France 2030. Collection: https://huggingface.co/collections/OpenLLM-France/luciole-llm. Verified via primary model/dataset cards, August 2026 (HF direct fetch blocked by egress proxy; figures cross-checked via search snippets of the live cards).
Openness
5 high confidence- weights
- open(Apache-2.0, safetensors on HF for 1B/8B/23B base + Instruct-1.1)
- data
- open(Luciole-Training-Dataset on HF, ~4.65T tokens, "only corpora under open licenses")
- code
- open(Luciole-Training pretraining + data-processing pipeline on GitHub, GPL-3.0)
- checkpoints
- open(intermediate + final pretraining checkpoints released)
- license
- Apache-2.0(OSI)
Full open-source tier alongside Lucie / OLMo / Apertus: open weights, open training data, open training code, and intermediate checkpoints. Notably a cleaner full-openness exemplar than its Lucie predecessor — the Luciole-Training-Dataset is described as containing only openly licensed corpora, where Lucie-Training-Dataset carried a non-commercial cc-by-nc-sa license. The weights are Apache-2.0 (OSI) across every distributed size, so the data + code + OSI-license rule lands at 5.
- https://huggingface.co/OpenLLM-France/Luciole-23B-Instruct-1.1 recorded 2026-08-07
Apache-2.0 license; downloadable safetensors weights; fine-tuned + DPO-aligned version of Luciole-23B-Base; developed by LINAGORA and the OpenLLM-France consortium (France 2030 / BPI France)
- https://huggingface.co/datasets/OpenLLM-France/Luciole-Training-Dataset recorded 2026-08-07
released multilingual pretraining corpus (~4.65T tokens; EN 53.4% / FR 16.3% / DE / ES / IT / PT / NL / AR + regional), "contains only corpora under open licenses"
- https://github.com/OpenLLM-France/Luciole-Training recorded 2026-08-07
training pipeline for the Luciole series — data prep, training recipe, and configs (GPL-3.0)
- https://huggingface.co/collections/OpenLLM-France/luciole-llm recorded 2026-08-07
official Luciole LLM collection (1B/8B/23B base + Instruct-1.1, training dataset), with final and intermediate pretraining checkpoints shared for research and continual pretraining
Adoption
1 medium confidenceNiche sovereign / French open-LLM audience, and newer and smaller-reach than Lucie-7B (level 2). Summed monthly HF downloads across the shipped sizes sit in the low thousands (23B ~640, 8B ~944, 1B ~842 as of Aug 2026), well under the 10K level-2 threshold. Adoption may be recomputed from the linked artifacts by the warehouse.
- https://huggingface.co/OpenLLM-France/Luciole-23B-Instruct-1.1 recorded 2026-08-07
~640 downloads/month, 13 likes (as of 2026-08-07)
- https://huggingface.co/collections/OpenLLM-France/luciole-llm recorded 2026-08-07
family of 1B / 8B / 23B (base + Instruct-1.1); per-size monthly downloads in the hundreds
Capability
2 low confidenceMultilingual French–English family up to 23B, prioritizing French cultural representation and open data rights over frontier performance. Larger than Lucie-7B (capability 2) but the same sovereign-model tier — mid-scale and well below the 2026 open-weight frontier (trillion-param MoEs). Positioned like Alia-40b (capability 2); kept conservative pending published benchmarks. Value is openness + FR/EU language coverage.
- https://huggingface.co/OpenLLM-France/Luciole-23B-Base recorded 2026-08-07
23B multilingual base model, ~30% French pretraining data, French/English focus
- https://huggingface.co/collections/OpenLLM-France/luciole-llm recorded 2026-08-07
1B / 8B / 23B base + Instruct-1.1 lineup; Instruct trained via SFT + DPO on math, science, coding, chat, RAG, and translation
Unchanged since 2026-08-07 (last edited, not re-checked)