AI Potluck
Model components / Base / pretrained models

Luciole

Linagora

Linagora / OpenLLM-France's fully open multilingual (French–English) LLM family and the successor to Lucie, released in 1B, 8B, and 23B sizes (base + Instruct-1.1). Apache-2.0 weights, the redistributable Luciole-Training-Dataset (~4.65T tokens drawn only from openly licensed corpora), the Luciole-Training pipeline on GitHub, and intermediate + final pretraining checkpoints — clearing the full open-source bar. Base models are pretrained on roughly 30% French data (English 53%, French 16%, plus German/Spanish/Italian/Portuguese/Dutch/Arabic and regional languages); the Instruct-1.1 models are SFT + DPO aligned on math, science, coding, chat, RAG, and translation. Built by LINAGORA and the OpenLLM-France consortium, funded by BPI France through the France 2030 program.

Luciole family (1B / 8B / 23B, base + Instruct-1.1) from LINAGORA / OpenLLM-France, successor to Lucie-7B. Apache-2.0 weights; open Luciole-Training-Dataset (~4.65T tokens, open-licensed corpora only — cleaner terms than Lucie's cc-by-nc-sa corpus); Luciole-Training pipeline on GitHub (GPL-3.0); intermediate + final checkpoints published. Funded by BPI France / France 2030. Collection: https://huggingface.co/collections/OpenLLM-France/luciole-llm. Verified via primary model/dataset cards, August 2026 (HF direct fetch blocked by egress proxy; figures cross-checked via search snippets of the live cards).

Openness

5 high confidence
5.0
weights
open(Apache-2.0, safetensors on HF for 1B/8B/23B base + Instruct-1.1)
data
open(Luciole-Training-Dataset on HF, ~4.65T tokens, "only corpora under open licenses")
code
open(Luciole-Training pretraining + data-processing pipeline on GitHub, GPL-3.0)
checkpoints
open(intermediate + final pretraining checkpoints released)
license
Apache-2.0(OSI)

Full open-source tier alongside Lucie / OLMo / Apertus: open weights, open training data, open training code, and intermediate checkpoints. Notably a cleaner full-openness exemplar than its Lucie predecessor — the Luciole-Training-Dataset is described as containing only openly licensed corpora, where Lucie-Training-Dataset carried a non-commercial cc-by-nc-sa license. The weights are Apache-2.0 (OSI) across every distributed size, so the data + code + OSI-license rule lands at 5.

Adoption

1 medium confidence
1.0

Niche sovereign / French open-LLM audience, and newer and smaller-reach than Lucie-7B (level 2). Summed monthly HF downloads across the shipped sizes sit in the low thousands (23B ~640, 8B ~944, 1B ~842 as of Aug 2026), well under the 10K level-2 threshold. Adoption may be recomputed from the linked artifacts by the warehouse.

Capability

2 low confidence
2.0

Multilingual French–English family up to 23B, prioritizing French cultural representation and open data rights over frontier performance. Larger than Lucie-7B (capability 2) but the same sovereign-model tier — mid-scale and well below the 2026 open-weight frontier (trillion-param MoEs). Positioned like Alia-40b (capability 2); kept conservative pending published benchmarks. Value is openness + FR/EU language coverage.

Unchanged since 2026-08-07 (last edited, not re-checked)