AI Potluck
Model components / Training & synthetic datasets

SmolTalk

Hugging Face

Hugging Face's supervised fine-tuning mixture behind the SmolLM instruct models, combining new synthetic sets such as Smol-Magpie-Ultra with established public datasets. Marin 8B Instruct also trains on it. The record covers both SmolTalk and its successor SmolTalk2.

Component datasets keep the terms of their own sources, which the card documents. Verified 2026-08-13 via the HuggingFaceTB/smoltalk dataset card on Hugging Face.

Openness

4 high confidence
4.0
access
ungated
license
apache-2.0(new-subsets)+per-component
dataset_card
present

Ungated; newly created subsets are Apache-2.0, incorporated public datasets keep their own licenses.

Adoption

3 high confidence
3.0

30,281 downloads in the trailing 30 days for HuggingFaceTB/smoltalk, which bands at 10K-100K, level 3 on the dataset adoption scale.

Capability

5 high confidence
5.0

Leads head-to-head SFT-mixture ablations (beats OpenHermes-2.5, UltraChat, MagPie-Pro) and is the current default across SmolLM2/3 and Marin 8B Instruct.

Verified 2026-08-13