SmolTalk
Hugging FaceHugging Face's supervised fine-tuning mixture behind the SmolLM instruct models, combining new synthetic sets such as Smol-Magpie-Ultra with established public datasets. Marin 8B Instruct also trains on it. The record covers both SmolTalk and its successor SmolTalk2.
Component datasets keep the terms of their own sources, which the card documents. Verified 2026-08-13 via the HuggingFaceTB/smoltalk dataset card on Hugging Face.
Openness
4 high confidence- access
- ungated
- license
- apache-2.0(new-subsets)+per-component
- dataset_card
- present
Ungated; newly created subsets are Apache-2.0, incorporated public datasets keep their own licenses.
- https://huggingface.co/datasets/HuggingFaceTB/smoltalk recorded 2026-08-13
ungated, mixed licensing documented on card
Adoption
3 high confidence30,281 downloads in the trailing 30 days for HuggingFaceTB/smoltalk, which bands at 10K-100K, level 3 on the dataset adoption scale.
- https://huggingface.co/api/datasets/HuggingFaceTB/smoltalk recorded 2026-08-13
30,281 downloads in the trailing 30 days for HuggingFaceTB/smoltalk
Capability
5 high confidenceLeads head-to-head SFT-mixture ablations (beats OpenHermes-2.5, UltraChat, MagPie-Pro) and is the current default across SmolLM2/3 and Marin 8B Instruct.
- https://arxiv.org/html/2502.02737v1 recorded 2026-08-13
SmolLM2 paper: SmolTalk surpasses OpenHermes-2.5, UltraChat, and MagPie-Pro
- https://huggingface.co/marin-community/marin-8b-instruct recorded 2026-08-13
Marin 8B Instruct lists SmolTalk in its SFT mix
Verified 2026-08-13