AI Potluck
Model components / Training & synthetic datasets

UltraChat 200k

Hugging Face

Hugging Face H4's filtered 200,000-dialogue slice of the ChatGPT-distilled UltraChat corpus, prepared for supervised fine-tuning of chat models. It anchored the Zephyr recipe and is also used by TinyLlama-Chat, and remains a common starting point for open chat alignment work.

Tulu 3 does not use it as an SFT source. Verified 2026-08-13 via the HuggingFaceH4/ultrachat_200k dataset card on Hugging Face.

Openness

5 high confidence
5.0
license
MIT
gated
false
dataset_card
present

MIT-licensed, ungated, full dataset card.

Adoption

3 high confidence
3.0

75,198 downloads in the trailing 30 days for HuggingFaceH4/ultrachat_200k, which bands at 10K-100K, level 3 on the dataset adoption scale.

Capability

3 high confidence
3.0

Clean SFT before/after in the Zephyr paper (MT-Bench ~6.64), but dropped from 2024-25 default SFT mixes (absent from SmolTalk and Tulu 3).

Verified 2026-08-13