AI Potluck
Model components / Training & synthetic datasets

UltraChat 200k

Hugging Face

UltraChat 200k is Hugging Face H4's filtered 200k-dialogue slice of the ChatGPT-distilled UltraChat corpus, released under MIT for supervised fine-tuning. It anchored the Zephyr recipe and became a default open SFT set for chat alignment, also used by TinyLlama-Chat.

Openness

5 high confidence
5.0
license
MIT
gated
false
dataset_card
present

MIT-licensed, ungated, full dataset card.

Adoption

3 high confidence
3.0

57,823 monthly downloads on Hugging Face, graded on the training-corpus bands.

Capability

3 high confidence
3.0

Clean SFT before/after in the Zephyr paper (MT-Bench ~6.64), but dropped from 2024-25 default SFT mixes (absent from SmolTalk and Tulu 3).

Unchanged since 2026-07-04 (last edited, not re-checked)