UltraChat 200k
Hugging FaceUltraChat 200k is Hugging Face H4's filtered 200k-dialogue slice of the ChatGPT-distilled UltraChat corpus, released under MIT for supervised fine-tuning. It anchored the Zephyr recipe and became a default open SFT set for chat alignment, also used by TinyLlama-Chat.
Openness
5 high confidence- license
- MIT
- gated
- false
- dataset_card
- present
MIT-licensed, ungated, full dataset card.
- https://huggingface.co/datasets/HuggingFaceH4/ultrachat_200k recorded 2026-07-04
MIT license, ungated access, dataset card
Adoption
3 high confidence57,823 monthly downloads on Hugging Face, graded on the training-corpus bands.
- https://huggingface.co/api/datasets/HuggingFaceH4/ultrachat_200k recorded 2026-07-04
downloads field = 57,823 (30-day)
Capability
3 high confidenceClean SFT before/after in the Zephyr paper (MT-Bench ~6.64), but dropped from 2024-25 default SFT mixes (absent from SmolTalk and Tulu 3).
- https://ar5iv.labs.arxiv.org/html/2310.16944 recorded 2026-07-04
Zephyr: dSFT on UltraChat reaches MT-Bench 6.64
- https://huggingface.co/datasets/allenai/tulu-3-sft-mixture recorded 2026-07-04
UltraChat-200k is not used as an SFT source in Tulu 3
Unchanged since 2026-07-04 (last edited, not re-checked)