UltraChat 200k
Hugging FaceHugging Face H4's filtered 200,000-dialogue slice of the ChatGPT-distilled UltraChat corpus, prepared for supervised fine-tuning of chat models. It anchored the Zephyr recipe and is also used by TinyLlama-Chat, and remains a common starting point for open chat alignment work.
Tulu 3 does not use it as an SFT source. Verified 2026-08-13 via the HuggingFaceH4/ultrachat_200k dataset card on Hugging Face.
Openness
5 high confidence- license
- MIT
- gated
- false
- dataset_card
- present
MIT-licensed, ungated, full dataset card.
- https://huggingface.co/datasets/HuggingFaceH4/ultrachat_200k recorded 2026-08-13
MIT license, ungated access, dataset card
Adoption
3 high confidence75,198 downloads in the trailing 30 days for HuggingFaceH4/ultrachat_200k, which bands at 10K-100K, level 3 on the dataset adoption scale.
- https://huggingface.co/api/datasets/HuggingFaceH4/ultrachat_200k recorded 2026-08-13
75,198 downloads in the trailing 30 days for HuggingFaceH4/ultrachat_200k
Capability
3 high confidenceClean SFT before/after in the Zephyr paper (MT-Bench ~6.64), but dropped from 2024-25 default SFT mixes (absent from SmolTalk and Tulu 3).
- https://ar5iv.labs.arxiv.org/html/2310.16944 recorded 2026-08-13
Zephyr: dSFT on UltraChat reaches MT-Bench 6.64
- https://huggingface.co/datasets/allenai/tulu-3-sft-mixture recorded 2026-08-13
UltraChat-200k is not used as an SFT source in Tulu 3
Verified 2026-08-13