OpenHermes 2.5
TekniumOpenHermes 2.5 is a large synthetic instruction-following dataset of ~1M conversational examples compiled and generated primarily from GPT-4 outputs across math, code, roleplay, and general-knowledge tasks. Created by Teknium as an expansion of OpenHermes 1, it was used to train the popular OpenHermes 2.5 fine-tunes and is widely used as a base SFT mix for open instruction-tuned models.
Verified live 2026-06-22 via primary sources. Publicly downloadable and ungated with a full card, though an explicit license string is not stated on the card.
Openness
5 medium confidence- license
- not-clearly-stated-on-card(commonly cited as permissive)
- card
- present
- ungated
- yes
- format
- JSON/Parquet
- rows
- ~1M
Publicly downloadable and ungated with a full card, though an explicit license string is not stated on the card.
- https://huggingface.co/datasets/teknium/OpenHermes-2.5 recorded 2026-06-22
Public ungated dataset card, ~1M rows, synthetic/GPT-4 tags
Adoption
3 high confidence14,982 monthly downloads on Hugging Face, graded on the training-corpus bands.
- https://huggingface.co/api/datasets/teknium/OpenHermes-2.5 recorded 2026-07-04
downloads field = 14,982 (30-day)
Capability
3 high confidenceAttributed gains inside SmolTalk (MMLU/BBH/WinoGrande) but no standalone ablation; superseded by Hermes 3 and survives as a 100k SmolTalk slice.
- https://arxiv.org/html/2502.02737v1 recorded 2026-07-04
SmolLM2 paper credits a 100k OpenHermes-2.5 subset with MMLU/WinoGrande/BBH gains
Unchanged since 2026-07-04 (last edited, not re-checked)