AI Potluck
Model components / Training & synthetic datasets

OpenHermes 2.5

Teknium

Synthetic instruction-following dataset of roughly a million conversational examples, compiled and generated primarily from GPT-4 outputs across math, code, roleplay and general-knowledge tasks. Created by Teknium as an expansion of OpenHermes 1, it trained the OpenHermes 2.5 fine-tunes and is used as a base supervised-tuning mix by open instruction models.

The dataset card declares no license and no license file ships with the data, which the openness axis resolves. Verified 2026-08-13 via the teknium/OpenHermes-2.5 dataset card and its raw README.

Openness

3 medium confidence
3.0
license
not-clearly-stated-on-card(no license field in the card YAML, no license section in the card body, no LICENSE file in the repository and no license tag on the Hub API)
card
present
ungated
yes
format
JSON/Parquet
rows
~1M

Publicly downloadable and ungated with a full card, but with no license anywhere the publisher controls: the card front matter carries language, pretty_name and tags and no license key, the 126-line body runs from the source list through the format description to the citation without a terms section, the repository holds only .gitattributes, README.md and openhermes2_5.json so no LICENSE file ships with the data, and the Hub API returns no license either. Files that download freely without a grant of rights are not open, and not closed either, which is what the gated class records here.

Adoption

3 high confidence
3.0

17,420 downloads in the trailing 30 days for teknium/OpenHermes-2.5, which bands at 10K-100K, level 3 on the dataset adoption scale.

Capability

3 high confidence
3.0

Attributed gains inside SmolTalk (MMLU/BBH/WinoGrande) but no standalone ablation; superseded by Hermes 3 and survives as a 100k SmolTalk slice.

Verified 2026-08-13