AI Potluck
Model components / Training & synthetic datasets

OpenThoughts-114k

Open Thoughts

Synthetic reasoning dataset of roughly 114,000 examples spanning math, science, code and puzzles, built by curating problem sets, generating reasoning traces with DeepSeek-R1 and then verifying correctness. It was created by the Open Thoughts team and used to train the OpenThinker 7B and 32B models.

OpenThoughts3-1.2M is the recommended successor. Verified 2026-08-13 via the open-thoughts/OpenThoughts-114k dataset card on Hugging Face.

Openness

5 high confidence
5.0
license
apache-2.0
access
ungated
dataset_card
present

Apache-2.0 licensed, ungated, with dataset card, code repo, and paper.

Adoption

4 high confidence
4.0

The Hugging Face datasets API reports 107,269 downloads in the trailing 30 days for open-thoughts/OpenThoughts-114k, ungated and with 901 likes, which bands at 100K-1M, level 4 on the dataset adoption scale.

Capability

3 high confidence
3.0

Strong early-2025 reasoning results (OpenThinker-32B led open-data MATH500/GPQA-D) but superseded by OpenThoughts3-1.2M. The OpenThinker-32B card reports MATH500 90.6 and GPQA-Diamond 61.6 against LIMO-32B, s1-32B and s1.1-32B, and the OpenThoughts blog now recommends OpenThoughts3-1.2M.

Verified 2026-08-13