PersonaHub
Tencent AI LabPersona-driven synthetic data resource from Tencent AI Lab, backed by the paper "Scaling Synthetic Data Creation with 1,000,000,000 Personas". It releases a large set of distilled personas - a 200,000-persona preview alongside roughly 370 million elite personas - plus seed synthetic samples including 50,000 math problems, 50,000 logical-reasoning problems and 50,000 instructions.
Tulu 3's persona-math set expands the same methodology. Verified 2026-08-13 via the proj-persona/PersonaHub dataset card on Hugging Face.
Openness
2 high confidence- license
- cc-by-nc-sa-4.0(non-commercial)
- card
- present
- ungated
- yes
- scope
- ~370M personas+seed samples
Publicly downloadable and ungated, and licensed CC-BY-NC-SA-4.0, which permits redistribution for non-commercial research and forbids commercial use. Openness turns on the commercial-use test rather than on redistributability alone, and the Open Definition agrees: data is not open unless it may be reused commercially. That leaves this corpus between open and gated, which is what restricted records.
- https://huggingface.co/datasets/proj-persona/PersonaHub recorded 2026-08-13
CC-BY-NC-SA-4.0 license, ungated card, persona counts and paper reference
Adoption
2 high confidence9,346 downloads in the trailing 30 days for proj-persona/PersonaHub, which bands at 1K-10K, level 2 on the dataset adoption scale.
- https://huggingface.co/api/datasets/proj-persona/PersonaHub recorded 2026-08-13
9,346 downloads in the trailing 30 days for proj-persona/PersonaHub
Capability
4 medium confidenceThe paper's headline training result is the basis: Qwen2-7B fine-tuned on ~1.07M persona-generated math problems reaches 64.9% on MATH, rivaling GPT-4-turbo-preview at the time. Persona Hub is 1 billion personas curated from web data and used to synthesize mathematical and logical reasoning problems at scale, and the lineage carries downstream: Tulu 3's SFT persona-math set conditions on ~250K of these personas and credits the persona methodology for its 149,960 examples, so the corpus is an ingredient in a mixture already scored 4 here. It sits at 4 rather than 5 because the 5s in this category - SmolTalk, FineWeb-Edu, DCLM-Baseline - both lead head-to-head ablations and are the default others adopt wholesale, whereas PersonaHub is a generation method plus seed samples whose adoption runs through derivative sets. That is the same shape as tulu-3-sft-mixture and fineweb, both also 4.
- https://arxiv.org/abs/2406.20094 recorded 2026-08-13
Abstract - Persona Hub is 1 billion personas curated from web data, showcased on synthesised mathematical and logical reasoning problems. The Qwen2-7B MATH result is in the paper body, not on this page
- https://huggingface.co/datasets/allenai/tulu-3-sft-personas-math recorded 2026-08-13
Card - 149,960 synthetic math examples built by expanding the persona methodology of Ge et al. 2024
Verified 2026-08-13