AI Potluck
Model components / Training & synthetic datasets

PersonaHub

Tencent AI Lab

PersonaHub is a persona-driven synthetic data resource from Tencent AI Lab, backed by the paper 'Scaling Synthetic Data Creation with 1,000,000,000 Personas' (arXiv 2406.20094). The repo releases a large set of distilled personas (200K preview plus ~370M elite personas) plus seed synthetic samples such as 50K math problems, 50K logical-reasoning problems, and 50K instructions. It is a method and dataset for generating diverse, scalable synthetic instruction and reasoning data.

Verified live 2026-06-22 via primary sources. Publicly downloadable and ungated under CC-BY-NC-SA-4.0, which permits redistribution for non-commercial research.

Openness

2 high confidence
2.0
license
cc-by-nc-sa-4.0(non-commercial)
card
present
ungated
yes
scope
~370M personas+seed samples

Publicly downloadable and ungated, and licensed CC-BY-NC-SA-4.0, which permits redistribution for non-commercial research and forbids commercial use. Moved from 5/open on 2026-08-01 by the universal license scale. The map had answered the non-commercial question two ways: `command-r` is 2/restricted under CC-BY-NC while this was 5/open under CC-BY-NC-SA, on the reasoning that a corpus's openness question is redistributability rather than commercial reuse. The scale settles it on the commercial-use test, and the Open Definition agrees - data is not open unless it may be reused commercially. `restricted` is new to the dataset class vocabulary for this: it was the only product type with no word between `open` and `gated`.

Adoption

2 high confidence
2.0

Publicly downloadable and ungated, and licensed CC-BY-NC-SA-4.0, which permits redistribution for non-commercial research and forbids commercial use. Moved from 5/open on 2026-08-01 by the universal license scale. The map had answered the non-commercial question two ways: `command-r` is 2/restricted under CC-BY-NC while this was 5/open under CC-BY-NC-SA, on the reasoning that a corpus's openness question is redistributability rather than commercial reuse. The scale settles it on the commercial-use test, and the Open Definition agrees - data is not open unless it may be reused commercially. `restricted` is new to the dataset class vocabulary for this: it was the only product type with no word between `open` and `gated`.

Capability

2 high confidence
2.0

Publicly downloadable and ungated, and licensed CC-BY-NC-SA-4.0, which permits redistribution for non-commercial research and forbids commercial use. Moved from 5/open on 2026-08-01 by the universal license scale. The map had answered the non-commercial question two ways: `command-r` is 2/restricted under CC-BY-NC while this was 5/open under CC-BY-NC-SA, on the reasoning that a corpus's openness question is redistributability rather than commercial reuse. The scale settles it on the commercial-use test, and the Open Definition agrees - data is not open unless it may be reused commercially. `restricted` is new to the dataset class vocabulary for this: it was the only product type with no word between `open` and `gated`.

Unchanged since 2026-08-01 (last edited, not re-checked)