TruthfulQA
UnknownTruthfulQA, 817 questions across 38 categories crafted to elicit imitative falsehoods/misconceptions; generation + multiple_choice configs. Paper arXiv:2109.07958 (2021). Now relatively dated and saturated by frontier models but still a widely-cited truthfulness benchmark. HF dataset truthfulqa/truthful_qa verified live June 2026.
TruthfulQA, 817 questions across 38 categories crafted to elicit imitative falsehoods/misconceptions; generation + multiple_choice configs. Paper arXiv:2109.07958 (2021). Now relatively dated and saturated by frontier models but still a widely-cited truthfulness benchmark. HF dataset truthfulqa/truthful_qa verified live June 2026.
Openness
5 high confidence- license
- Apache-2.0(OSI/open,redistributable)
- access
- public(not gated)
- datasheet
- present(card+creation methodology+citation)
- splits
- public(817 Qs, two configs, no hidden held-out)
Fully open: Apache-2.0, ungated, redistributable, complete card with creation methodology.
- https://huggingface.co/datasets/truthfulqa/truthful_qa recorded 2026-06-04
Apache-2.0 license, not gated, 817 questions, dataset card present
- https://arxiv.org/abs/2109.07958 recorded 2026-06-04
TruthfulQA paper documenting 817 questions, 38 categories
Adoption
4 medium confidence~97,282 HF downloads in the trailing month (primary); long-standing widely-cited truthfulness benchmark. Medium confidence on the level because the benchmark is older (2021) and increasingly saturated, so download volume may overstate its current role in frontier evaluation vs GPQA/MMLU-Pro.
- https://huggingface.co/datasets/truthfulqa/truthful_qa recorded 2026-06-04
97,282 downloads last month
Capability
not assessedDataset, not 'capable'; openness and adoption carry this category per recipe.
Unchanged since 2026-06-09 (last edited, not re-checked)