AI Potluck
Model components / Benchmark / eval datasets

TruthfulQA

Unknown

TruthfulQA, 817 questions across 38 categories crafted to elicit imitative falsehoods/misconceptions; generation + multiple_choice configs. Paper arXiv:2109.07958 (2021). Now relatively dated and saturated by frontier models but still a widely-cited truthfulness benchmark. HF dataset truthfulqa/truthful_qa verified live June 2026.

TruthfulQA, 817 questions across 38 categories crafted to elicit imitative falsehoods/misconceptions; generation + multiple_choice configs. Paper arXiv:2109.07958 (2021). Now relatively dated and saturated by frontier models but still a widely-cited truthfulness benchmark. HF dataset truthfulqa/truthful_qa verified live June 2026.

Openness

5 high confidence
5.0
license
Apache-2.0(OSI/open,redistributable)
access
public(not gated)
datasheet
present(card+creation methodology+citation)
splits
public(817 Qs, two configs, no hidden held-out)

Fully open: Apache-2.0, ungated, redistributable, complete card with creation methodology.

Adoption

4 medium confidence
4.0

~97,282 HF downloads in the trailing month (primary); long-standing widely-cited truthfulness benchmark. Medium confidence on the level because the benchmark is older (2021) and increasingly saturated, so download volume may overstate its current role in frontier evaluation vs GPQA/MMLU-Pro.

Capability

not assessed

Dataset, not 'capable'; openness and adoption carry this category per recipe.

Unchanged since 2026-06-09 (last edited, not re-checked)