AI Potluck
Model components / Benchmark / eval datasets

TruthfulQA

Unknown

817 questions across 38 categories, written to elicit imitative falsehoods - the misconceptions a model picks up from its training text rather than errors of reasoning. It ships generation and multiple-choice configurations and dates from 2021.

Relatively dated and largely saturated by frontier models, though still widely reported. Verified 2026-08-13 via the truthfulqa/truthful_qa dataset card on Hugging Face and the paper abstract.

Openness

5 high confidence
5.0
license
Apache-2.0(OSI/open,redistributable)
access
public(not gated)
datasheet
present(card+creation methodology+citation)
splits
public(817 Qs, two configs, no hidden held-out)

Fully open: Apache-2.0, ungated, redistributable, complete card with creation methodology.

Adoption

4 medium confidence
4.0

109,020 Hugging Face downloads in the trailing 30 days for truthfulqa/truthful_qa, which puts it in the 100K-1M band, level 4 on the dataset adoption scale. That scale tops out at >1M, an order of magnitude below the ones used for software and models, because no dataset in this corpus has ever passed 10M downloads. TruthfulQA is a long-standing truthfulness benchmark, though an older one and increasingly saturated.

Capability

not assessed

A dataset is not 'capable', so this axis is left unscored; openness and adoption carry this category.

Verified 2026-08-12