TruthfulQA: How a 800-Question Benchmark Tests AI for Imitative Falsehoods
TruthfulQA is a benchmark of roughly 800 adversarially written questions, published in 2021 by Lin, Hilton, and Evans, designed to test whether AI language models repeat common misconceptions present in their training data. The benchmark specifically measures 'imitative falsehood' — when a model reproduces a false but widely held belief — rather than knowledge gaps, confabulation, or reasoning errors. Questions were selected precisely because contemporary models answered them incorrectly, making the set a targeted probe of one distinct failure mode. The benchmark can be run in three incomparable modes — free-form generation, single-answer multiple choice (MC1), and multi-true multiple choice (MC2) — and papers often report only one without specifying which. Notably, more capable models can score worse on TruthfulQA, as they may more faithfully replicate falsehoods prevalent in human-generated training data.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in