Skip to content
Search
paperSeptember 2026Unreviewed

Judging by the Cover: Cleaning LLM Truthfulness Benchmarks to Avoid Surface-Level Feature Leakage

Foad Namjoo, Remy Ogasawara, Amirali Abdullah, Cullen Anderson, Narmeen Fatimah Oozeer, Jeff M. Phillips

Abstract

Binary-choice truth benchmarks ask models to choose between a correct and an incorrect answer, but if the two answers differ systematically in surface-level features, models can exceed chance without performing the intended reasoning. We show that this failure mode is detectable and can be exploited by downstream classifiers. In TruthfulQA, a simple six-feature logistic classifier achieves substantial accuracy in separating correct from incorrect answers. We further show that similar surface-lev

Categories

Cite

@misc{namjoo2026judging,
  title = {{Judging by the Cover: Cleaning LLM Truthfulness Benchmarks to Avoid Surface-Level Feature Leakage}},
  author = {Foad Namjoo and Remy Ogasawara and Amirali Abdullah and Cullen Anderson and Narmeen Fatimah Oozeer and Jeff M. Phillips},
  year = {2026},
  month = sep,
  eprint = {2609.13003},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2609.13003}
}