September 2026Unreviewed
Judging by the Cover: Cleaning LLM Truthfulness Benchmarks to Avoid Surface-Level Feature Leakage
Foad Namjoo, Remy Ogasawara, Amirali Abdullah, Cullen Anderson, Narmeen Fatimah Oozeer, Jeff M. Phillips
Abstract
Binary-choice truth benchmarks ask models to choose between a correct and an incorrect answer, but if the two answers differ systematically in surface-level features, models can exceed chance without performing the intended reasoning. We show that this failure mode is detectable and can be exploited by downstream classifiers. In TruthfulQA, a simple six-feature logistic classifier achieves substantial accuracy in separating correct from incorrect answers. We further show that similar surface-lev
Categories
Cite
@misc{namjoo2026judging,
title = {{Judging by the Cover: Cleaning LLM Truthfulness Benchmarks to Avoid Surface-Level Feature Leakage}},
author = {Foad Namjoo and Remy Ogasawara and Amirali Abdullah and Cullen Anderson and Narmeen Fatimah Oozeer and Jeff M. Phillips},
year = {2026},
month = sep,
eprint = {2609.13003},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2609.13003}
}