September 2026Unreviewed
EvalDetectBench: A Benchmark for Measuring Evaluation Awareness in Frontier Language Models
Xinning Li, Kemunto Ochwang'i, Aryasomayajula Ram Bharadwaj, Alexandra Souly, Robert Kirk
Abstract
Frontier large language models can often recognize when they are being evaluated, a capability known as evaluation awareness. If models behave differently in evaluations than in deployment, this undermines the validity of evaluation results, which are a crucial component of current AI safety frameworks. We introduce EvalDetectBench, an open pipeline and benchmark for measuring evaluation awareness that works with any Inspect-compatible evaluation, allowing practitioners to test against current a
Categories
Cite
@misc{li2026evaldetectbench,
title = {{EvalDetectBench: A Benchmark for Measuring Evaluation Awareness in Frontier Language Models}},
author = {Xinning Li and Kemunto Ochwang'i and Aryasomayajula Ram Bharadwaj and Alexandra Souly and Robert Kirk},
year = {2026},
month = sep,
eprint = {2609.01611},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2609.01611}
}