July 2026Unreviewed
The Poisoned Chalice of LLM Evaluation Report
Jonathan Katzy, Ali Al-Kaswan, Razvan Mihai Popescu, Zhou Yang
Abstract
Large language models are increasingly used to evaluate and support software engineering tasks, yet the validity of these evaluations is often undermined by uncertainty about whether benchmark instances were seen during pretraining. This can lead to data contamination, which may inflate performance and result in misleading conclusions about model capability. Despite this, the training corpora of many modern models are only partially disclosed, making direct decontamination infeasible. This creat
Categories
Cite
@misc{katzy2026poisoned,
title = {{The Poisoned Chalice of LLM Evaluation Report}},
author = {Jonathan Katzy and Ali Al-Kaswan and Razvan Mihai Popescu and Zhou Yang},
year = {2026},
month = jul,
eprint = {2607.07481},
archivePrefix = {arXiv},
doi = {10.1145/3803437.3807733},
url = {https://arxiv.org/abs/2607.07481}
}