Skip to content
Search
paperJuly 2026Unreviewed

The Poisoned Chalice of LLM Evaluation Report

Jonathan Katzy, Ali Al-Kaswan, Razvan Mihai Popescu, Zhou Yang

Abstract

Large language models are increasingly used to evaluate and support software engineering tasks, yet the validity of these evaluations is often undermined by uncertainty about whether benchmark instances were seen during pretraining. This can lead to data contamination, which may inflate performance and result in misleading conclusions about model capability. Despite this, the training corpora of many modern models are only partially disclosed, making direct decontamination infeasible. This creat

Categories

Cite

@misc{katzy2026poisoned,
  title = {{The Poisoned Chalice of LLM Evaluation Report}},
  author = {Jonathan Katzy and Ali Al-Kaswan and Razvan Mihai Popescu and Zhou Yang},
  year = {2026},
  month = jul,
  eprint = {2607.07481},
  archivePrefix = {arXiv},
  doi = {10.1145/3803437.3807733},
  url = {https://arxiv.org/abs/2607.07481}
}