Skip to content
Search
paperSeptember 2026Unreviewed

Contamination Inflates Scores but Rarely Reorders Large Language Model Leaderboards

Xingyao Xiao, Yihong Cheng

Abstract

Benchmark contamination, the leakage of test items into training data, is widely described as a threat to the reliability of large language model (LLM) leaderboards. We argue that this concern conflates two distinct questions: whether contamination inflates absolute scores, and whether it reorders the ranking of models. We recast contamination as a violation of anchor-item invariance and measure it through the differential functioning of original versus semantically equivalent paraphrased items,

Categories

Cite

@misc{xiao2026contamination,
  title = {{Contamination Inflates Scores but Rarely Reorders Large Language Model Leaderboards}},
  author = {Xingyao Xiao and Yihong Cheng},
  year = {2026},
  month = sep,
  eprint = {2609.02899},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2609.02899}
}