Skip to content
Search
paperSeptember 2026Unreviewed

Keep Evaluation Fair: Detecting Data Leakage in Code Generation Benchmarks via Membership Inference Attacks

Dong-Dong Zhao, Jian Chen, Guan-Cheng Lin, Jianwen Xiang, J. Keung, Xiao Yu

Abstract

Code generation benchmarks are widely used to evaluate Large Language Models (LLMs), but benchmark data leakage into training sets can inflate performance and undermine evaluation validity. DetectLeak, a method specifically designed for code generation benchmark leakage detection, relies on perplexity scores to identify likely leaked samples. However, perplexity mainly reflects general familiarity with code patterns and may perform poorly on complex or rare samples. It also overlooks other usefu

Categories

Framework mappings

OWASP Top 10 for LLM Applications
  • LLM02Sensitive Information Disclosure
MITRE ATLAS
  • AML.T0024.000Infer Training Data Membership

Suggested from the entry's categories.

Cite

@misc{zhao2026keep,
  title = {{Keep Evaluation Fair: Detecting Data Leakage in Code Generation Benchmarks via Membership Inference Attacks}},
  author = {Dong-Dong Zhao and Jian Chen and Guan-Cheng Lin and Jianwen Xiang and J. Keung and Xiao Yu},
  year = {2026},
  month = sep,
  eprint = {2609.09865},
  archivePrefix = {arXiv},
  url = {https://www.semanticscholar.org/paper/b35edac8aca5fbcd8a42b330512e3a64fd1214cc}
}