September 2026Unreviewed
Keep Evaluation Fair: Detecting Data Leakage in Code Generation Benchmarks via Membership Inference Attacks
Dong-Dong Zhao, Jian Chen, Guan-Cheng Lin, Jianwen Xiang, J. Keung, Xiao Yu
Abstract
Code generation benchmarks are widely used to evaluate Large Language Models (LLMs), but benchmark data leakage into training sets can inflate performance and undermine evaluation validity. DetectLeak, a method specifically designed for code generation benchmark leakage detection, relies on perplexity scores to identify likely leaked samples. However, perplexity mainly reflects general familiarity with code patterns and may perform poorly on complex or rare samples. It also overlooks other usefu
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM02Sensitive Information Disclosure
MITRE ATLAS
- AML.T0024.000Infer Training Data Membership
Suggested from the entry's categories.
Cite
@misc{zhao2026keep,
title = {{Keep Evaluation Fair: Detecting Data Leakage in Code Generation Benchmarks via Membership Inference Attacks}},
author = {Dong-Dong Zhao and Jian Chen and Guan-Cheng Lin and Jianwen Xiang and J. Keung and Xiao Yu},
year = {2026},
month = sep,
eprint = {2609.09865},
archivePrefix = {arXiv},
url = {https://www.semanticscholar.org/paper/b35edac8aca5fbcd8a42b330512e3a64fd1214cc}
}