May 2026Unreviewed
Pretraining Data Exposure in Large Language Models: A Survey of Membership Inference, Data Contamination, and Security Implications
Ziyi Tong, Feifei Sun, Le Minh Nguyen
Abstract
Large Language Models (LLMs) have become the predominant paradigm in NLP, advancing both research and industry. As model sizes and pretraining data grow, concerns about Pretraining Data Exposure (PDE) increase due to the scale and opacity of training datasets. PDE refers to determining whether specific data appeared in an LLM's pretraining corpus. It is critical for ensuring evaluation integrity and protecting privacy, intersecting two key areas: data contamination and membership inference. Thou
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM02Sensitive Information Disclosure
MITRE ATLAS
- AML.T0024.000Infer Training Data Membership
Suggested from the entry's categories.
Cite
@misc{tong2026pretraining,
title = {{Pretraining Data Exposure in Large Language Models: A Survey of Membership Inference, Data Contamination, and Security Implications}},
author = {Ziyi Tong and Feifei Sun and Le Minh Nguyen},
year = {2026},
month = may,
eprint = {2605.26133},
archivePrefix = {arXiv},
doi = {10.1007/978-3-031-97144-0_14},
url = {https://arxiv.org/abs/2605.26133}
}