Skip to content
Search
paperOctober 2026UnreviewedOpen access

LLMs Leak Training Data Beyond Verbatim Memorization: Extraction via Membership Decoding

Zi-Tai Chen, Reza Shokri

Proceedings on Privacy Enhancing Technologies

Abstract

Extracting training data from large language models (LLMs) is a serious privacy breach that exposes (potentially private) data without data owners' consent. Existing extractions follow the generation-then-audit paradigm, where the greedy decoding method in generation limits the extraction scope and only verbatim memorized data is under audits. A majority of partially memorized member data (around 90%) remains unexplored, of which LLMs could memorize almost all tokens but fail to rank the trainin

Categories

Framework mappings

OWASP Top 10 for LLM Applications
  • LLM02Sensitive Information Disclosure
MITRE ATLAS
  • AML.T0024.000Infer Training Data Membership

Suggested from the entry's categories.

Cite

@inproceedings{chen2026llms,
  title = {{LLMs Leak Training Data Beyond Verbatim Memorization: Extraction via Membership Decoding}},
  author = {Zi-Tai Chen and Reza Shokri},
  year = {2026},
  month = oct,
  booktitle = {Proceedings on Privacy Enhancing Technologies},
  doi = {10.56553/popets-2026-0139},
  url = {https://www.semanticscholar.org/paper/c52e919714f66688c0620093ff6b9898bd00edc9}
}