2024ReviewedOpen access
Pandora's White-Box: Precise Training Data Detection and Extraction in Large Language Models
Jeffrey G. Wang, Jason Wang, Marvin Li, Seth Neel
arXiv preprint
Abstract
Develops precise methods for detecting and extracting training data from LLMs when white-box access is available, with implications for copyright and privacy.
Categories
#training-data-detection#white-box#copyright
Framework mappings
OWASP Top 10 for LLM Applications
- LLM02Sensitive Information Disclosure
MITRE ATLAS
- AML.T0024Exfiltration via AI Inference API
Cite
@misc{wang2024pandoras,
title = {{Pandora's White-Box: Precise Training Data Detection and Extraction in Large Language Models}},
author = {Jeffrey G. Wang and Jason Wang and Marvin Li and Seth Neel},
year = {2024},
eprint = {2402.17012},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2402.17012}
}