Skip to content
Search
paper2024ReviewedOpen access

Pandora's White-Box: Precise Training Data Detection and Extraction in Large Language Models

Jeffrey G. Wang, Jason Wang, Marvin Li, Seth Neel

arXiv preprint

Abstract

Develops precise methods for detecting and extracting training data from LLMs when white-box access is available, with implications for copyright and privacy.

Categories

#training-data-detection#white-box#copyright

Framework mappings

OWASP Top 10 for LLM Applications
  • LLM02Sensitive Information Disclosure
MITRE ATLAS
  • AML.T0024Exfiltration via AI Inference API

Cite

@misc{wang2024pandoras,
  title = {{Pandora's White-Box: Precise Training Data Detection and Extraction in Large Language Models}},
  author = {Jeffrey G. Wang and Jason Wang and Marvin Li and Seth Neel},
  year = {2024},
  eprint = {2402.17012},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2402.17012}
}