April 2023ReviewedOpen access
Multi-step Jailbreaking Privacy Attacks on ChatGPT
Haoran Li, Dadi Guo, Wei Fan, Mingshi Xu, Jie Huang, Fanpu Meng, Yangqiu Song
EMNLP 2023 Findings
Abstract
Demonstrates multi-step jailbreaking attacks to extract personal information from ChatGPT, showing how sequential prompting can bypass safety measures.
Categories
#multi-step#privacy#PII-extraction
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
- LLM02Sensitive Information Disclosure
MITRE ATLAS
- AML.T0054LLM Jailbreak
- AML.T0057LLM Data Leakage
Cite
@inproceedings{li2023multistep,
title = {{Multi-step Jailbreaking Privacy Attacks on ChatGPT}},
author = {Haoran Li and Dadi Guo and Wei Fan and Mingshi Xu and Jie Huang and Fanpu Meng and Yangqiu Song},
year = {2023},
month = apr,
booktitle = {EMNLP 2023 Findings},
eprint = {2304.05197},
archivePrefix = {arXiv},
doi = {10.18653/v1/2023.findings-emnlp.272},
url = {https://arxiv.org/abs/2304.05197}
}