Skip to content
Search
paperApril 2023ReviewedOpen access

Multi-step Jailbreaking Privacy Attacks on ChatGPT

Haoran Li, Dadi Guo, Wei Fan, Mingshi Xu, Jie Huang, Fanpu Meng, Yangqiu Song

EMNLP 2023 Findings

Abstract

Demonstrates multi-step jailbreaking attacks to extract personal information from ChatGPT, showing how sequential prompting can bypass safety measures.

Categories

#multi-step#privacy#PII-extraction

Framework mappings

OWASP Top 10 for LLM Applications
  • LLM01Prompt Injection
  • LLM02Sensitive Information Disclosure
MITRE ATLAS
  • AML.T0054LLM Jailbreak
  • AML.T0057LLM Data Leakage

Cite

@inproceedings{li2023multistep,
  title = {{Multi-step Jailbreaking Privacy Attacks on ChatGPT}},
  author = {Haoran Li and Dadi Guo and Wei Fan and Mingshi Xu and Jie Huang and Fanpu Meng and Yangqiu Song},
  year = {2023},
  month = apr,
  booktitle = {EMNLP 2023 Findings},
  eprint = {2304.05197},
  archivePrefix = {arXiv},
  doi = {10.18653/v1/2023.findings-emnlp.272},
  url = {https://arxiv.org/abs/2304.05197}
}