Skip to content
Search
paper2024ReviewedOpen access

AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models

Xiaogeng Liu, Nan Xu, Muhao Chen, Chaowei Xiao

ICLR 2024

Abstract

Proposes AutoDAN, a method for automatically generating stealthy jailbreak prompts that are semantically meaningful and can bypass perplexity-based defenses.

Categories

#automated-jailbreak#stealthy#genetic-algorithm

Framework mappings

MITRE ATLAS
  • AML.T0054LLM Jailbreak

Cite

@inproceedings{liu2024autodan,
  title = {{AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models}},
  author = {Xiaogeng Liu and Nan Xu and Muhao Chen and Chaowei Xiao},
  year = {2024},
  booktitle = {ICLR 2024},
  eprint = {2310.04451},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2310.04451}
}