2024ReviewedOpen access
AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models
Xiaogeng Liu, Nan Xu, Muhao Chen, Chaowei Xiao
ICLR 2024
Abstract
Proposes AutoDAN, a method for automatically generating stealthy jailbreak prompts that are semantically meaningful and can bypass perplexity-based defenses.
Categories
#automated-jailbreak#stealthy#genetic-algorithm
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0054LLM Jailbreak
Cite
@inproceedings{liu2024autodan,
title = {{AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models}},
author = {Xiaogeng Liu and Nan Xu and Muhao Chen and Chaowei Xiao},
year = {2024},
booktitle = {ICLR 2024},
eprint = {2310.04451},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2310.04451}
}