2024ReviewedOpen access
Tree of Attacks: Jailbreaking Black-Box LLMs with Auto-Generated Subtrees
Anay Mehrotra, Manolis Zampetakis, Paul Kassianik, Blaine Nelson, Hyrum Anderson, Yaron Singer, Amin Karbasi
NeurIPS 2024
Abstract
Introduces TAP using an LLM to iteratively refine jailbreak prompts against black-box target models with high success rates.
Categories
#black-box#automated#tree-search
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0054LLM Jailbreak
Cite
@inproceedings{mehrotra2024tree,
title = {{Tree of Attacks: Jailbreaking Black-Box LLMs with Auto-Generated Subtrees}},
author = {Anay Mehrotra and Manolis Zampetakis and Paul Kassianik and Blaine Nelson and Hyrum Anderson and Yaron Singer and Amin Karbasi},
year = {2024},
booktitle = {NeurIPS 2024},
eprint = {2312.02119},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2312.02119}
}