Skip to content
Search
paper2025ReviewedOpen access

Jailbreaking Black Box Large Language Models in Twenty Queries

Patrick Chao, Alexander Robey, Edgar Dobriban, Hamed Hassani, George J. Pappas, Eric Wong

IEEE SaTML 2025

Abstract

Uses an attacker LLM to automatically generate jailbreak prompts through iterative refinement achieving high success with only black-box access.

Categories

#automated#iterative#black-box

Framework mappings

MITRE ATLAS
  • AML.T0054LLM Jailbreak

Cite

@article{chao2025jailbreaking,
  title = {{Jailbreaking Black Box Large Language Models in Twenty Queries}},
  author = {Patrick Chao and Alexander Robey and Edgar Dobriban and Hamed Hassani and George J. Pappas and Eric Wong},
  year = {2025},
  journal = {IEEE SaTML 2025},
  eprint = {2310.08419},
  archivePrefix = {arXiv},
  doi = {10.1109/SaTML64287.2025.00010},
  url = {https://arxiv.org/abs/2310.08419}
}