June 2026Unreviewed
JailbreakOPT: Tool-Assisted Iterative Jailbreak Prompt Optimization
Ge Shi, Jun Yin, Donglin Xie, Fangyi Liu, Yucan Li, Menglin Liu
Abstract
Jailbreak attacks expose persistent safety weaknesses in large language models (LLMs), but existing stateless single-turn methods face a trade-off: hand-crafted prompts are expressive but static, while iterative prompt optimization can adapt but often relies on low-level mutations that require many target queries. We propose JailbreakOPT, a tool-assisted framework for improving iterative single-turn jailbreak prompt optimization. JailbreakOPT organizes diverse atomic jailbreak prompts into an at
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0054LLM Jailbreak
Suggested from the entry's categories.
Cite
@misc{shi2026jailbreakopt,
title = {{JailbreakOPT: Tool-Assisted Iterative Jailbreak Prompt Optimization}},
author = {Ge Shi and Jun Yin and Donglin Xie and Fangyi Liu and Yucan Li and Menglin Liu},
year = {2026},
month = jun,
eprint = {2606.11425},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2606.11425}
}