May 2026Unreviewed
SoK: Robustness in Large Language Models against Jailbreak Attacks
Feiyue Xu, Hongsheng Hu, Chaoxiang He, Sheng Hang, Hanqing Hu, Xiuming Liu, Yubo Zhao, Zhengyan Zhou, Bin Benjamin Zhu, Shi-Feng Sun, Dawu Gu, Shuo Wang
Abstract
Large Language Models (LLMs) have achieved remarkable success but remain highly susceptible to jailbreak attacks, in which adversarial prompts coerce models into generating harmful, unethical, or policy-violating outputs. Such attacks pose real-world risks, eroding safety, trust, and regulatory compliance in high-stakes applications. Although a variety of attack and defense methods have been proposed, existing evaluation practices are inadequate, often relying on narrow metrics like attack succe
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0054LLM Jailbreak
Suggested from the entry's categories.
Cite
@misc{xu2026sok,
title = {{SoK: Robustness in Large Language Models against Jailbreak Attacks}},
author = {Feiyue Xu and Hongsheng Hu and Chaoxiang He and Sheng Hang and Hanqing Hu and Xiuming Liu and Yubo Zhao and Zhengyan Zhou and Bin Benjamin Zhu and Shi-Feng Sun and Dawu Gu and Shuo Wang},
year = {2026},
month = may,
eprint = {2605.05058},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2605.05058}
}