← Back to search
paper llmsec-2026-00083

SoK: Robustness in Large Language Models against Jailbreak Attacks

Feiyue Xu, Hongsheng Hu, Chaoxiang He, Sheng Hang, Hanqing Hu, Xiuming Liu, Yubo Zhao, Zhengyan Zhou, Bin Benjamin Zhu, Shi-Feng Sun, Dawu Gu, Shuo Wang

2026-05

Abstract

Large Language Models (LLMs) have achieved remarkable success but remain highly susceptible to jailbreak attacks, in which adversarial prompts coerce models into generating harmful, unethical, or policy-violating outputs. Such attacks pose real-world risks, eroding safety, trust, and regulatory compliance in high-stakes applications. Although a variety of attack and defense methods have been proposed, existing evaluation practices are inadequate, often relying on narrow metrics like attack succe

Categories

Cite This Resource

@article{llmsec202600083,
  title = {SoK: Robustness in Large Language Models against Jailbreak Attacks},
  author = {Feiyue Xu and Hongsheng Hu and Chaoxiang He and Sheng Hang and Hanqing Hu and Xiuming Liu and Yubo Zhao and Zhengyan Zhou and Bin Benjamin Zhu and Shi-Feng Sun and Dawu Gu and Shuo Wang},
  year = {2026},
  url = {https://arxiv.org/abs/2605.05058},
}

Metadata

Added
2026-05-17
Added by
automation
Source
arxiv
arxiv_id
2605.05058