May 2026Unreviewed
Guaranteed Jailbreaking Defense via Disrupt-and-Rectify Smoothing
Zheng Lin, Zhenxing Niu, Haoxuan Ji, Haichang Gao
Abstract
This paper proposes a guaranteed defense method for large language models (LLMs) to safeguard against jailbreaking attacks. Drawing inspiration from the denoised-smoothing approach in the adversarial defense domain, we propose a novel smoothing-based defense method, termed Disrupt-and-Rectify Smoothing (DR-Smoothing). Specifically, we integrate a two-stage prompt processing scheme-first disrupting the input prompt, then rectifying it-into the conventional smoothing defense framework. This disrup
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0054LLM Jailbreak
Suggested from the entry's categories.
Cite
@misc{lin2026guaranteed,
title = {{Guaranteed Jailbreaking Defense via Disrupt-and-Rectify Smoothing}},
author = {Zheng Lin and Zhenxing Niu and Haoxuan Ji and Haichang Gao},
year = {2026},
month = may,
eprint = {2605.10582},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2605.10582}
}