← Back to search
paper llmsec-2026-00104

Guaranteed Jailbreaking Defense via Disrupt-and-Rectify Smoothing

Zheng Lin, Zhenxing Niu, Haoxuan Ji, Haichang Gao

2026-05

Abstract

This paper proposes a guaranteed defense method for large language models (LLMs) to safeguard against jailbreaking attacks. Drawing inspiration from the denoised-smoothing approach in the adversarial defense domain, we propose a novel smoothing-based defense method, termed Disrupt-and-Rectify Smoothing (DR-Smoothing). Specifically, we integrate a two-stage prompt processing scheme-first disrupting the input prompt, then rectifying it-into the conventional smoothing defense framework. This disrup

Categories

Cite This Resource

@article{llmsec202600104,
  title = {Guaranteed Jailbreaking Defense via Disrupt-and-Rectify Smoothing},
  author = {Zheng Lin and Zhenxing Niu and Haoxuan Ji and Haichang Gao},
  year = {2026},
  url = {https://arxiv.org/abs/2605.10582},
}

Metadata

Added
2026-05-17
Added by
automation
Source
arxiv
arxiv_id
2605.10582