← Back to search
paper llmsec-2026-00103

Re-Triggering Safeguards within LLMs for Jailbreak Detection

Zheng Lin, Zhenxing Niu, Haoxuan Ji, Yuzhe Huang, Haichang Gao

2026-05

Abstract

This paper proposes a jailbreaking prompt detection method for large language models (LLMs) to defend against jailbreak attacks. Although recent LLMs are equipped with built-in safeguards, it remains possible to craft jailbreaking prompts that bypass them. We argue that such jailbreaking prompts are inherently fragile, and thus introduce an embedding disruption method to re-activate the safeguards within LLMs. Unlike previous defense methods that aim to serve as standalone solutions, our approac

Categories

Cite This Resource

@article{llmsec202600103,
  title = {Re-Triggering Safeguards within LLMs for Jailbreak Detection},
  author = {Zheng Lin and Zhenxing Niu and Haoxuan Ji and Yuzhe Huang and Haichang Gao},
  year = {2026},
  url = {https://arxiv.org/abs/2605.10611},
}

Metadata

Added
2026-05-17
Added by
automation
Source
arxiv
arxiv_id
2605.10611