May 2026Unreviewed
Re-Triggering Safeguards within LLMs for Jailbreak Detection
Zheng Lin, Zhenxing Niu, Haoxuan Ji, Yuzhe Huang, Haichang Gao
Abstract
This paper proposes a jailbreaking prompt detection method for large language models (LLMs) to defend against jailbreak attacks. Although recent LLMs are equipped with built-in safeguards, it remains possible to craft jailbreaking prompts that bypass them. We argue that such jailbreaking prompts are inherently fragile, and thus introduce an embedding disruption method to re-activate the safeguards within LLMs. Unlike previous defense methods that aim to serve as standalone solutions, our approac
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0054LLM Jailbreak
Suggested from the entry's categories.
Cite
@misc{lin2026retriggering,
title = {{Re-Triggering Safeguards within LLMs for Jailbreak Detection}},
author = {Zheng Lin and Zhenxing Niu and Haoxuan Ji and Yuzhe Huang and Haichang Gao},
year = {2026},
month = may,
eprint = {2605.10611},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2605.10611}
}