October 2025Unreviewed
Proactive defense against LLM Jailbreak
Weiliang Zhao, Jinjun Peng, Daniel Ben-Levi, Zhou Yu, Junfeng Yang
arXiv.org
Abstract
The proliferation of powerful large language models (LLMs) has necessitated robust safety alignment, yet these models remain vulnerable to evolving adversarial attacks, including multi-turn jailbreaks that iteratively search for successful queries. Current defenses, which are primarily reactive and static, often fail to handle these iterative attacks. In this paper, we introduce ProAct, a novel proactive defense framework designed to disrupt and mislead these iterative search jailbreak methods.
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0043Craft Adversarial Data
- AML.T0054LLM Jailbreak
Suggested from the entry's categories.
Cite
@misc{zhao2025proactive,
title = {{Proactive defense against LLM Jailbreak}},
author = {Weiliang Zhao and Jinjun Peng and Daniel Ben-Levi and Zhou Yu and Junfeng Yang},
year = {2025},
month = oct,
eprint = {2510.05052},
archivePrefix = {arXiv},
doi = {10.48550/arXiv.2510.05052},
url = {https://www.semanticscholar.org/paper/a5cb42e27c0207971690d47adbae8b633842fd3c}
}