Skip to content
Search
paperOctober 2025Unreviewed

Proactive defense against LLM Jailbreak

Weiliang Zhao, Jinjun Peng, Daniel Ben-Levi, Zhou Yu, Junfeng Yang

arXiv.org

Abstract

The proliferation of powerful large language models (LLMs) has necessitated robust safety alignment, yet these models remain vulnerable to evolving adversarial attacks, including multi-turn jailbreaks that iteratively search for successful queries. Current defenses, which are primarily reactive and static, often fail to handle these iterative attacks. In this paper, we introduce ProAct, a novel proactive defense framework designed to disrupt and mislead these iterative search jailbreak methods.

Categories

Framework mappings

MITRE ATLAS
  • AML.T0043Craft Adversarial Data
  • AML.T0054LLM Jailbreak

Suggested from the entry's categories.

Cite

@misc{zhao2025proactive,
  title = {{Proactive defense against LLM Jailbreak}},
  author = {Weiliang Zhao and Jinjun Peng and Daniel Ben-Levi and Zhou Yu and Junfeng Yang},
  year = {2025},
  month = oct,
  eprint = {2510.05052},
  archivePrefix = {arXiv},
  doi = {10.48550/arXiv.2510.05052},
  url = {https://www.semanticscholar.org/paper/a5cb42e27c0207971690d47adbae8b633842fd3c}
}