Skip to content
Search
paperAugust 2026Unreviewed

A Self-Evolving Multi-Agent Framework Defense against LLM Jailbreak Attacks

Tongyan Hu, Bryan Hooi

Abstract

Large language models (LLMs) remain vulnerable to jailbreak attacks that exploit techniques such as role-playing, obfuscation, code transformation, and multi-step indirection to elicit harmful outputs. As jailbreak strategies keep emerging, defenses have proliferated in an ongoing cat-and-mouse game, yet most remain static: their safety behavior is fixed at deployment, so they cannot accumulate defensive experience or adapt to unseen strategies. We propose a self-evolving test-time defense built

Categories

Framework mappings

MITRE ATLAS
  • AML.T0054LLM Jailbreak

Suggested from the entry's categories.

Cite

@misc{hu2026selfevolving,
  title = {{A Self-Evolving Multi-Agent Framework Defense against LLM Jailbreak Attacks}},
  author = {Tongyan Hu and Bryan Hooi},
  year = {2026},
  month = aug,
  eprint = {2608.26008},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2608.26008}
}