August 2026Unreviewed
A Self-Evolving Multi-Agent Framework Defense against LLM Jailbreak Attacks
Tongyan Hu, Bryan Hooi
Abstract
Large language models (LLMs) remain vulnerable to jailbreak attacks that exploit techniques such as role-playing, obfuscation, code transformation, and multi-step indirection to elicit harmful outputs. As jailbreak strategies keep emerging, defenses have proliferated in an ongoing cat-and-mouse game, yet most remain static: their safety behavior is fixed at deployment, so they cannot accumulate defensive experience or adapt to unseen strategies. We propose a self-evolving test-time defense built
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0054LLM Jailbreak
Suggested from the entry's categories.
Cite
@misc{hu2026selfevolving,
title = {{A Self-Evolving Multi-Agent Framework Defense against LLM Jailbreak Attacks}},
author = {Tongyan Hu and Bryan Hooi},
year = {2026},
month = aug,
eprint = {2608.26008},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2608.26008}
}