Skip to content
Search
paperMay 2026Unreviewed

Adversarial-Test-Driven Multi-Agent LLM Defense: A Self-Evolving Framework via Inference-Time Prompt Optimization

Yang Qu, Yuwei He, Lei Cao, Juzheng Wang, Sulei Li, Hongxi Chen

Electronics

Abstract

Large Language Models (LLMs) remain highly susceptible to jailbreak attacks that bypass safety alignments through sophisticated prompt manipulation. While multi-agent defense systems have emerged as a promising countermeasure, existing frameworks predominantly rely on static agent designs, which struggle to adapt to evolving adversarial strategies. To bridge this gap, we propose an Adversarial-Test-Driven Multi-Agent Defense framework that shifts the focus from model-level fine-tuning to system-

Categories

Framework mappings

MITRE ATLAS
  • AML.T0054LLM Jailbreak

Suggested from the entry's categories.

Cite

@article{qu2026adversarialtestdriven,
  title = {{Adversarial-Test-Driven Multi-Agent LLM Defense: A Self-Evolving Framework via Inference-Time Prompt Optimization}},
  author = {Yang Qu and Yuwei He and Lei Cao and Juzheng Wang and Sulei Li and Hongxi Chen},
  year = {2026},
  month = may,
  journal = {Electronics},
  doi = {10.3390/electronics15112365},
  url = {https://www.semanticscholar.org/paper/cdc9ba46a28837d6157b523d6b784f9e2f82cca8}
}