May 2026Unreviewed
Adversarial-Test-Driven Multi-Agent LLM Defense: A Self-Evolving Framework via Inference-Time Prompt Optimization
Yang Qu, Yuwei He, Lei Cao, Juzheng Wang, Sulei Li, Hongxi Chen
Electronics
Abstract
Large Language Models (LLMs) remain highly susceptible to jailbreak attacks that bypass safety alignments through sophisticated prompt manipulation. While multi-agent defense systems have emerged as a promising countermeasure, existing frameworks predominantly rely on static agent designs, which struggle to adapt to evolving adversarial strategies. To bridge this gap, we propose an Adversarial-Test-Driven Multi-Agent Defense framework that shifts the focus from model-level fine-tuning to system-
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0054LLM Jailbreak
Suggested from the entry's categories.
Cite
@article{qu2026adversarialtestdriven,
title = {{Adversarial-Test-Driven Multi-Agent LLM Defense: A Self-Evolving Framework via Inference-Time Prompt Optimization}},
author = {Yang Qu and Yuwei He and Lei Cao and Juzheng Wang and Sulei Li and Hongxi Chen},
year = {2026},
month = may,
journal = {Electronics},
doi = {10.3390/electronics15112365},
url = {https://www.semanticscholar.org/paper/cdc9ba46a28837d6157b523d6b784f9e2f82cca8}
}