August 2026Unreviewed
RedEvoAgent: Automatic Red-Teaming Agent with Experience-Driven Skill Evolution
Junjie Zhang, Hui Liu, Kecheng Chen, Xianbo Mo, Changsheng Chen, Haoliang Li
Abstract
LLM-based agents are increasingly deployed in product-level execution harnesses, where jailbreaks can trigger harmful tool use and persistent state changes, creating greater risks than unsafe text generation alone. Existing automatic red-teaming methods often rely on fixed attacks, while recent agentic attackers coordinate multiple jailbreak tools and show stronger potential through trajectory-based retrieval. However, such retrieval can reuse misleading experiences due to retrieval bias and unc
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
OWASP Top 10 for Agentic Applications
- ASI02Tool Misuse & Exploitation
MITRE ATLAS
- AML.T0053AI Agent Tool Invocation
- AML.T0054LLM Jailbreak
NIST AI Risk Management Framework
- MEASUREMeasure
Suggested from the entry's categories.
Cite
@misc{zhang2026redevoagent,
title = {{RedEvoAgent: Automatic Red-Teaming Agent with Experience-Driven Skill Evolution}},
author = {Junjie Zhang and Hui Liu and Kecheng Chen and Xianbo Mo and Changsheng Chen and Haoliang Li},
year = {2026},
month = aug,
eprint = {2608.27439},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2608.27439}
}