August 2026Unreviewed
Agent Against Agent: An Agentic System for Automatic Prompt Injection Red Teaming
Yanting Wang, Chenlong Yin, Runpeng Geng, Jinyuan Jia
Abstract
Prompt injection poses significant security risks to LLM agents. Efficient and effective red-teaming is therefore critical, both for evaluating these risks and for collecting training data to improve defenses. Existing state-of-the-art prompt injection red-teaming methods primarily rely on reinforcement learning (RL), producing attacker models that often generalize poorly to new target LLMs. In this work, we develop PIMiner, an agentic system for prompt injection red-teaming. During training, PI
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0051LLM Prompt Injection
NIST AI Risk Management Framework
- MEASUREMeasure
Suggested from the entry's categories.
Cite
@misc{wang2026agent,
title = {{Agent Against Agent: An Agentic System for Automatic Prompt Injection Red Teaming}},
author = {Yanting Wang and Chenlong Yin and Runpeng Geng and Jinyuan Jia},
year = {2026},
month = aug,
eprint = {2608.05108},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2608.05108}
}