Skip to content
Search
paperAugust 2026Unreviewed

Agent Against Agent: An Agentic System for Automatic Prompt Injection Red Teaming

Yanting Wang, Chenlong Yin, Runpeng Geng, Jinyuan Jia

Abstract

Prompt injection poses significant security risks to LLM agents. Efficient and effective red-teaming is therefore critical, both for evaluating these risks and for collecting training data to improve defenses. Existing state-of-the-art prompt injection red-teaming methods primarily rely on reinforcement learning (RL), producing attacker models that often generalize poorly to new target LLMs. In this work, we develop PIMiner, an agentic system for prompt injection red-teaming. During training, PI

Categories

Framework mappings

MITRE ATLAS
  • AML.T0051LLM Prompt Injection

Suggested from the entry's categories.

Cite

@misc{wang2026agent,
  title = {{Agent Against Agent: An Agentic System for Automatic Prompt Injection Red Teaming}},
  author = {Yanting Wang and Chenlong Yin and Runpeng Geng and Jinyuan Jia},
  year = {2026},
  month = aug,
  eprint = {2608.05108},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2608.05108}
}