Skip to content
Search
paperJuly 2026Unreviewed

GPT-Red: Automated Red Teaming via Self-Play at Scale

Eric Wallace, Christopher A. Choquette-Choo, Nikhil Kandpal, Sam Toyer, Dylan Hunn, Stephanie Lin, Yuxin Wen, Xiangyu Qi, Christopher Wolff, Zizhao Wang, Milad Nasr, Sicheng Zhu, Chuan Guo, Juan Felipe Cerón Uribe, Kaiwen Wang, Aiden Low, Kai Xiao, Kai Chen

Abstract

We introduce \textbf{GPT-Red}, an automated red-teaming agent that is trained to discover novel prompt injection attacks against frontier LLMs. The goal of this model is to evaluate and improve the robustness of our production systems. To this end, we use it to adversarially train GPT-5.6, our most robust model to prompt injections to date. To create GPT-Red, we design a scalable self-play algorithm where the model is tasked with attacking a diverse population of simultaneously-trained defender

Categories

Framework mappings

MITRE ATLAS
  • AML.T0051LLM Prompt Injection

Suggested from the entry's categories.

Cite

@misc{wallace2026gptred,
  title = {{GPT-Red: Automated Red Teaming via Self-Play at Scale}},
  author = {Eric Wallace and Christopher A. Choquette-Choo and Nikhil Kandpal and Sam Toyer and Dylan Hunn and Stephanie Lin and Yuxin Wen and Xiangyu Qi and Christopher Wolff and Zizhao Wang and Milad Nasr and Sicheng Zhu and Chuan Guo and Juan Felipe Cerón Uribe and Kaiwen Wang and Aiden Low and Kai Xiao and Kai Chen},
  year = {2026},
  month = jul,
  eprint = {2607.26115},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2607.26115}
}