Skip to content
Search
paperAugust 2026Unreviewed

REDAgentBench: Executable Red Teaming and Faithful Measurement of LLM Agent Systems

Zixing Chen, Xingyuan Liu, Jie Zhu, Huaixia Dou, Shuo Jiang, Junhui Li, Lifan Guo, Feng Chen, Chi Zhang

Abstract

Large language model (LLM) agents combine language-based reasoning with external tools to perform complex tasks. Adversarial inputs can exploit interactions between the agent and its environment, causing the agent to violate safety policies during execution. Yet existing evaluations often reduce agent safety to a single attack success rate (ASR), collapsing exposure, execution, observation, and adjudication and potentially conflating actual violations with evidence visibility. We introduce REDAg

Categories

Framework mappings

Suggested from the entry's categories.

Cite

@misc{chen2026redagentbench,
  title = {{REDAgentBench: Executable Red Teaming and Faithful Measurement of LLM Agent Systems}},
  author = {Zixing Chen and Xingyuan Liu and Jie Zhu and Huaixia Dou and Shuo Jiang and Junhui Li and Lifan Guo and Feng Chen and Chi Zhang},
  year = {2026},
  month = aug,
  eprint = {2608.10669},
  archivePrefix = {arXiv},
  url = {https://www.semanticscholar.org/paper/4529381cbe510e8c5911b678654a62b46b3a1327}
}