August 2026Unreviewed
REDAgentBench: Executable Red Teaming and Faithful Measurement of LLM Agent Systems
Zixing Chen, Xingyuan Liu, Jie Zhu, Huaixia Dou, Shuo Jiang, Junhui Li, Lifan Guo, Feng Chen, Chi Zhang
Abstract
Large language model (LLM) agents combine language-based reasoning with external tools to perform complex tasks. Adversarial inputs can exploit interactions between the agent and its environment, causing the agent to violate safety policies during execution. Yet existing evaluations often reduce agent safety to a single attack success rate (ASR), collapsing exposure, execution, observation, and adjudication and potentially conflating actual violations with evidence visibility. We introduce REDAg
Categories
Framework mappings
NIST AI Risk Management Framework
- MEASUREMeasure
Suggested from the entry's categories.
Cite
@misc{chen2026redagentbench,
title = {{REDAgentBench: Executable Red Teaming and Faithful Measurement of LLM Agent Systems}},
author = {Zixing Chen and Xingyuan Liu and Jie Zhu and Huaixia Dou and Shuo Jiang and Junhui Li and Lifan Guo and Feng Chen and Chi Zhang},
year = {2026},
month = aug,
eprint = {2608.10669},
archivePrefix = {arXiv},
url = {https://www.semanticscholar.org/paper/4529381cbe510e8c5911b678654a62b46b3a1327}
}