May 2026Unreviewed
DecodingTrust-Agent Platform (DTap): A Controllable and Interactive Red-Teaming Platform for AI Agents
Zhaorun Chen, Xun Liu, Haibo Tong, Chengquan Guo, Yuzhou Nie, Jiawei Zhang, Mintong Kang, Chejian Xu, Qichang Liu, Xiaogeng Liu, Tianneng Shi, Chaowei Xiao, Sanmi Koyejo, Percy Liang, Wenbo Guo, Dawn Song, Bo Li
Abstract
AI agents are increasingly deployed across diverse domains to automate complex workflows through long-horizon and high-stakes action executions. Due to their high capability and flexibility, such agents raise significant security and safety concerns. A growing number of real-world incidents have shown that adversaries can easily manipulate agents into performing harmful actions, such as leaking API keys, deleting user data, or initiating unauthorized transactions. Evaluating agent security is in
Categories
Framework mappings
NIST AI Risk Management Framework
- MEASUREMeasure
Suggested from the entry's categories.
Cite
@misc{chen2026decodingtrustagent,
title = {{DecodingTrust-Agent Platform (DTap): A Controllable and Interactive Red-Teaming Platform for AI Agents}},
author = {Zhaorun Chen and Xun Liu and Haibo Tong and Chengquan Guo and Yuzhou Nie and Jiawei Zhang and Mintong Kang and Chejian Xu and Qichang Liu and Xiaogeng Liu and Tianneng Shi and Chaowei Xiao and Sanmi Koyejo and Percy Liang and Wenbo Guo and Dawn Song and Bo Li},
year = {2026},
month = may,
eprint = {2605.04808},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2605.04808}
}