Skip to content
Search
paperJune 2026Unreviewed

AgentCanary: A Security Evaluation Framework for Autonomous AI Agents in Real Executable Environments

Peiyang Li, Songping Wang, Yi Huang, Yanhua Shi, Chenhao Zhang, Qi Li, Yueming Lyu, Caifeng Shan, Fengting Li, Chao Feng, Chuanqun Zhu, Liang Chen

Abstract

Autonomous AI agents have driven the transition from conversation to task execution, shifting security failures from textual deception to system compromise. Although security evaluation is crucial for proactive risk prevention, prior work is constrained by fundamental bottlenecks, including fragmented risk coverage, static or low-fidelity execution environments, and single-dimensional and coarse-grained assessment metrics. To address these challenges, we propose AgentCanary, a comprehensive secu

Categories

Cite

@misc{li2026agentcanary,
  title = {{AgentCanary: A Security Evaluation Framework for Autonomous AI Agents in Real Executable Environments}},
  author = {Peiyang Li and Songping Wang and Yi Huang and Yanhua Shi and Chenhao Zhang and Qi Li and Yueming Lyu and Caifeng Shan and Fengting Li and Chao Feng and Chuanqun Zhu and Liang Chen},
  year = {2026},
  month = jun,
  eprint = {2606.10484},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2606.10484}
}