May 2026Unreviewed
AIRGuard: Guarding Agent Actions with Runtime Authority Control
Suliu Qin, Haomin Zhuang, Yujun Zhou, Yufei Han, Xiangliang Zhang
Abstract
Tool-using language agents turn model decisions into external side effects: they read files, run scripts, call APIs, send messages, and invoke Model Context Protocol tools. This makes agent attacks different from jailbreaks. The harmful step is often not an obviously forbidden output, but an ordinary executable action that becomes unsafe because attacker-controlled context steers authorized access against the user's interest. We identify this failure mode as authority confusion: untrusted resour
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0054LLM Jailbreak
Suggested from the entry's categories.
Cite
@misc{qin2026airguard,
title = {{AIRGuard: Guarding Agent Actions with Runtime Authority Control}},
author = {Suliu Qin and Haomin Zhuang and Yujun Zhou and Yufei Han and Xiangliang Zhang},
year = {2026},
month = may,
eprint = {2605.28914},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2605.28914}
}