Skip to content
Search
paperMay 2026Unreviewed

AIRGuard: Guarding Agent Actions with Runtime Authority Control

Suliu Qin, Haomin Zhuang, Yujun Zhou, Yufei Han, Xiangliang Zhang

Abstract

Tool-using language agents turn model decisions into external side effects: they read files, run scripts, call APIs, send messages, and invoke Model Context Protocol tools. This makes agent attacks different from jailbreaks. The harmful step is often not an obviously forbidden output, but an ordinary executable action that becomes unsafe because attacker-controlled context steers authorized access against the user's interest. We identify this failure mode as authority confusion: untrusted resour

Categories

Framework mappings

MITRE ATLAS
  • AML.T0054LLM Jailbreak

Suggested from the entry's categories.

Cite

@misc{qin2026airguard,
  title = {{AIRGuard: Guarding Agent Actions with Runtime Authority Control}},
  author = {Suliu Qin and Haomin Zhuang and Yujun Zhou and Yufei Han and Xiangliang Zhang},
  year = {2026},
  month = may,
  eprint = {2605.28914},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2605.28914}
}