August 2026Unreviewed
What Guides the Agent? Adjudicating Unauthorized Behavior via Localizing Behavior-Guiding Instructions
Yichao Gao, Yumo Zhang, Yunhao Yao, Haohua Du, Pu-Han Luo, Ruiqi Li, Zhiqiang Wang
Abstract
LLM agents integrated with external resources gain complex task capabilities, yet the unified natural-language context channel makes them vulnerable to injection attacks: untrusted external data may be dynamically parsed as behavior-guiding instructions during LLM inference, thereby subverting the agent's decision. Existing defenses focus on static detection or isolation of malicious content at the input/output level, remains insufficient for detecting such dynamic inducements that arise during
Categories
Cite
@misc{gao2026what,
title = {{What Guides the Agent? Adjudicating Unauthorized Behavior via Localizing Behavior-Guiding Instructions}},
author = {Yichao Gao and Yumo Zhang and Yunhao Yao and Haohua Du and Pu-Han Luo and Ruiqi Li and Zhiqiang Wang},
year = {2026},
month = aug,
eprint = {2608.24022},
archivePrefix = {arXiv},
url = {https://www.semanticscholar.org/paper/db973d82aa1cd4dd5cfc5628be6838c59c4a86a9}
}