Skip to content
Search
paperAugust 2026Unreviewed

What Guides the Agent? Adjudicating Unauthorized Behavior via Localizing Behavior-Guiding Instructions

Yichao Gao, Yumo Zhang, Yunhao Yao, Haohua Du, Pu-Han Luo, Ruiqi Li, Zhiqiang Wang

Abstract

LLM agents integrated with external resources gain complex task capabilities, yet the unified natural-language context channel makes them vulnerable to injection attacks: untrusted external data may be dynamically parsed as behavior-guiding instructions during LLM inference, thereby subverting the agent's decision. Existing defenses focus on static detection or isolation of malicious content at the input/output level, remains insufficient for detecting such dynamic inducements that arise during

Categories

Cite

@misc{gao2026what,
  title = {{What Guides the Agent? Adjudicating Unauthorized Behavior via Localizing Behavior-Guiding Instructions}},
  author = {Yichao Gao and Yumo Zhang and Yunhao Yao and Haohua Du and Pu-Han Luo and Ruiqi Li and Zhiqiang Wang},
  year = {2026},
  month = aug,
  eprint = {2608.24022},
  archivePrefix = {arXiv},
  url = {https://www.semanticscholar.org/paper/db973d82aa1cd4dd5cfc5628be6838c59c4a86a9}
}