Skip to content
Search
paperApril 2026Unreviewed

RouteGuard: Internal-Signal Detection of Skill Poisoning in LLM Agents

Wenjie Xiao, Xuehai Tang, Biyu Zhou, Songlin Hu, Jizhong Han

Abstract

Agent skills introduce a new and more severe form of indirect injection for LLM agents: unlike traditional indirect prompt injection, attackers can hide malicious instructions inside a dense, action-oriented skill that already functions as a legitimate instruction source. We study pre-execution skill-poison detection and show that successful skill poisoning induces a structured internal effect, attention hijacking, in which response-time attention shifts from trusted context to malicious skill s

Categories

Framework mappings

OWASP Top 10 for LLM Applications
  • LLM01Prompt Injection
  • LLM04Data and Model Poisoning
MITRE ATLAS
  • AML.T0020Poison Training Data
  • AML.T0051LLM Prompt Injection

Suggested from the entry's categories.

Cite

@misc{xiao2026routeguard,
  title = {{RouteGuard: Internal-Signal Detection of Skill Poisoning in LLM Agents}},
  author = {Wenjie Xiao and Xuehai Tang and Biyu Zhou and Songlin Hu and Jizhong Han},
  year = {2026},
  month = apr,
  eprint = {2604.22888},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2604.22888}
}