April 2026Unreviewed
RouteGuard: Internal-Signal Detection of Skill Poisoning in LLM Agents
Wenjie Xiao, Xuehai Tang, Biyu Zhou, Songlin Hu, Jizhong Han
Abstract
Agent skills introduce a new and more severe form of indirect injection for LLM agents: unlike traditional indirect prompt injection, attackers can hide malicious instructions inside a dense, action-oriented skill that already functions as a legitimate instruction source. We study pre-execution skill-poison detection and show that successful skill poisoning induces a structured internal effect, attention hijacking, in which response-time attention shifts from trusted context to malicious skill s
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
- LLM04Data and Model Poisoning
MITRE ATLAS
- AML.T0020Poison Training Data
- AML.T0051LLM Prompt Injection
Suggested from the entry's categories.
Cite
@misc{xiao2026routeguard,
title = {{RouteGuard: Internal-Signal Detection of Skill Poisoning in LLM Agents}},
author = {Wenjie Xiao and Xuehai Tang and Biyu Zhou and Songlin Hu and Jizhong Han},
year = {2026},
month = apr,
eprint = {2604.22888},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2604.22888}
}