Skip to content
Search
paperMay 2026Unreviewed

AgentShield: Deception-based Compromise Detection for Tool-using LLM Agents

Yassin H. Rassul, Tarik A. Rashid

Abstract

Defenses against indirect prompt injection (IPI) in tool-using LLM agents share two structural weaknesses. First, they all attempt to prevent attacks rather than detect the compromises that slip through. Second, they have only been evaluated in English, leaving users of low-resource languages such as Kurdish and Arabic without tested protection. This paper addresses both gaps with AgentShield, a deception-based detection framework that places three layers of traps inside the agent's tool interfa

Categories

Framework mappings

MITRE ATLAS
  • AML.T0051LLM Prompt Injection

Suggested from the entry's categories.

Cite

@misc{rassul2026agentshield,
  title = {{AgentShield: Deception-based Compromise Detection for Tool-using LLM Agents}},
  author = {Yassin H. Rassul and Tarik A. Rashid},
  year = {2026},
  month = may,
  eprint = {2605.11026},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2605.11026}
}