Skip to content
Search
paperAugust 2026UnreviewedOpen access

Evaluating Indirect Prompt Injection Defenses in Tool-Using LLM Agents: Security, Utility, and Replication

Adil Khan, Khaled AlKhanbashi, Azza Mohamed

Computers

Abstract

Large language model (LLM) agents that retrieve external content and use tools are vulnerable to indirect prompt injection, in which untrusted content contains instructions intended to influence agent behavior. We evaluated four defenses and an undefended control across GPT-5.4, GPT-5.4-mini, and Claude Sonnet 4.6 on the AgentDojo banking benchmark (Tool Filter was evaluated only for the OpenAI models), reporting attack success rate (ASR), benign utility, utility under attack, operational measur

Categories

Framework mappings

MITRE ATLAS
  • AML.T0051LLM Prompt Injection

Suggested from the entry's categories.

Cite

@article{khan2026evaluating,
  title = {{Evaluating Indirect Prompt Injection Defenses in Tool-Using LLM Agents: Security, Utility, and Replication}},
  author = {Adil Khan and Khaled AlKhanbashi and Azza Mohamed},
  year = {2026},
  month = aug,
  journal = {Computers},
  doi = {10.3390/computers15090570},
  url = {https://www.semanticscholar.org/paper/857a1c7d30ff72cce69c0067e8c4daa739bf41f3}
}