August 2026UnreviewedOpen access
Evaluating Indirect Prompt Injection Defenses in Tool-Using LLM Agents: Security, Utility, and Replication
Adil Khan, Khaled AlKhanbashi, Azza Mohamed
Computers
Abstract
Large language model (LLM) agents that retrieve external content and use tools are vulnerable to indirect prompt injection, in which untrusted content contains instructions intended to influence agent behavior. We evaluated four defenses and an undefended control across GPT-5.4, GPT-5.4-mini, and Claude Sonnet 4.6 on the AgentDojo banking benchmark (Tool Filter was evaluated only for the OpenAI models), reporting attack success rate (ASR), benign utility, utility under attack, operational measur
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0051LLM Prompt Injection
Suggested from the entry's categories.
Cite
@article{khan2026evaluating,
title = {{Evaluating Indirect Prompt Injection Defenses in Tool-Using LLM Agents: Security, Utility, and Replication}},
author = {Adil Khan and Khaled AlKhanbashi and Azza Mohamed},
year = {2026},
month = aug,
journal = {Computers},
doi = {10.3390/computers15090570},
url = {https://www.semanticscholar.org/paper/857a1c7d30ff72cce69c0067e8c4daa739bf41f3}
}