June 2026Unreviewed
Defending against Adaptive Prompt Injection Attacks via Reasoning-enabled Task Alignment
Lipeng He, Yihan Wang, Jiawen Zhang, N. Asokan
Abstract
Indirect prompt injection attacks hijack LLM-based agents by embedding malicious instructions in third-party data that the agent retrieves during task execution. Existing defenses report near-zero attack success rate on static benchmarks, yet recent adaptive evaluations show that these results collapse once the attacker is allowed to optimize against the deployed defense. In this work, we trace this collapse to two failure modes. First, existing defense methods are confined to recognizing specif
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0051LLM Prompt Injection
Suggested from the entry's categories.
Cite
@misc{he2026defending,
title = {{Defending against Adaptive Prompt Injection Attacks via Reasoning-enabled Task Alignment}},
author = {Lipeng He and Yihan Wang and Jiawen Zhang and N. Asokan},
year = {2026},
month = jun,
eprint = {2606.15441},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2606.15441}
}