Skip to content
Search
paperMay 2026Unreviewed

AI Agents May Always Fall for Prompt Injections

Sahar Abdelnabi, Eugene Bagdasarian

Abstract

Prompt injection is the most critical vulnerability in deployed AI agents. Despite recent progress, we show that the prevailing defense paradigm (data-instruction separation) both fails to detect attacks that operate through contextual manipulation and degrades contextually appropriate behavior. We then recast prompt injection via the lens of Contextual Integrity (CI), a privacy theory that judges information flow compliance with contextual norms. This explains types of attacks that current defe

Categories

Framework mappings

MITRE ATLAS
  • AML.T0051LLM Prompt Injection

Suggested from the entry's categories.

Cite

@misc{abdelnabi2026ai,
  title = {{AI Agents May Always Fall for Prompt Injections}},
  author = {Sahar Abdelnabi and Eugene Bagdasarian},
  year = {2026},
  month = may,
  eprint = {2605.17634},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2605.17634}
}