Skip to content
Search
paperAugust 2026Unreviewed

Beyond Over-Refusal: Defending Indirect Prompt Injection via Latent Instruction Manifolds

Jiahao Chen, Rui Yin, Xinfeng Li, Qianli Ma, Tianyu Du, Zhihui Fu, Jun Wang, Zhaoxiang Wang, Shouling Ji

Abstract

Large Language Models (LLMs) have been integrated into complex ecosystems (e.g., Code Agents), while Indirect Prompt Injection (IPI) attacks have emerged as critical barriers to their safe deployment. Attackers exploit LLMs' indistinguishability between "instructions" and "data" to manipulate LLMs via maliciously injected instructions. Existing defenses, however, face an intractable safety-utility trade-off: most guardrails either incur high latency or suffer from severe over-refusal. In this pa

Categories

Framework mappings

MITRE ATLAS
  • AML.T0051LLM Prompt Injection

Suggested from the entry's categories.

Cite

@misc{chen2026beyond,
  title = {{Beyond Over-Refusal: Defending Indirect Prompt Injection via Latent Instruction Manifolds}},
  author = {Jiahao Chen and Rui Yin and Xinfeng Li and Qianli Ma and Tianyu Du and Zhihui Fu and Jun Wang and Zhaoxiang Wang and Shouling Ji},
  year = {2026},
  month = aug,
  eprint = {2608.22248},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2608.22248}
}