August 2026Unreviewed
Beyond Over-Refusal: Defending Indirect Prompt Injection via Latent Instruction Manifolds
Jiahao Chen, Rui Yin, Xinfeng Li, Qianli Ma, Tianyu Du, Zhihui Fu, Jun Wang, Zhaoxiang Wang, Shouling Ji
Abstract
Large Language Models (LLMs) have been integrated into complex ecosystems (e.g., Code Agents), while Indirect Prompt Injection (IPI) attacks have emerged as critical barriers to their safe deployment. Attackers exploit LLMs' indistinguishability between "instructions" and "data" to manipulate LLMs via maliciously injected instructions. Existing defenses, however, face an intractable safety-utility trade-off: most guardrails either incur high latency or suffer from severe over-refusal. In this pa
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0051LLM Prompt Injection
Suggested from the entry's categories.
Cite
@misc{chen2026beyond,
title = {{Beyond Over-Refusal: Defending Indirect Prompt Injection via Latent Instruction Manifolds}},
author = {Jiahao Chen and Rui Yin and Xinfeng Li and Qianli Ma and Tianyu Du and Zhihui Fu and Jun Wang and Zhaoxiang Wang and Shouling Ji},
year = {2026},
month = aug,
eprint = {2608.22248},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2608.22248}
}