Skip to content
Search
paperJuly 2026Unreviewed

PVDetector: Detecting Prompt Injection Attacks on Purpose-Specific LLM Agents through Policy-Violation Concept Analysis

Junhui Wang, Hangtao Zhang, Zhirun Zheng, Li Zeng, Jiejun Xiao, Xi Luo, Lihua Yin, Saiqin Long

Abstract

Large language models (LLMs) are increasingly deployed as purpose-specific agents to handle domain-specific tasks such as customer service and code generation. These agents are expected to comply with not only generic safety guardrails but also purpose-specific restrictions tailored to their designated roles. Such additional restrictions enlarge the attack surface, particularly to prompt injection (PI) attacks. To defend against such attacks, existing detection methods primarily rely on analyzin

Categories

Framework mappings

MITRE ATLAS
  • AML.T0051LLM Prompt Injection

Suggested from the entry's categories.

Cite

@misc{wang2026pvdetector,
  title = {{PVDetector: Detecting Prompt Injection Attacks on Purpose-Specific LLM Agents through Policy-Violation Concept Analysis}},
  author = {Junhui Wang and Hangtao Zhang and Zhirun Zheng and Li Zeng and Jiejun Xiao and Xi Luo and Lihua Yin and Saiqin Long},
  year = {2026},
  month = jul,
  eprint = {2607.12624},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2607.12624}
}