Skip to content
Search
paperMay 2026Unreviewed

A Cross-Modal Prompt Injection Attack against Large Vision-Language Models with Image-Only Perturbation

Hao Yang, Zhuo Ma, Yang Liu, Yilong Yang, Guancheng Wang, JianFeng Ma

Abstract

Large vision-language models (LVLMs) have emerged as a powerful paradigm for multimodal intelligence, but their growing deployment also expands the attack surface of prompt injection. Despite this growing concern, existing attacks still suffer from a critical limitation: the injected prompt for one modality only steers the model's interpretation of that singular input. Alternatively, these attacks remain multimodal but fail to achieve cross-modal prompt perturbation. To bridge this gap, we intro

Categories

Framework mappings

MITRE ATLAS
  • AML.T0051LLM Prompt Injection

Suggested from the entry's categories.

Cite

@misc{yang2026crossmodal,
  title = {{A Cross-Modal Prompt Injection Attack against Large Vision-Language Models with Image-Only Perturbation}},
  author = {Hao Yang and Zhuo Ma and Yang Liu and Yilong Yang and Guancheng Wang and JianFeng Ma},
  year = {2026},
  month = may,
  eprint = {2605.16090},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2605.16090}
}