May 2026Unreviewed
A Cross-Modal Prompt Injection Attack against Large Vision-Language Models with Image-Only Perturbation
Hao Yang, Zhuo Ma, Yang Liu, Yilong Yang, Guancheng Wang, JianFeng Ma
Abstract
Large vision-language models (LVLMs) have emerged as a powerful paradigm for multimodal intelligence, but their growing deployment also expands the attack surface of prompt injection. Despite this growing concern, existing attacks still suffer from a critical limitation: the injected prompt for one modality only steers the model's interpretation of that singular input. Alternatively, these attacks remain multimodal but fail to achieve cross-modal prompt perturbation. To bridge this gap, we intro
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0051LLM Prompt Injection
Suggested from the entry's categories.
Cite
@misc{yang2026crossmodal,
title = {{A Cross-Modal Prompt Injection Attack against Large Vision-Language Models with Image-Only Perturbation}},
author = {Hao Yang and Zhuo Ma and Yang Liu and Yilong Yang and Guancheng Wang and JianFeng Ma},
year = {2026},
month = may,
eprint = {2605.16090},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2605.16090}
}