August 2026Unreviewed
Vis-Poison: Poisoning Visual Knowledge in Multimodal Retrieval-Augmented Generation
Ru-Jin Liang, Zhong-Pu Chen, Yuhao Lei, Xin Miao
Abstract
While multimodal retrieval-augmented generation (RAG) systems increasingly rely on images as external knowledge sources, the introduction of poisoned visual evidence can severely compromise multimodal large language model (MLLM) generation. Unlike prior attacks that rely on altering textual metadata, we introduce Vis-Poison, a novel visual knowledge poisoning attack where the poisoned image itself is the attacker-controlled payload, without manipulating captions, summaries, metadata, or other as
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM04Data and Model Poisoning
MITRE ATLAS
- AML.T0020Poison Training Data
Suggested from the entry's categories.
Cite
@misc{liang2026vispoison,
title = {{Vis-Poison: Poisoning Visual Knowledge in Multimodal Retrieval-Augmented Generation}},
author = {Ru-Jin Liang and Zhong-Pu Chen and Yuhao Lei and Xin Miao},
year = {2026},
month = aug,
eprint = {2608.20756},
archivePrefix = {arXiv},
url = {https://www.semanticscholar.org/paper/cbe9563fd7efd2df8017dd23c0e4a6e31b9fc5da}
}