Skip to content
Search
paperAugust 2026Unreviewed

Vis-Poison: Poisoning Visual Knowledge in Multimodal Retrieval-Augmented Generation

Ru-Jin Liang, Zhong-Pu Chen, Yuhao Lei, Xin Miao

Abstract

While multimodal retrieval-augmented generation (RAG) systems increasingly rely on images as external knowledge sources, the introduction of poisoned visual evidence can severely compromise multimodal large language model (MLLM) generation. Unlike prior attacks that rely on altering textual metadata, we introduce Vis-Poison, a novel visual knowledge poisoning attack where the poisoned image itself is the attacker-controlled payload, without manipulating captions, summaries, metadata, or other as

Categories

Framework mappings

OWASP Top 10 for LLM Applications
  • LLM04Data and Model Poisoning
MITRE ATLAS
  • AML.T0020Poison Training Data

Suggested from the entry's categories.

Cite

@misc{liang2026vispoison,
  title = {{Vis-Poison: Poisoning Visual Knowledge in Multimodal Retrieval-Augmented Generation}},
  author = {Ru-Jin Liang and Zhong-Pu Chen and Yuhao Lei and Xin Miao},
  year = {2026},
  month = aug,
  eprint = {2608.20756},
  archivePrefix = {arXiv},
  url = {https://www.semanticscholar.org/paper/cbe9563fd7efd2df8017dd23c0e4a6e31b9fc5da}
}