← Back to search
paper llmsec-2026-00156

Adversarial Attacks on Multimodal Large Language Models: A Comprehensive Survey

Bhavuk Jain, Sercan Ö. Arık, Hardeo K. Thakur

2026-03

Abstract

Multimodal large language models (MLLMs) integrate information from multiple modalities such as text, images, audio, and video, enabling complex capabilities such as visual question answering and audio translation. While powerful, this increased expressiveness introduces new and amplified vulnerabilities to adversarial manipulation. This survey provides a comprehensive and systematic analysis of adversarial threats to MLLMs, moving beyond enumerating attack techniques to explain the underlying c

Cite This Resource

@article{llmsec202600156,
  title = {Adversarial Attacks on Multimodal Large Language Models: A Comprehensive Survey},
  author = {Bhavuk Jain and Sercan Ö. Arık and Hardeo K. Thakur},
  year = {2026},
  url = {https://arxiv.org/abs/2603.27918},
}

Metadata

Added
2026-05-17
Added by
automation
Source
arxiv
arxiv_id
2603.27918