Skip to content
Search
paperMarch 2026Unreviewed

Adversarial Attacks on Multimodal Large Language Models: A Comprehensive Survey

Bhavuk Jain, Sercan Ö. Arık, Hardeo K. Thakur

Abstract

Multimodal large language models (MLLMs) integrate information from multiple modalities such as text, images, audio, and video, enabling complex capabilities such as visual question answering and audio translation. While powerful, this increased expressiveness introduces new and amplified vulnerabilities to adversarial manipulation. This survey provides a comprehensive and systematic analysis of adversarial threats to MLLMs, moving beyond enumerating attack techniques to explain the underlying c

Categories

Framework mappings

MITRE ATLAS
  • AML.T0043Craft Adversarial Data

Suggested from the entry's categories.

Cite

@misc{jain2026adversarial,
  title = {{Adversarial Attacks on Multimodal Large Language Models: A Comprehensive Survey}},
  author = {Bhavuk Jain and Sercan Ö. Arık and Hardeo K. Thakur},
  year = {2026},
  month = mar,
  eprint = {2603.27918},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2603.27918}
}