March 2026Unreviewed
Adversarial Attacks on Multimodal Large Language Models: A Comprehensive Survey
Bhavuk Jain, Sercan Ö. Arık, Hardeo K. Thakur
Abstract
Multimodal large language models (MLLMs) integrate information from multiple modalities such as text, images, audio, and video, enabling complex capabilities such as visual question answering and audio translation. While powerful, this increased expressiveness introduces new and amplified vulnerabilities to adversarial manipulation. This survey provides a comprehensive and systematic analysis of adversarial threats to MLLMs, moving beyond enumerating attack techniques to explain the underlying c
Categories
Framework mappings
MITRE ATLAS
- AML.T0043Craft Adversarial Data
Suggested from the entry's categories.
Cite
@misc{jain2026adversarial,
title = {{Adversarial Attacks on Multimodal Large Language Models: A Comprehensive Survey}},
author = {Bhavuk Jain and Sercan Ö. Arık and Hardeo K. Thakur},
year = {2026},
month = mar,
eprint = {2603.27918},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2603.27918}
}