July 2026Unreviewed
Evolving Safety Landscape of Multi-modal Large Language Models: A Survey of Emerging Threats and Safeguards
Xi Li, Shu Zhao, Xiaohan Zou, Fei Zhao, Fuxiao Liu, Yusen Zhang, Cheng Han, Yushun Dong, Jiaqi Wang
Abstract
Multi-modal large language models (MLLMs) integrate heterogeneous modalities through modality alignment and fusion, enabling stronger understanding and reasoning. However, this architectural shift reshapes the safety landscape of machine learning. Increased model complexity and cross-modal interactions give rise to novel threats, including compromised modality integration, modality misalignment, and fused safety risks, reflecting shifts in threat modeling beyond uni-modal assumptions. These shif
Categories
Cite
@misc{li2026evolving,
title = {{Evolving Safety Landscape of Multi-modal Large Language Models: A Survey of Emerging Threats and Safeguards}},
author = {Xi Li and Shu Zhao and Xiaohan Zou and Fei Zhao and Fuxiao Liu and Yusen Zhang and Cheng Han and Yushun Dong and Jiaqi Wang},
year = {2026},
month = jul,
eprint = {2608.07535},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2608.07535}
}