Skip to content
Search
paperJuly 2026Unreviewed

Evolving Safety Landscape of Multi-modal Large Language Models: A Survey of Emerging Threats and Safeguards

Xi Li, Shu Zhao, Xiaohan Zou, Fei Zhao, Fuxiao Liu, Yusen Zhang, Cheng Han, Yushun Dong, Jiaqi Wang

Abstract

Multi-modal large language models (MLLMs) integrate heterogeneous modalities through modality alignment and fusion, enabling stronger understanding and reasoning. However, this architectural shift reshapes the safety landscape of machine learning. Increased model complexity and cross-modal interactions give rise to novel threats, including compromised modality integration, modality misalignment, and fused safety risks, reflecting shifts in threat modeling beyond uni-modal assumptions. These shif

Categories

Cite

@misc{li2026evolving,
  title = {{Evolving Safety Landscape of Multi-modal Large Language Models: A Survey of Emerging Threats and Safeguards}},
  author = {Xi Li and Shu Zhao and Xiaohan Zou and Fei Zhao and Fuxiao Liu and Yusen Zhang and Cheng Han and Yushun Dong and Jiaqi Wang},
  year = {2026},
  month = jul,
  eprint = {2608.07535},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2608.07535}
}