Skip to content
Search
paperAugust 2026Unreviewed

MMAligner: Safeguarding Multimodal Large Language Models through Representation Calibration

Shenyi Zhang, Keyan Guo, Zihao Wang, Xuebin Li, Lingchen Zhao, Hongxin Hu, Chao Shen, Qian Wang

Abstract

Multimodal large language models (MLLMs) often refuse unsafe text prompts yet generate harmful responses to semantically equivalent multimodal inputs. Existing defenses either rely on external guardrails, which add inference overhead without repairing intrinsic flaws, or safety fine-tuning, which treats alignment as black-box optimization and may sacrifice utility or require large multimodal datasets. To identify the cause of this safety disparity, we analyze MLLM representations geometrically.

Categories

Cite

@misc{zhang2026mmaligner,
  title = {{MMAligner: Safeguarding Multimodal Large Language Models through Representation Calibration}},
  author = {Shenyi Zhang and Keyan Guo and Zihao Wang and Xuebin Li and Lingchen Zhao and Hongxin Hu and Chao Shen and Qian Wang},
  year = {2026},
  month = aug,
  eprint = {2608.05909},
  archivePrefix = {arXiv},
  url = {https://www.semanticscholar.org/paper/c3be36d17b18c2decb13515aef81d47f81e3366a}
}