Skip to content
Search
paperAugust 2026Unreviewed

COMIC: Reference-Aware Safety Gating for Multimodal Large Language Models

Md Abdullahil Oaphy, Anhao Xiang, Zongxing Xie, Huayue Gu, Chenyu Wang, Honghui Xu

Abstract

Multimodal large language models (MLLMs) are increasingly used to interact with screenshots, scanned documents, diagrams, and other visually grounded inputs. This shift introduces a new safety risk: in many multimodal jailbreaks, neither the prompt nor the image is harmful in isolation. Unsafe behavior emerges only when the model binds an apparently benign operation, such as summarizing, translating, or following, to a localized visual target. This reveals a structural weakness in current multim

Categories

Framework mappings

MITRE ATLAS
  • AML.T0054LLM Jailbreak

Suggested from the entry's categories.

Cite

@misc{oaphy2026comic,
  title = {{COMIC: Reference-Aware Safety Gating for Multimodal Large Language Models}},
  author = {Md Abdullahil Oaphy and Anhao Xiang and Zongxing Xie and Huayue Gu and Chenyu Wang and Honghui Xu},
  year = {2026},
  month = aug,
  eprint = {2608.17234},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2608.17234}
}