Skip to content
Search
paperAugust 2026Unreviewed

Do VLMs Share Safety Neurons Across Modalities?

Jia-Xuan Li, Jia-Hao Zhang, D. Vo, H. Nguyen, Pride Kavumba, Koki Wataoka

Abstract

Vision-language models (VLMs) can comply with harmful requests delivered through images, even when their LLM backbones would refuse the same content in text. While prior work characterizes these jailbreaks empirically or at the representation level, how visual inputs perturb safety pathways at the neuron level remains uncharted. We close this gap with a causal, neuron-level analysis of safety mechanisms in 10 VLMs. We propose a two-stage detection pipeline with iterative ablation that accounts f

Categories

Framework mappings

MITRE ATLAS
  • AML.T0054LLM Jailbreak

Suggested from the entry's categories.

Cite

@misc{li2026do,
  title = {{Do VLMs Share Safety Neurons Across Modalities?}},
  author = {Jia-Xuan Li and Jia-Hao Zhang and D. Vo and H. Nguyen and Pride Kavumba and Koki Wataoka},
  year = {2026},
  month = aug,
  eprint = {2608.30750},
  archivePrefix = {arXiv},
  url = {https://www.semanticscholar.org/paper/8a8800ed3022d7532c7daf5782a597449f10b2b0}
}