August 2026Unreviewed
Do VLMs Share Safety Neurons Across Modalities?
Jia-Xuan Li, Jia-Hao Zhang, D. Vo, H. Nguyen, Pride Kavumba, Koki Wataoka
Abstract
Vision-language models (VLMs) can comply with harmful requests delivered through images, even when their LLM backbones would refuse the same content in text. While prior work characterizes these jailbreaks empirically or at the representation level, how visual inputs perturb safety pathways at the neuron level remains uncharted. We close this gap with a causal, neuron-level analysis of safety mechanisms in 10 VLMs. We propose a two-stage detection pipeline with iterative ablation that accounts f
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0054LLM Jailbreak
Suggested from the entry's categories.
Cite
@misc{li2026do,
title = {{Do VLMs Share Safety Neurons Across Modalities?}},
author = {Jia-Xuan Li and Jia-Hao Zhang and D. Vo and H. Nguyen and Pride Kavumba and Koki Wataoka},
year = {2026},
month = aug,
eprint = {2608.30750},
archivePrefix = {arXiv},
url = {https://www.semanticscholar.org/paper/8a8800ed3022d7532c7daf5782a597449f10b2b0}
}