Skip to content
Search
paperJune 2026Unreviewed

Benign Inputs, Harmful Outputs: Cross-Modal Jailbreaking via Distributed Semantic Recomposition

Yani Wang, Yilong Yang, Yang Liu, Zhuzhu Wang, Zuobin Ying, Zhuo Ma

Abstract

Multimodal Large Language Models (MLLMs) have recently demonstrated remarkable capabilities in content synthesis and autonomous reasoning. Previous safety guardrails are primarily designed for unimodal textual input interception, leaving them vulnerable to cross-modal jailbreak attacks. However, regardless unimodal textual attack or cross-modal jailbreak, typically inclusive part of explicit harmful or sensitive content at the input level, which is called Harm-Bearing. It allow the model's safet

Categories

Framework mappings

MITRE ATLAS
  • AML.T0054LLM Jailbreak

Suggested from the entry's categories.

Cite

@misc{wang2026benign,
  title = {{Benign Inputs, Harmful Outputs: Cross-Modal Jailbreaking via Distributed Semantic Recomposition}},
  author = {Yani Wang and Yilong Yang and Yang Liu and Zhuzhu Wang and Zuobin Ying and Zhuo Ma},
  year = {2026},
  month = jun,
  eprint = {2606.01837},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2606.01837}
}