Skip to content
Search
paperSeptember 2026Unreviewed

Understanding In-Context Multimodal Jailbreaks via Posterior Reweighting

Xu Zhang, Dev Mistry, Xiang Xu, Ren Wang

Abstract

In-context learning (ICL) jailbreaks reveal a critical vulnerability in multimodal large language models (MLLMs): harmful demonstrations in the prompt can induce unsafe outputs without modifying model parameters. Despite extensive empirical evidence, existing work lacks a principled understanding of why such jailbreaks reliably succeed or how their effectiveness scales with context composition. We propose a posterior reweighting framework that models a safety-aligned MLLM as implicitly operating

Categories

Framework mappings

MITRE ATLAS
  • AML.T0054LLM Jailbreak

Suggested from the entry's categories.

Cite

@misc{zhang2026understanding,
  title = {{Understanding In-Context Multimodal Jailbreaks via Posterior Reweighting}},
  author = {Xu Zhang and Dev Mistry and Xiang Xu and Ren Wang},
  year = {2026},
  month = sep,
  eprint = {2609.10613},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2609.10613}
}