September 2026Unreviewed
Understanding In-Context Multimodal Jailbreaks via Posterior Reweighting
Xu Zhang, Dev Mistry, Xiang Xu, Ren Wang
Abstract
In-context learning (ICL) jailbreaks reveal a critical vulnerability in multimodal large language models (MLLMs): harmful demonstrations in the prompt can induce unsafe outputs without modifying model parameters. Despite extensive empirical evidence, existing work lacks a principled understanding of why such jailbreaks reliably succeed or how their effectiveness scales with context composition. We propose a posterior reweighting framework that models a safety-aligned MLLM as implicitly operating
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0054LLM Jailbreak
Suggested from the entry's categories.
Cite
@misc{zhang2026understanding,
title = {{Understanding In-Context Multimodal Jailbreaks via Posterior Reweighting}},
author = {Xu Zhang and Dev Mistry and Xiang Xu and Ren Wang},
year = {2026},
month = sep,
eprint = {2609.10613},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2609.10613}
}