Skip to content
Search
paperJuly 2026Unreviewed

Overloading Large Vision-Language Models for Jailbreaking

Haoyu Zhang, Yangyang Guo, Mohan Kankanhalli

Abstract

Large Vision-Language Models (LVLMs) exhibit remarkable vision-language capabilities and are increasingly deployed in real-world applications such as personal assistants, document analysis systems, and embodied agents. However, their dual-modal attack surfaces make them vulnerable to jailbreak attacks. Existing LVLM jailbreaks rely on simple designs, e.g., short text and out-of-distribution images. Nevertheless, recent advancements in both large language model backbones and multimodal mechanisms

Categories

Framework mappings

MITRE ATLAS
  • AML.T0054LLM Jailbreak

Suggested from the entry's categories.

Cite

@misc{zhang2026overloading,
  title = {{Overloading Large Vision-Language Models for Jailbreaking}},
  author = {Haoyu Zhang and Yangyang Guo and Mohan Kankanhalli},
  year = {2026},
  month = jul,
  eprint = {2607.02961},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2607.02961}
}