July 2026Unreviewed
Overloading Large Vision-Language Models for Jailbreaking
Haoyu Zhang, Yangyang Guo, Mohan Kankanhalli
Abstract
Large Vision-Language Models (LVLMs) exhibit remarkable vision-language capabilities and are increasingly deployed in real-world applications such as personal assistants, document analysis systems, and embodied agents. However, their dual-modal attack surfaces make them vulnerable to jailbreak attacks. Existing LVLM jailbreaks rely on simple designs, e.g., short text and out-of-distribution images. Nevertheless, recent advancements in both large language model backbones and multimodal mechanisms
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0054LLM Jailbreak
Suggested from the entry's categories.
Cite
@misc{zhang2026overloading,
title = {{Overloading Large Vision-Language Models for Jailbreaking}},
author = {Haoyu Zhang and Yangyang Guo and Mohan Kankanhalli},
year = {2026},
month = jul,
eprint = {2607.02961},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2607.02961}
}