July 2023Unreviewed
Jailbreak in pieces: Compositional Adversarial Attacks on Multi-Modal Language Models
Erfan Shayegani, Yue Dong, Nael B. Abu-Ghazaleh
International Conference on Learning Representations
Abstract
We introduce new jailbreak attacks on vision language models (VLMs), which use aligned LLMs and are resilient to text-only jailbreak attacks. Specifically, we develop cross-modality attacks on alignment where we pair adversarial images going through the vision encoder with textual prompts to break the alignment of the language model. Our attacks employ a novel compositional strategy that combines an image, adversarially targeted towards toxic embeddings, with generic prompts to accomplish the ja
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0043Craft Adversarial Data
- AML.T0054LLM Jailbreak
Suggested from the entry's categories.
Cite
@inproceedings{shayegani2023jailbreak,
title = {{Jailbreak in pieces: Compositional Adversarial Attacks on Multi-Modal Language Models}},
author = {Erfan Shayegani and Yue Dong and Nael B. Abu-Ghazaleh},
year = {2023},
month = jul,
booktitle = {International Conference on Learning Representations},
eprint = {2307.14539},
archivePrefix = {arXiv},
url = {https://www.semanticscholar.org/paper/92b9d8b8c81c4c53ea62000c0924500b2dd11bce}
}