February 2024Unreviewed
Agent Smith: A Single Image Can Jailbreak One Million Multimodal LLM Agents Exponentially Fast
Xiangming Gu, Xiaosen Zheng, Tianyu Pang, Chao Du, Qian Liu, Ye Wang, Jing Jiang, Min Lin
International Conference on Machine Learning
Abstract
A multimodal large language model (MLLM) agent can receive instructions, capture images, retrieve histories from memory, and decide which tools to use. Nonetheless, red-teaming efforts have revealed that adversarial images/prompts can jailbreak an MLLM and cause unaligned behaviors. In this work, we report an even more severe safety issue in multi-agent environments, referred to as infectious jailbreak. It entails the adversary simply jailbreaking a single agent, and without any further interven
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0054LLM Jailbreak
NIST AI Risk Management Framework
- MEASUREMeasure
Suggested from the entry's categories.
Cite
@inproceedings{gu2024agent,
title = {{Agent Smith: A Single Image Can Jailbreak One Million Multimodal LLM Agents Exponentially Fast}},
author = {Xiangming Gu and Xiaosen Zheng and Tianyu Pang and Chao Du and Qian Liu and Ye Wang and Jing Jiang and Min Lin},
year = {2024},
month = feb,
booktitle = {International Conference on Machine Learning},
eprint = {2402.08567},
archivePrefix = {arXiv},
doi = {10.48550/arXiv.2402.08567},
url = {https://www.semanticscholar.org/paper/b0ada492ba48e85016cbbfd95ec7180fb7e79648}
}