2024ReviewedOpen access
Are Aligned Neural Networks Adversarially Aligned?
Nicholas Carlini, Milad Nasr, Christopher A. Choquette-Choo, Matthew Jagielski, Irena Gao, Anas Awadalla, Pang Wei Koh, Daphne Ippolito, Katherine Lee, Florian Tramer, Ludwig Schmidt
NeurIPS 2023
Abstract
Evaluates whether multimodal LLMs aligned to refuse harmful text requests also refuse harmful image-based requests, finding significant gaps.
Categories
#multimodal#alignment-gap#image-attacks
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
Cite
@inproceedings{carlini2024are,
title = {{Are Aligned Neural Networks Adversarially Aligned?}},
author = {Nicholas Carlini and Milad Nasr and Christopher A. Choquette-Choo and Matthew Jagielski and Irena Gao and Anas Awadalla and Pang Wei Koh and Daphne Ippolito and Katherine Lee and Florian Tramer and Ludwig Schmidt},
year = {2024},
booktitle = {NeurIPS 2023},
eprint = {2306.15447},
archivePrefix = {arXiv},
doi = {10.52202/075280-2687},
url = {https://arxiv.org/abs/2306.15447}
}