July 2023ReviewedOpen access
Universal and Transferable Adversarial Attacks on Aligned Language Models
Andy Zou, Zifan Wang, Nicholas Carlini, Milad Nasr, J. Zico Kolter, Matt Fredrikson
arXiv preprint
Abstract
Proposes an automated method (GCG) to generate adversarial suffixes that cause aligned LLMs to produce harmful content, with attacks transferring across models including ChatGPT and Claude.
Categories
#GCG#adversarial-suffix#transferability
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0043Craft Adversarial Data
- AML.T0054LLM Jailbreak
Cite
@misc{zou2023universal,
title = {{Universal and Transferable Adversarial Attacks on Aligned Language Models}},
author = {Andy Zou and Zifan Wang and Nicholas Carlini and Milad Nasr and J. Zico Kolter and Matt Fredrikson},
year = {2023},
month = jul,
eprint = {2307.15043},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2307.15043}
}