January 2024ReviewedOpen access
Adversarial Attacks and Defenses in Large Language Models: Old and New Threats
Leo Schwinn, David Dobre, Stephan Gunnemann, Gauthier Gidel
arXiv preprint
Abstract
Systematizes adversarial attacks and defenses for LLMs, connecting them to the classical adversarial ML literature while identifying LLM-specific threats.
Categories
#survey#adversarial-ML#systematization
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0043Craft Adversarial Data
- AML.T0054LLM Jailbreak
Cite
@misc{schwinn2024adversarial,
title = {{Adversarial Attacks and Defenses in Large Language Models: Old and New Threats}},
author = {Leo Schwinn and David Dobre and Stephan Gunnemann and Gauthier Gidel},
year = {2024},
month = jan,
eprint = {2310.19737},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2310.19737}
}