Skip to content
Search
paperJanuary 2024ReviewedOpen access

Adversarial Attacks and Defenses in Large Language Models: Old and New Threats

Leo Schwinn, David Dobre, Stephan Gunnemann, Gauthier Gidel

arXiv preprint

Abstract

Systematizes adversarial attacks and defenses for LLMs, connecting them to the classical adversarial ML literature while identifying LLM-specific threats.

Categories

#survey#adversarial-ML#systematization

Framework mappings

MITRE ATLAS
  • AML.T0043Craft Adversarial Data
  • AML.T0054LLM Jailbreak

Cite

@misc{schwinn2024adversarial,
  title = {{Adversarial Attacks and Defenses in Large Language Models: Old and New Threats}},
  author = {Leo Schwinn and David Dobre and Stephan Gunnemann and Gauthier Gidel},
  year = {2024},
  month = jan,
  eprint = {2310.19737},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2310.19737}
}