Skip to content
Search
paperMay 2025Unreviewed

Attack and defense techniques in large language models: A survey and new perspectives

Zhiyu Liao, Kang Chen, Y. Lin, Kangkang Li, Yunxuan Liu, Hefeng Chen, Xingwang Huang, Yuanhui Yu

Neural Networks

Abstract

Large Language Models (LLMs) have become central to numerous natural language processing tasks, but their vulnerabilities present significant security and ethical challenges. This systematic survey explores the evolving landscape of attack and defense techniques in LLMs. We classify attacks into adversarial prompt attacks, optimized attacks, model theft, as well as attacks on LLM applications, detailing their mechanisms and implications. Consequently, we analyze defense strategies, such as preve

Categories

Framework mappings

MITRE ATLAS
  • AML.T0024.002Extract AI Model

Suggested from the entry's categories.

Cite

@article{liao2025attack,
  title = {{Attack and defense techniques in large language models: A survey and new perspectives}},
  author = {Zhiyu Liao and Kang Chen and Y. Lin and Kangkang Li and Yunxuan Liu and Hefeng Chen and Xingwang Huang and Yuanhui Yu},
  year = {2025},
  month = may,
  journal = {Neural Networks},
  eprint = {2505.00976},
  archivePrefix = {arXiv},
  doi = {10.48550/arXiv.2505.00976},
  url = {https://www.semanticscholar.org/paper/5bce864b579b376c028ec40a8fec0f999b005d0e}
}