May 2025Unreviewed
Attack and defense techniques in large language models: A survey and new perspectives
Zhiyu Liao, Kang Chen, Y. Lin, Kangkang Li, Yunxuan Liu, Hefeng Chen, Xingwang Huang, Yuanhui Yu
Neural Networks
Abstract
Large Language Models (LLMs) have become central to numerous natural language processing tasks, but their vulnerabilities present significant security and ethical challenges. This systematic survey explores the evolving landscape of attack and defense techniques in LLMs. We classify attacks into adversarial prompt attacks, optimized attacks, model theft, as well as attacks on LLM applications, detailing their mechanisms and implications. Consequently, we analyze defense strategies, such as preve
Categories
Framework mappings
MITRE ATLAS
- AML.T0024.002Extract AI Model
Suggested from the entry's categories.
Cite
@article{liao2025attack,
title = {{Attack and defense techniques in large language models: A survey and new perspectives}},
author = {Zhiyu Liao and Kang Chen and Y. Lin and Kangkang Li and Yunxuan Liu and Hefeng Chen and Xingwang Huang and Yuanhui Yu},
year = {2025},
month = may,
journal = {Neural Networks},
eprint = {2505.00976},
archivePrefix = {arXiv},
doi = {10.48550/arXiv.2505.00976},
url = {https://www.semanticscholar.org/paper/5bce864b579b376c028ec40a8fec0f999b005d0e}
}