Skip to content
Search
paperOctober 2025Unreviewed

SoK: Systematizing LLM Prompt Security: Taxonomies, Datasets, and Unified Evaluation of Attacks and Defenses

Hanbin Hong, Shuangqiao Wu, Shuya Feng, Nima Naderloui, Shenao Yan, Jingyue Zhang, Ali Arastehfard, Heqing Huang, Yuan Hong

Abstract

Large Language Models (LLMs) are increasingly used as interfaces to information, code, and real-world services, making prompt-level security failures a practical concern. Although jailbreak attacks, defenses, datasets, and automated judgers have advanced rapidly, evaluation remains fragmented across threat models, access assumptions, cost budgets, datasets, and success criteria. This makes reported attack success rates and defense gains hard to compare. This SoK systematizes LLM prompt security

Categories

Framework mappings

MITRE ATLAS
  • AML.T0054LLM Jailbreak

Suggested from the entry's categories.

Cite

@misc{hong2025sok,
  title = {{SoK: Systematizing LLM Prompt Security: Taxonomies, Datasets, and Unified Evaluation of Attacks and Defenses}},
  author = {Hanbin Hong and Shuangqiao Wu and Shuya Feng and Nima Naderloui and Shenao Yan and Jingyue Zhang and Ali Arastehfard and Heqing Huang and Yuan Hong},
  year = {2025},
  month = oct,
  eprint = {2510.15476},
  archivePrefix = {arXiv},
  url = {https://www.semanticscholar.org/paper/9f8a3c6f1fca4aeb8364bdb767b2e14e92376d5d}
}