Skip to content
Search
paperAugust 2026Unreviewed

Prompt-Based Jailbreaking of Leading LLM Chatbots: A Survey of Attacks and Defenses

Brynn Knowlton, Jovani Campa, Davide Gallo, Khalil Dajani, Nabeel Alzahrani

IEEE Transactions on Artificial Intelligence

Abstract

Generative artificial intelligence (AI) systems—particularly large language models (LLMs)—remain vulnerable to jailbreak attacks: adversarial prompts that bypass safeguards and elicit unsafe or restricted outputs. This survey synthesizes jailbreak research from 2023–2025, covering attack methods, defense strategies, and evaluation frameworks. Jailbreak techniques are grouped into five main categories: prompt-based injections, role-play conditioning, multiturn dialogue, multilingual or multimodal

Categories

Framework mappings

MITRE ATLAS
  • AML.T0054LLM Jailbreak

Suggested from the entry's categories.

Cite

@article{knowlton2026promptbased,
  title = {{Prompt-Based Jailbreaking of Leading LLM Chatbots: A Survey of Attacks and Defenses}},
  author = {Brynn Knowlton and Jovani Campa and Davide Gallo and Khalil Dajani and Nabeel Alzahrani},
  year = {2026},
  month = aug,
  journal = {IEEE Transactions on Artificial Intelligence},
  doi = {10.1109/TAI.2026.3665656},
  url = {https://www.semanticscholar.org/paper/9b0b450d7b38f5524e7e6a0aad293f3a521be348}
}