August 2026Unreviewed
Prompt-Based Jailbreaking of Leading LLM Chatbots: A Survey of Attacks and Defenses
Brynn Knowlton, Jovani Campa, Davide Gallo, Khalil Dajani, Nabeel Alzahrani
IEEE Transactions on Artificial Intelligence
Abstract
Generative artificial intelligence (AI) systems—particularly large language models (LLMs)—remain vulnerable to jailbreak attacks: adversarial prompts that bypass safeguards and elicit unsafe or restricted outputs. This survey synthesizes jailbreak research from 2023–2025, covering attack methods, defense strategies, and evaluation frameworks. Jailbreak techniques are grouped into five main categories: prompt-based injections, role-play conditioning, multiturn dialogue, multilingual or multimodal
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0054LLM Jailbreak
Suggested from the entry's categories.
Cite
@article{knowlton2026promptbased,
title = {{Prompt-Based Jailbreaking of Leading LLM Chatbots: A Survey of Attacks and Defenses}},
author = {Brynn Knowlton and Jovani Campa and Davide Gallo and Khalil Dajani and Nabeel Alzahrani},
year = {2026},
month = aug,
journal = {IEEE Transactions on Artificial Intelligence},
doi = {10.1109/TAI.2026.3665656},
url = {https://www.semanticscholar.org/paper/9b0b450d7b38f5524e7e6a0aad293f3a521be348}
}