July 2025Unreviewed
Breaking the Shield: Adversarial Jailbreak Attacks and Defense Mechanisms in Large Language Models
Shreya Dubey, H. Lamkuche
2025 International Conference on Emerging Information Technology and Engineering Solutions (EITES)
Abstract
Large Language Models (LLMs) are increasingly used across diverse applications, raising security concerns, especially regarding jail-break attacks that bypass safety measures. This paper presents a systematic security evaluation of LLMs, exploring the effectiveness of jail-break detection methods, the security progression across model versions, the relationship between model size and vulnerability, and the impact of combined defense strategies. We evaluate both open-source (e.g., LLama, Mistral)
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0054LLM Jailbreak
Suggested from the entry's categories.
Cite
@inproceedings{dubey2025breaking,
title = {{Breaking the Shield: Adversarial Jailbreak Attacks and Defense Mechanisms in Large Language Models}},
author = {Shreya Dubey and H. Lamkuche},
year = {2025},
month = jul,
booktitle = {2025 International Conference on Emerging Information Technology and Engineering Solutions (EITES)},
doi = {10.1109/EITES66543.2025.00023},
url = {https://www.semanticscholar.org/paper/01f493a008c386b93901c788bee8908b5c81a79a}
}