Skip to content
Search
paperJuly 2025Unreviewed

Breaking the Shield: Adversarial Jailbreak Attacks and Defense Mechanisms in Large Language Models

Shreya Dubey, H. Lamkuche

2025 International Conference on Emerging Information Technology and Engineering Solutions (EITES)

Abstract

Large Language Models (LLMs) are increasingly used across diverse applications, raising security concerns, especially regarding jail-break attacks that bypass safety measures. This paper presents a systematic security evaluation of LLMs, exploring the effectiveness of jail-break detection methods, the security progression across model versions, the relationship between model size and vulnerability, and the impact of combined defense strategies. We evaluate both open-source (e.g., LLama, Mistral)

Categories

Framework mappings

MITRE ATLAS
  • AML.T0054LLM Jailbreak

Suggested from the entry's categories.

Cite

@inproceedings{dubey2025breaking,
  title = {{Breaking the Shield: Adversarial Jailbreak Attacks and Defense Mechanisms in Large Language Models}},
  author = {Shreya Dubey and H. Lamkuche},
  year = {2025},
  month = jul,
  booktitle = {2025 International Conference on Emerging Information Technology and Engineering Solutions (EITES)},
  doi = {10.1109/EITES66543.2025.00023},
  url = {https://www.semanticscholar.org/paper/01f493a008c386b93901c788bee8908b5c81a79a}
}