June 2025Unreviewed
Advancing Jailbreak Strategies: A Hybrid Approach to Exploiting LLM Vulnerabilities and Bypassing Modern Defenses
Mohamed Ahmed, Mohamed Abdelmouty, Mingyu Kim, Gunvanth Kandula, Alex Park, James C. Davis
arXiv.org
Abstract
The advancement of Pre-Trained Language Models (PTLMs) and Large Language Models (LLMs) has led to their widespread adoption across diverse applications. Despite their success, these models remain vulnerable to attacks that exploit their inherent weaknesses to bypass safety measures. Two primary inference-phase threats are token-level and prompt-level jailbreaks. Token-level attacks embed adversarial sequences that transfer well to black-box models like GPT but leave detectable patterns and rely
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0054LLM Jailbreak
Suggested from the entry's categories.
Cite
@misc{ahmed2025advancing,
title = {{Advancing Jailbreak Strategies: A Hybrid Approach to Exploiting LLM Vulnerabilities and Bypassing Modern Defenses}},
author = {Mohamed Ahmed and Mohamed Abdelmouty and Mingyu Kim and Gunvanth Kandula and Alex Park and James C. Davis},
year = {2025},
month = jun,
eprint = {2506.21972},
archivePrefix = {arXiv},
doi = {10.48550/arXiv.2506.21972},
url = {https://www.semanticscholar.org/paper/70359e58bc876a54f7b3694f50b4d6e52703124d}
}