Skip to content
Search
paperJune 2025Unreviewed

Advancing Jailbreak Strategies: A Hybrid Approach to Exploiting LLM Vulnerabilities and Bypassing Modern Defenses

Mohamed Ahmed, Mohamed Abdelmouty, Mingyu Kim, Gunvanth Kandula, Alex Park, James C. Davis

arXiv.org

Abstract

The advancement of Pre-Trained Language Models (PTLMs) and Large Language Models (LLMs) has led to their widespread adoption across diverse applications. Despite their success, these models remain vulnerable to attacks that exploit their inherent weaknesses to bypass safety measures. Two primary inference-phase threats are token-level and prompt-level jailbreaks. Token-level attacks embed adversarial sequences that transfer well to black-box models like GPT but leave detectable patterns and rely

Categories

Framework mappings

MITRE ATLAS
  • AML.T0054LLM Jailbreak

Suggested from the entry's categories.

Cite

@misc{ahmed2025advancing,
  title = {{Advancing Jailbreak Strategies: A Hybrid Approach to Exploiting LLM Vulnerabilities and Bypassing Modern Defenses}},
  author = {Mohamed Ahmed and Mohamed Abdelmouty and Mingyu Kim and Gunvanth Kandula and Alex Park and James C. Davis},
  year = {2025},
  month = jun,
  eprint = {2506.21972},
  archivePrefix = {arXiv},
  doi = {10.48550/arXiv.2506.21972},
  url = {https://www.semanticscholar.org/paper/70359e58bc876a54f7b3694f50b4d6e52703124d}
}