Skip to content
Search
paperMay 2026Unreviewed

The Art of the Jailbreak: Formulating Jailbreak Attacks for LLM Security Beyond Binary Scoring

Ismail Hossain, Tanzim Ahad, Md Jahangir Alam, Sai Puppala, Syed Bahauddin Alam, Sajedul Talukder

Abstract

Jailbreak attacks -- adversarial prompts that bypass LLM alignment through purely linguistic manipulation -- pose a growing operational security threat, yet the field lacks large-scale, reproducible infrastructure for generating, categorizing, and evaluating them systematically. This paper addresses that gap with three contributions. (1) Large-scale compositional jailbreak dataset. We construct 114,000 adversarial prompts by applying 912 composing strategies to 125 harmful seed prompts from Jail

Categories

Framework mappings

MITRE ATLAS
  • AML.T0054LLM Jailbreak

Suggested from the entry's categories.

Cite

@misc{hossain2026art,
  title = {{The Art of the Jailbreak: Formulating Jailbreak Attacks for LLM Security Beyond Binary Scoring}},
  author = {Ismail Hossain and Tanzim Ahad and Md Jahangir Alam and Sai Puppala and Syed Bahauddin Alam and Sajedul Talukder},
  year = {2026},
  month = may,
  eprint = {2605.09225},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2605.09225}
}