Skip to content
Search
paperMay 2026Unreviewed

MultiBreak: A Scalable and Diverse Multi-turn Jailbreak Benchmark for Evaluating LLM Safety

Jialin Song, Xiaodong Liu, Weiwei Yang, Wuyang Chen, Mingqian Feng, Xuekai Zhu, Jianfeng Gao

Abstract

We present MultiBreak, a scalable and diverse multi-turn jailbreak benchmark to evaluate large language model (LLM) safety. Multi-turn jailbreaks mimic natural conversational settings, making them easier to bypass safety-aligned LLM than single-turn jailbreaks. Existing multi-turn benchmarks are limited in size or rely heavily on templates, which restrict their diversity. To address this gap, we unify a wide range of harmful jailbreak intents, and introduce an active learning pipeline for expand

Categories

Framework mappings

MITRE ATLAS
  • AML.T0054LLM Jailbreak

Suggested from the entry's categories.

Cite

@misc{song2026multibreak,
  title = {{MultiBreak: A Scalable and Diverse Multi-turn Jailbreak Benchmark for Evaluating LLM Safety}},
  author = {Jialin Song and Xiaodong Liu and Weiwei Yang and Wuyang Chen and Mingqian Feng and Xuekai Zhu and Jianfeng Gao},
  year = {2026},
  month = may,
  eprint = {2605.01687},
  archivePrefix = {arXiv},
  url = {https://www.semanticscholar.org/paper/531a805a9f8b0646b82ef0181660b9d5967c6199}
}