Skip to content
Search
paperMay 2026Unreviewed

MT-JailBench: A Modular Benchmark for Understanding Multi-Turn Jailbreak Attacks

Xinkai Zhang, Zhipeng Wei, Huanli Gong, Jing Ting Zheng, Yuchen Zhang, Yue Dong, N. Benjamin Erichson

Abstract

Multi-turn jailbreaks exploit the ability of large language models to accumulate and act on conversational context. Instead of stating a harmful request directly, an attacker can gradually steer the conversation toward an unsafe answer. Recent methods demonstrate this risk, but they are usually evaluated as black-box pipelines with different budgets, judges, retry rules, and strategy generation procedures. As a result, it is often unclear whether reported gains reflect stronger attack mechanisms

Categories

Framework mappings

MITRE ATLAS
  • AML.T0054LLM Jailbreak

Suggested from the entry's categories.

Cite

@misc{zhang2026mtjailbench,
  title = {{MT-JailBench: A Modular Benchmark for Understanding Multi-Turn Jailbreak Attacks}},
  author = {Xinkai Zhang and Zhipeng Wei and Huanli Gong and Jing Ting Zheng and Yuchen Zhang and Yue Dong and N. Benjamin Erichson},
  year = {2026},
  month = may,
  eprint = {2605.11002},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2605.11002}
}