← Back to search
paper llmsec-2026-00108

MT-JailBench: A Modular Benchmark for Understanding Multi-Turn Jailbreak Attacks

Xinkai Zhang, Zhipeng Wei, Huanli Gong, Jing Ting Zheng, Yuchen Zhang, Yue Dong, N. Benjamin Erichson

2026-05

Abstract

Multi-turn jailbreaks exploit the ability of large language models to accumulate and act on conversational context. Instead of stating a harmful request directly, an attacker can gradually steer the conversation toward an unsafe answer. Recent methods demonstrate this risk, but they are usually evaluated as black-box pipelines with different budgets, judges, retry rules, and strategy generation procedures. As a result, it is often unclear whether reported gains reflect stronger attack mechanisms

Cite This Resource

@article{llmsec202600108,
  title = {MT-JailBench: A Modular Benchmark for Understanding Multi-Turn Jailbreak Attacks},
  author = {Xinkai Zhang and Zhipeng Wei and Huanli Gong and Jing Ting Zheng and Yuchen Zhang and Yue Dong and N. Benjamin Erichson},
  year = {2026},
  url = {https://arxiv.org/abs/2605.11002},
}

Metadata

Added
2026-05-17
Added by
automation
Source
arxiv
arxiv_id
2605.11002