← Back to search
paper llmsec-2026-00108
MT-JailBench: A Modular Benchmark for Understanding Multi-Turn Jailbreak Attacks
Xinkai Zhang, Zhipeng Wei, Huanli Gong, Jing Ting Zheng, Yuchen Zhang, Yue Dong, N. Benjamin Erichson
2026-05
Abstract
Multi-turn jailbreaks exploit the ability of large language models to accumulate and act on conversational context. Instead of stating a harmful request directly, an attacker can gradually steer the conversation toward an unsafe answer. Recent methods demonstrate this risk, but they are usually evaluated as black-box pipelines with different budgets, judges, retry rules, and strategy generation procedures. As a result, it is often unclear whether reported gains reflect stronger attack mechanisms
Categories
Cite This Resource
@article{llmsec202600108,
title = {MT-JailBench: A Modular Benchmark for Understanding Multi-Turn Jailbreak Attacks},
author = {Xinkai Zhang and Zhipeng Wei and Huanli Gong and Jing Ting Zheng and Yuchen Zhang and Yue Dong and N. Benjamin Erichson},
year = {2026},
url = {https://arxiv.org/abs/2605.11002},
} Metadata
- Added
- 2026-05-17
- Added by
- automation
- Source
- arxiv
- arxiv_id
- 2605.11002