May 2026Unreviewed
MT-JailBench: A Modular Benchmark for Understanding Multi-Turn Jailbreak Attacks
Xinkai Zhang, Zhipeng Wei, Huanli Gong, Jing Ting Zheng, Yuchen Zhang, Yue Dong, N. Benjamin Erichson
Abstract
Multi-turn jailbreaks exploit the ability of large language models to accumulate and act on conversational context. Instead of stating a harmful request directly, an attacker can gradually steer the conversation toward an unsafe answer. Recent methods demonstrate this risk, but they are usually evaluated as black-box pipelines with different budgets, judges, retry rules, and strategy generation procedures. As a result, it is often unclear whether reported gains reflect stronger attack mechanisms
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0054LLM Jailbreak
Suggested from the entry's categories.
Cite
@misc{zhang2026mtjailbench,
title = {{MT-JailBench: A Modular Benchmark for Understanding Multi-Turn Jailbreak Attacks}},
author = {Xinkai Zhang and Zhipeng Wei and Huanli Gong and Jing Ting Zheng and Yuchen Zhang and Yue Dong and N. Benjamin Erichson},
year = {2026},
month = may,
eprint = {2605.11002},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2605.11002}
}