Skip to content
Search
paperAugust 2026Unreviewed

Certified Multi-Turn Robustness for LLM Safety via Compositional Bounds and Safety Persistence

Yang Liu, Bin Chong, Wenkai Yang, Shuai Zhang, Yan-Cheng Chen, Fei Han, GuoZhen, Cheng Zhang, Huaibing Xie, Changze Lv, Shihan Dou, Pluto Zhou

Abstract

Large language models (LLMs) are vulnerable to multi-turn jailbreak attacks that progressively manipulate conversation context. Existing certified robustness methods are limited to single-turn inputs; naive multi-turn composition yields bounds that degrade exponentially in the number of turns. We introduce Multi-Turn Certified Robustness (MTCR), a framework that models conversational safety via State-Adversarial MDPs and defines $k$-turn certified robustness as the worst-case safety probability

Categories

Framework mappings

MITRE ATLAS
  • AML.T0054LLM Jailbreak

Suggested from the entry's categories.

Cite

@misc{liu2026certified,
  title = {{Certified Multi-Turn Robustness for LLM Safety via Compositional Bounds and Safety Persistence}},
  author = {Yang Liu and Bin Chong and Wenkai Yang and Shuai Zhang and Yan-Cheng Chen and Fei Han and GuoZhen and Cheng Zhang and Huaibing Xie and Changze Lv and Shihan Dou and Pluto Zhou},
  year = {2026},
  month = aug,
  eprint = {2608.20820},
  archivePrefix = {arXiv},
  url = {https://www.semanticscholar.org/paper/bffff5b3e618d29747dab5b25337420a0ec519b2}
}