Skip to content
Search
paperJune 2026Unreviewed

THRD: A Training-Free Multi-Turn Defense Framework for Jailbreak Attacks on Large Language Models

Zhiqing Ma, Zhonghao Xu, Dong Yu, Chen Kang, Changliang Li, Pengyuan Liu

Abstract

Multi-turn jailbreak attacks pose a growing threat to LLMs by exploiting conversational dynamics such as gradual escalation and cross-turn coordination. Existing defenses either rely on costly retraining -- often degrading model utility -- or apply single-turn analysis independently at each turn, failing to capture how risk accumulates along interaction trajectories. We observe that safety behavior in multi-turn interaction is trajectory-dependent: dialogue history continuously reshapes the mode

Categories

Framework mappings

MITRE ATLAS
  • AML.T0054LLM Jailbreak

Suggested from the entry's categories.

Cite

@misc{ma2026thrd,
  title = {{THRD: A Training-Free Multi-Turn Defense Framework for Jailbreak Attacks on Large Language Models}},
  author = {Zhiqing Ma and Zhonghao Xu and Dong Yu and Chen Kang and Changliang Li and Pengyuan Liu},
  year = {2026},
  month = jun,
  eprint = {2606.01738},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2606.01738}
}