June 2026Unreviewed
THRD: A Training-Free Multi-Turn Defense Framework for Jailbreak Attacks on Large Language Models
Zhiqing Ma, Zhonghao Xu, Dong Yu, Chen Kang, Changliang Li, Pengyuan Liu
Abstract
Multi-turn jailbreak attacks pose a growing threat to LLMs by exploiting conversational dynamics such as gradual escalation and cross-turn coordination. Existing defenses either rely on costly retraining -- often degrading model utility -- or apply single-turn analysis independently at each turn, failing to capture how risk accumulates along interaction trajectories. We observe that safety behavior in multi-turn interaction is trajectory-dependent: dialogue history continuously reshapes the mode
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0054LLM Jailbreak
Suggested from the entry's categories.
Cite
@misc{ma2026thrd,
title = {{THRD: A Training-Free Multi-Turn Defense Framework for Jailbreak Attacks on Large Language Models}},
author = {Zhiqing Ma and Zhonghao Xu and Dong Yu and Chen Kang and Changliang Li and Pengyuan Liu},
year = {2026},
month = jun,
eprint = {2606.01738},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2606.01738}
}