Skip to content
Search
paperJune 2026Unreviewed

Scalable Hierarchical Attention Transformers for Multi-Turn Jailbreak Detection in Long Conversations

Chenhui Hu, Muhammed Salih, Sudipto Guha, Subramanian Srinivasan

Abstract

Multi-turn jailbreaks can evade turn-level moderation by spreading unsafe intent across a dialogue through gradual escalation, reframing, and role manipulation. We address multi-turn jailbreak detection as a conversation-level classification problem and introduce an efficient hierarchical detector that avoids expensive long-context concatenation while retaining cross-turn reasoning. The model encodes individual turns to form compact turn representations and applies a lightweight conversation mod

Categories

Framework mappings

MITRE ATLAS
  • AML.T0054LLM Jailbreak

Suggested from the entry's categories.

Cite

@misc{hu2026scalable,
  title = {{Scalable Hierarchical Attention Transformers for Multi-Turn Jailbreak Detection in Long Conversations}},
  author = {Chenhui Hu and Muhammed Salih and Sudipto Guha and Subramanian Srinivasan},
  year = {2026},
  month = jun,
  eprint = {2606.21082},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2606.21082}
}