Skip to content
Search
paperAugust 2026Unreviewed

HiRoute: Hierarchical Routed Prompt Tuning for Safety Alignment of Large Language Models

Fang-Zhou Chen, Shiji Zhao, Mengyan Wang, Qihui Zhu, Ran-Jie Duan, Maoxun Yuan, Xingxing Wei

Abstract

Large language models (LLMs) remain vulnerable to harmful requests and jailbreak attacks. Parameter-efficient safety alignment methods based on prompt tuning typically rely on a single global prompt or externally selected prompt modules. Such static designs struggle to maintain a cross-category safety boundary while generating constructive responses tailored to specific risks and avoiding over-refusal of benign inputs. To address these limitations, we propose HiRoute, an input-adaptive hierarchi

Categories

Framework mappings

MITRE ATLAS
  • AML.T0054LLM Jailbreak

Suggested from the entry's categories.

Cite

@misc{chen2026hiroute,
  title = {{HiRoute: Hierarchical Routed Prompt Tuning for Safety Alignment of Large Language Models}},
  author = {Fang-Zhou Chen and Shiji Zhao and Mengyan Wang and Qihui Zhu and Ran-Jie Duan and Maoxun Yuan and Xingxing Wei},
  year = {2026},
  month = aug,
  eprint = {2608.12821},
  archivePrefix = {arXiv},
  url = {https://www.semanticscholar.org/paper/d697618a11df5206c8d73ecedf55ddf9db224a37}
}