← Back to search
paper llmsec-2026-00089

RouteHijack: Routing-Aware Attack on Mixture-of-Experts LLMs

Zhiyuan Xu, Joseph Gardiner, Sana Belguith, Lichao Wu

2026-05

Abstract

Safety alignment is critical for the responsible deployment of large language models (LLMs). As Mixture-of-Experts (MoE) architectures are increasingly adopted to scale model capacity, understanding their safety robustness becomes essential. Existing adversarial attacks, however, have notable limitations. Prompt-based jailbreaks rely on heuristic search and transfer poorly, model intervention methods require privileged access to internal representations, and optimization-based input attacks rema

Cite This Resource

@article{llmsec202600089,
  title = {RouteHijack: Routing-Aware Attack on Mixture-of-Experts LLMs},
  author = {Zhiyuan Xu and Joseph Gardiner and Sana Belguith and Lichao Wu},
  year = {2026},
  url = {https://arxiv.org/abs/2605.02946},
}

Metadata

Added
2026-05-17
Added by
automation
Source
arxiv
arxiv_id
2605.02946