May 2026Unreviewed
RouteHijack: Routing-Aware Attack on Mixture-of-Experts LLMs
Zhiyuan Xu, Joseph Gardiner, Sana Belguith, Lichao Wu
Abstract
Safety alignment is critical for the responsible deployment of large language models (LLMs). As Mixture-of-Experts (MoE) architectures are increasingly adopted to scale model capacity, understanding their safety robustness becomes essential. Existing adversarial attacks, however, have notable limitations. Prompt-based jailbreaks rely on heuristic search and transfer poorly, model intervention methods require privileged access to internal representations, and optimization-based input attacks rema
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0043Craft Adversarial Data
- AML.T0054LLM Jailbreak
Suggested from the entry's categories.
Cite
@misc{xu2026routehijack,
title = {{RouteHijack: Routing-Aware Attack on Mixture-of-Experts LLMs}},
author = {Zhiyuan Xu and Joseph Gardiner and Sana Belguith and Lichao Wu},
year = {2026},
month = may,
eprint = {2605.02946},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2605.02946}
}