Skip to content
Search
paperSeptember 2026Unreviewed

SEAL: Reinforcing Global Safety in Mixture-of-Experts through Shared Expert ALignment

Qingyu Meng, Yiwei Zha, Jiahuan Pei, Koen Hindriks, Herbert Bos, Min Chen

Abstract

Mixture-of-Experts (MoE) is a scaling architecture for large language models that activates only a small subset of expert modules per token, enabling massive parameter growth with nearly constant computation. Recent Hybrid MoE architecture adds \textit{shared experts} to capture consistently useful representations, further improving stability and generalization. MoE now powers many flagship open-source and commercial models, yet remains vulnerable to adversarial attacks. Specifically, sparse rou

Categories

Framework mappings

MITRE ATLAS
  • AML.T0043Craft Adversarial Data

Suggested from the entry's categories.

Cite

@misc{meng2026seal,
  title = {{SEAL: Reinforcing Global Safety in Mixture-of-Experts through Shared Expert ALignment}},
  author = {Qingyu Meng and Yiwei Zha and Jiahuan Pei and Koen Hindriks and Herbert Bos and Min Chen},
  year = {2026},
  month = sep,
  eprint = {2609.02293},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2609.02293}
}