September 2026Unreviewed
SEAL: Reinforcing Global Safety in Mixture-of-Experts through Shared Expert ALignment
Qingyu Meng, Yiwei Zha, Jiahuan Pei, Koen Hindriks, Herbert Bos, Min Chen
Abstract
Mixture-of-Experts (MoE) is a scaling architecture for large language models that activates only a small subset of expert modules per token, enabling massive parameter growth with nearly constant computation. Recent Hybrid MoE architecture adds \textit{shared experts} to capture consistently useful representations, further improving stability and generalization. MoE now powers many flagship open-source and commercial models, yet remains vulnerable to adversarial attacks. Specifically, sparse rou
Categories
Framework mappings
MITRE ATLAS
- AML.T0043Craft Adversarial Data
Suggested from the entry's categories.
Cite
@misc{meng2026seal,
title = {{SEAL: Reinforcing Global Safety in Mixture-of-Experts through Shared Expert ALignment}},
author = {Qingyu Meng and Yiwei Zha and Jiahuan Pei and Koen Hindriks and Herbert Bos and Min Chen},
year = {2026},
month = sep,
eprint = {2609.02293},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2609.02293}
}