September 2026Unreviewed
How Fragile Is Safety Alignment at Frontier Scale? A Single-Direction Attack on a 320B MoE
Yi Shi, Tanyu Chen, Kai Shen
Abstract
Directional ablation removes an aligned language model's ability to refuse by projecting a single "refusal direction" out of the weights that write the residual stream. It needs no gradient-based training and no optimization, only a few hundred contrastive prompts, which makes it the canonical white-box attack on open-weight alignment. However, it has been established only on dense models up to roughly 70B parameters. We study whether it survives the shift to frontier mixture-of-experts (MoE) mo
Categories
Cite
@misc{shi2026how,
title = {{How Fragile Is Safety Alignment at Frontier Scale? A Single-Direction Attack on a 320B MoE}},
author = {Yi Shi and Tanyu Chen and Kai Shen},
year = {2026},
month = sep,
eprint = {2609.09793},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2609.09793}
}