Skip to content
Search
paperSeptember 2026Unreviewed

How Fragile Is Safety Alignment at Frontier Scale? A Single-Direction Attack on a 320B MoE

Yi Shi, Tanyu Chen, Kai Shen

Abstract

Directional ablation removes an aligned language model's ability to refuse by projecting a single "refusal direction" out of the weights that write the residual stream. It needs no gradient-based training and no optimization, only a few hundred contrastive prompts, which makes it the canonical white-box attack on open-weight alignment. However, it has been established only on dense models up to roughly 70B parameters. We study whether it survives the shift to frontier mixture-of-experts (MoE) mo

Categories

Cite

@misc{shi2026how,
  title = {{How Fragile Is Safety Alignment at Frontier Scale? A Single-Direction Attack on a 320B MoE}},
  author = {Yi Shi and Tanyu Chen and Kai Shen},
  year = {2026},
  month = sep,
  eprint = {2609.09793},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2609.09793}
}