Skip to content
Search
paperAugust 2026Unreviewed

No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks

Simiao Xie, Chuancheng Shi, Shangze Li, Wenhua Wu, Fei Shen, Ying Zhou, Zhiyong Wang, Tat-Seng Chua

Abstract

With the rapid release of open-weight large foundation models, safety threats are shifting from black-box jailbreaks to neuron-level white-box attacks that directly identify and manipulate safety-related neurons. Existing alignment methods often investigate the safety behavior on a small number of neurons, creating fragile single point of failure with limited redundancy. To address this issue, we propose distributed safety alignment (DSA), which redundantly encodes safety capabilities across mul

Categories

Framework mappings

MITRE ATLAS
  • AML.T0054LLM Jailbreak

Suggested from the entry's categories.

Cite

@misc{xie2026no,
  title = {{No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks}},
  author = {Simiao Xie and Chuancheng Shi and Shangze Li and Wenhua Wu and Fei Shen and Ying Zhou and Zhiyong Wang and Tat-Seng Chua},
  year = {2026},
  month = aug,
  eprint = {2608.01414},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2608.01414}
}