August 2026Unreviewed
No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks
Simiao Xie, Chuancheng Shi, Shangze Li, Wenhua Wu, Fei Shen, Ying Zhou, Zhiyong Wang, Tat-Seng Chua
Abstract
With the rapid release of open-weight large foundation models, safety threats are shifting from black-box jailbreaks to neuron-level white-box attacks that directly identify and manipulate safety-related neurons. Existing alignment methods often investigate the safety behavior on a small number of neurons, creating fragile single point of failure with limited redundancy. To address this issue, we propose distributed safety alignment (DSA), which redundantly encodes safety capabilities across mul
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0054LLM Jailbreak
Suggested from the entry's categories.
Cite
@misc{xie2026no,
title = {{No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks}},
author = {Simiao Xie and Chuancheng Shi and Shangze Li and Wenhua Wu and Fei Shen and Ying Zhou and Zhiyong Wang and Tat-Seng Chua},
year = {2026},
month = aug,
eprint = {2608.01414},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2608.01414}
}