July 2026Unreviewed
Mask2Shield: Strengthening LLM Safety against Neuron-Pruning Attacks
Ying JinCheng, Minghui Xu, Yinhao Xiao, Xiuzhen Cheng, Wencheng Yang
Abstract
Large language models (LLMs) are safety-aligned before deployment to reduce harmful content generation. Yet neuron-level pruning attacks show that refusal can depend on a small set of removable units: disabling them can remove safety behavior while leaving much of the model usable. To address this problem, we introduce Mask2Shield (M2S), a masked-forward alignment method that trains a model under this functional pruning. The masked student must recover a safe refusal through the remaining comput
Categories
Cite
@misc{jincheng2026mask2shield,
title = {{Mask2Shield: Strengthening LLM Safety against Neuron-Pruning Attacks}},
author = {Ying JinCheng and Minghui Xu and Yinhao Xiao and Xiuzhen Cheng and Wencheng Yang},
year = {2026},
month = jul,
eprint = {2607.23015},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2607.23015}
}