Skip to content
Search
paperJuly 2026Unreviewed

Mask2Shield: Strengthening LLM Safety against Neuron-Pruning Attacks

Ying JinCheng, Minghui Xu, Yinhao Xiao, Xiuzhen Cheng, Wencheng Yang

Abstract

Large language models (LLMs) are safety-aligned before deployment to reduce harmful content generation. Yet neuron-level pruning attacks show that refusal can depend on a small set of removable units: disabling them can remove safety behavior while leaving much of the model usable. To address this problem, we introduce Mask2Shield (M2S), a masked-forward alignment method that trains a model under this functional pruning. The masked student must recover a safe refusal through the remaining comput

Categories

Cite

@misc{jincheng2026mask2shield,
  title = {{Mask2Shield: Strengthening LLM Safety against Neuron-Pruning Attacks}},
  author = {Ying JinCheng and Minghui Xu and Yinhao Xiao and Xiuzhen Cheng and Wencheng Yang},
  year = {2026},
  month = jul,
  eprint = {2607.23015},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2607.23015}
}