Skip to content
Search
paperAugust 2026Unreviewed

NeuronGuard: Robust LLM Safety Alignment via Ablation-Aware Safety Signal Redistribution

Anjun Gao, Yueyang Quan, Yufei Xia, Zhuqing Liu, Minghong Fang

Abstract

Safety alignment in large language models (LLMs) remains brittle against a growing spectrum of attacks. Jailbreak attacks bypass safety mechanisms through crafted prompts, while neuron-level attacks directly prune safety-critical neurons post-deployment. Both exploit a common weakness: safety-relevant information concentrates in a sparse neuron subset. We present NeuronGuard, a fine-tuning-stage defense that simultaneously hardens LLMs against both attack classes by redistributing safety signals

Categories

Framework mappings

MITRE ATLAS
  • AML.T0054LLM Jailbreak

Suggested from the entry's categories.

Cite

@misc{gao2026neuronguard,
  title = {{NeuronGuard: Robust LLM Safety Alignment via Ablation-Aware Safety Signal Redistribution}},
  author = {Anjun Gao and Yueyang Quan and Yufei Xia and Zhuqing Liu and Minghong Fang},
  year = {2026},
  month = aug,
  eprint = {2608.23959},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2608.23959}
}