Skip to content
Search
paperSeptember 2026Unreviewed

When Safety Routing Breaks: Understanding Alignment Fragility under Benign Fine-Tuning

Yitong Guo, Xiaoyi Chen, Siyuan Zhang, Xiaofeng Wang, Haixu Tang

Abstract

Benign fine-tuning severely weakens the safety alignment of large language models (LLMs), so we study why refusal behavior is so fragile. While prior work often attributes this failure to gradient conflict, we propose a fundamentally different Fisher-geometric explanation: safety Fisher is low-rank, and alignment makes the safety geometry flatter while preserving an output-routing pathway. After 100 benign fine-tuning examples, this pathway is selectively re-sharpened in output-side MLP modules,

Categories

Cite

@misc{guo2026when,
  title = {{When Safety Routing Breaks: Understanding Alignment Fragility under Benign Fine-Tuning}},
  author = {Yitong Guo and Xiaoyi Chen and Siyuan Zhang and Xiaofeng Wang and Haixu Tang},
  year = {2026},
  month = sep,
  eprint = {2609.01455},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2609.01455}
}