September 2026Unreviewed
When Safety Routing Breaks: Understanding Alignment Fragility under Benign Fine-Tuning
Yitong Guo, Xiaoyi Chen, Siyuan Zhang, Xiaofeng Wang, Haixu Tang
Abstract
Benign fine-tuning severely weakens the safety alignment of large language models (LLMs), so we study why refusal behavior is so fragile. While prior work often attributes this failure to gradient conflict, we propose a fundamentally different Fisher-geometric explanation: safety Fisher is low-rank, and alignment makes the safety geometry flatter while preserving an output-routing pathway. After 100 benign fine-tuning examples, this pathway is selectively re-sharpened in output-side MLP modules,
Categories
Cite
@misc{guo2026when,
title = {{When Safety Routing Breaks: Understanding Alignment Fragility under Benign Fine-Tuning}},
author = {Yitong Guo and Xiaoyi Chen and Siyuan Zhang and Xiaofeng Wang and Haixu Tang},
year = {2026},
month = sep,
eprint = {2609.01455},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2609.01455}
}