Skip to content
Search
paperSeptember 2026Unreviewed

From Detection to Refusal: Safer LLMs via Circuit-Guided Weight Scaling

Kuan-Lin Chu, Chung-En Sun, Tsui-Wei Weng

Abstract

Despite extensive alignment efforts, Large Language Models (LLMs) remain vulnerable to generating unsafe content under adversarial prompting, yet the internal mechanisms by which safety behaviors are implemented remain poorly understood. We study LLM safety from a mechanistic interpretability perspective and characterize a multi-stage *safety circuit* that organizes refusal behavior, consisting of (i) $\textbf{Harmful Detection Heads}$ that respond to harmful inputs, (ii) $\textbf{Safety Neurons

Categories

Cite

@misc{chu2026fromb,
  title = {{From Detection to Refusal: Safer LLMs via Circuit-Guided Weight Scaling}},
  author = {Kuan-Lin Chu and Chung-En Sun and Tsui-Wei Weng},
  year = {2026},
  month = sep,
  eprint = {2609.00051},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2609.00051}
}