September 2026Unreviewed
From Detection to Refusal: Safer LLMs via Circuit-Guided Weight Scaling
Kuan-Lin Chu, Chung-En Sun, Tsui-Wei Weng
Abstract
Despite extensive alignment efforts, Large Language Models (LLMs) remain vulnerable to generating unsafe content under adversarial prompting, yet the internal mechanisms by which safety behaviors are implemented remain poorly understood. We study LLM safety from a mechanistic interpretability perspective and characterize a multi-stage *safety circuit* that organizes refusal behavior, consisting of (i) $\textbf{Harmful Detection Heads}$ that respond to harmful inputs, (ii) $\textbf{Safety Neurons
Categories
Cite
@misc{chu2026fromb,
title = {{From Detection to Refusal: Safer LLMs via Circuit-Guided Weight Scaling}},
author = {Kuan-Lin Chu and Chung-En Sun and Tsui-Wei Weng},
year = {2026},
month = sep,
eprint = {2609.00051},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2609.00051}
}