Skip to content
Search
paperJune 2026Unreviewed

The Geometry of Refusal: Linear Instability in Safety-Aligned LLMs

Shivam Ratnakar, Kartikeya Vats

Abstract

Modern Large Language Models (LLMs) rely on extensive safety alignment, yet the mechanistic basis of refusal remains opaque. In this work, we investigate whether safety compliance is a deep semantic decision or a manipulable linear feature. We introduce Contrastive Logit Steering (CLS), a zero-optimization framework that isolates the "refusal direction" by contrasting hidden states derived from safe and unrestricted system prompts. Unlike representation engineering methods that intervene on inte

Categories

Cite

@misc{ratnakar2026geometry,
  title = {{The Geometry of Refusal: Linear Instability in Safety-Aligned LLMs}},
  author = {Shivam Ratnakar and Kartikeya Vats},
  year = {2026},
  month = jun,
  eprint = {2606.22686},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2606.22686}
}