Skip to content
Search
paperAugust 2026Unreviewed

Refusal geometry reflects refusal training: diverse refusal prefixes can raise stable rank and weaken refusal vector ablation attacks

Andrey Labunets

Abstract

Refusal training protects AI models from jailbreaks by training models to decline unsafe queries, reducing the risk of misuse. Recent work finds that refusal behavior in aligned language models can be mediated by a single activation direction or a low-dimensional refusal subspace shared across harmful prompts: ablating those directions suppresses refusals while largely preserves other model capabilities. Yet it remains unclear why safety-critical features in a wide range of models emerge and con

Categories

Framework mappings

MITRE ATLAS
  • AML.T0054LLM Jailbreak

Suggested from the entry's categories.

Cite

@misc{labunets2026refusal,
  title = {{Refusal geometry reflects refusal training: diverse refusal prefixes can raise stable rank and weaken refusal vector ablation attacks}},
  author = {Andrey Labunets},
  year = {2026},
  month = aug,
  eprint = {2608.25390},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2608.25390}
}