Skip to content
Search
paperJuly 2026Unreviewed

HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment

Shei Pern Chua, Fangzhao Wu

Abstract

Understanding how aligned LLMs internally represent safety is critical for diagnosing alignment vulnerabilities, as it explains why jailbreaks succeed and informs the design of robust alignment strategies. Prior work shows that aligned LLMs encode harmfulness and refusal as separable directions in the residual stream at prompt-side token positions. We show that jailbreaks succeed at prompt encoding by suppressing either the refusal or harmfulness direction before any token is generated, with dis

Categories

Framework mappings

MITRE ATLAS
  • AML.T0054LLM Jailbreak

Suggested from the entry's categories.

Cite

@misc{chua2026harc,
  title = {{HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment}},
  author = {Shei Pern Chua and Fangzhao Wu},
  year = {2026},
  month = jul,
  eprint = {2607.00572},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2607.00572}
}