Skip to content
Search
paperSeptember 2026Unreviewed

The Geometry of Refusal: Why Post-Hoc Safety Is Fragile and Pretraining-Time Safety Persists

Srikanth Malla, Chiho Choi, Joon Choi

Abstract

Post-hoc safety training (RLHF, DPO) is the dominant way to align large language models, yet jailbreaks (Zou et al., 2023b), fine-tuning attacks (Qi et al., 2024), and activation-space probes (Arditi et al., 2024) keep recovering the behaviors it was meant to remove. We give this fragility one geometric explanation and trace it to when, during pretraining, safety can take hold. We measure the safety update $\Delta = W_{\text{safe}} - W_{\text{base}}$ against the curvature of the model's capabili

Categories

Framework mappings

MITRE ATLAS
  • AML.T0054LLM Jailbreak

Suggested from the entry's categories.

Cite

@misc{malla2026geometry,
  title = {{The Geometry of Refusal: Why Post-Hoc Safety Is Fragile and Pretraining-Time Safety Persists}},
  author = {Srikanth Malla and Chiho Choi and Joon Choi},
  year = {2026},
  month = sep,
  eprint = {2609.06934},
  archivePrefix = {arXiv},
  url = {https://www.semanticscholar.org/paper/6db11455cdf61ce01b1e966dcdd60768ac4e03d4}
}