September 2026Unreviewed
The Geometry of Refusal: Why Post-Hoc Safety Is Fragile and Pretraining-Time Safety Persists
Srikanth Malla, Chiho Choi, Joon Choi
Abstract
Post-hoc safety training (RLHF, DPO) is the dominant way to align large language models, yet jailbreaks (Zou et al., 2023b), fine-tuning attacks (Qi et al., 2024), and activation-space probes (Arditi et al., 2024) keep recovering the behaviors it was meant to remove. We give this fragility one geometric explanation and trace it to when, during pretraining, safety can take hold. We measure the safety update $\Delta = W_{\text{safe}} - W_{\text{base}}$ against the curvature of the model's capabili
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0054LLM Jailbreak
Suggested from the entry's categories.
Cite
@misc{malla2026geometry,
title = {{The Geometry of Refusal: Why Post-Hoc Safety Is Fragile and Pretraining-Time Safety Persists}},
author = {Srikanth Malla and Chiho Choi and Joon Choi},
year = {2026},
month = sep,
eprint = {2609.06934},
archivePrefix = {arXiv},
url = {https://www.semanticscholar.org/paper/6db11455cdf61ce01b1e966dcdd60768ac4e03d4}
}