Skip to content
Search
paperSeptember 2026Unreviewed

Beyond Shallow Alignment: How Post-Training Methods Determine Refusal Circuits And Steering Robustness

Hoang Cuong Nguyen, Mark Dras, Usman Naseem

Abstract

How do the methods used to train language models to refuse harmful requests shape how that refusal actually works inside the model? We compare three post-training methods - supervised fine-tuning, reasoning-augmented fine-tuning (training on reasoning chains that justify a safety decision), and preference optimization (ORPO) - across three architecturally distinct models (Llama-3.1-8B, Gemma-2-9B, Qwen3-8B). We find that training method, not just data, reshapes how refusal is computed internally

Categories

Cite

@misc{nguyen2026beyond,
  title = {{Beyond Shallow Alignment: How Post-Training Methods Determine Refusal Circuits And Steering Robustness}},
  author = {Hoang Cuong Nguyen and Mark Dras and Usman Naseem},
  year = {2026},
  month = sep,
  eprint = {2609.03887},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2609.03887}
}