Skip to content
Search
paperApril 2026Unreviewed

Dynamic Adversarial Fine-Tuning Reorganizes Refusal Geometry

Wenhao Lan, Shan Li, Junbin Yang, Haihua Shen, Yijun Yang

Abstract

Safety-aligned language models must refuse harmful requests without collapsing into broad over-refusal, but the training-time mechanisms behind this tradeoff remain unclear. Prior work characterizes refusal directions and jailbreak robustness, yet does not explain how dynamic adversarial fine-tuning changes refusal carriers across training. We present a measurement-driven mechanism study, not a new defense, on one 7B backbone under supervised fine-tuning (SFT) and R2D2-style dynamic adversarial

Categories

Framework mappings

MITRE ATLAS
  • AML.T0054LLM Jailbreak

Suggested from the entry's categories.

Cite

@misc{lan2026dynamic,
  title = {{Dynamic Adversarial Fine-Tuning Reorganizes Refusal Geometry}},
  author = {Wenhao Lan and Shan Li and Junbin Yang and Haihua Shen and Yijun Yang},
  year = {2026},
  month = apr,
  eprint = {2604.27019},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2604.27019}
}