April 2026Unreviewed
Dynamic Adversarial Fine-Tuning Reorganizes Refusal Geometry
Wenhao Lan, Shan Li, Junbin Yang, Haihua Shen, Yijun Yang
Abstract
Safety-aligned language models must refuse harmful requests without collapsing into broad over-refusal, but the training-time mechanisms behind this tradeoff remain unclear. Prior work characterizes refusal directions and jailbreak robustness, yet does not explain how dynamic adversarial fine-tuning changes refusal carriers across training. We present a measurement-driven mechanism study, not a new defense, on one 7B backbone under supervised fine-tuning (SFT) and R2D2-style dynamic adversarial
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0054LLM Jailbreak
Suggested from the entry's categories.
Cite
@misc{lan2026dynamic,
title = {{Dynamic Adversarial Fine-Tuning Reorganizes Refusal Geometry}},
author = {Wenhao Lan and Shan Li and Junbin Yang and Haihua Shen and Yijun Yang},
year = {2026},
month = apr,
eprint = {2604.27019},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2604.27019}
}