← Back to search
paper llmsec-2026-00074

Few-Shot Truly Benign DPO Attack for Jailbreaking LLMs

Sangyeon Yoon, Wonje Jeung, Yoonjun Cho, Dongjae Jeon, Albert No

2026-05

Abstract

Fine-tuning APIs make frontier LLMs easy to customize, but they can also weaken safety alignment during fine-tuning. While prior work shows that benign supervised fine-tuning (SFT) can reduce refusal behavior, deployed fine-tuning pipelines increasingly support preference-based objectives, whose safety risks remain less understood. We show that Direct Preference Optimization (DPO) introduces a stronger and harder-to-audit failure mode. We propose a truly benign DPO attack using only 10 harmless

Cite This Resource

@article{llmsec202600074,
  title = {Few-Shot Truly Benign DPO Attack for Jailbreaking LLMs},
  author = {Sangyeon Yoon and Wonje Jeung and Yoonjun Cho and Dongjae Jeon and Albert No},
  year = {2026},
  url = {https://arxiv.org/abs/2605.10998},
}

Metadata

Added
2026-05-17
Added by
automation
Source
arxiv
arxiv_id
2605.10998