← Back to search
paper llmsec-2026-00074
Few-Shot Truly Benign DPO Attack for Jailbreaking LLMs
Sangyeon Yoon, Wonje Jeung, Yoonjun Cho, Dongjae Jeon, Albert No
2026-05
Abstract
Fine-tuning APIs make frontier LLMs easy to customize, but they can also weaken safety alignment during fine-tuning. While prior work shows that benign supervised fine-tuning (SFT) can reduce refusal behavior, deployed fine-tuning pipelines increasingly support preference-based objectives, whose safety risks remain less understood. We show that Direct Preference Optimization (DPO) introduces a stronger and harder-to-audit failure mode. We propose a truly benign DPO attack using only 10 harmless
Categories
Cite This Resource
@article{llmsec202600074,
title = {Few-Shot Truly Benign DPO Attack for Jailbreaking LLMs},
author = {Sangyeon Yoon and Wonje Jeung and Yoonjun Cho and Dongjae Jeon and Albert No},
year = {2026},
url = {https://arxiv.org/abs/2605.10998},
} Metadata
- Added
- 2026-05-17
- Added by
- automation
- Source
- arxiv
- arxiv_id
- 2605.10998