Skip to content
Search
paperMay 2026Unreviewed

Few-Shot Truly Benign DPO Attack for Jailbreaking LLMs

Sangyeon Yoon, Wonje Jeung, Yoonjun Cho, Dongjae Jeon, Albert No

Abstract

Fine-tuning APIs make frontier LLMs easy to customize, but they can also weaken safety alignment during fine-tuning. While prior work shows that benign supervised fine-tuning (SFT) can reduce refusal behavior, deployed fine-tuning pipelines increasingly support preference-based objectives, whose safety risks remain less understood. We show that Direct Preference Optimization (DPO) introduces a stronger and harder-to-audit failure mode. We propose a truly benign DPO attack using only 10 harmless

Categories

Framework mappings

MITRE ATLAS
  • AML.T0054LLM Jailbreak

Suggested from the entry's categories.

Cite

@misc{yoon2026fewshot,
  title = {{Few-Shot Truly Benign DPO Attack for Jailbreaking LLMs}},
  author = {Sangyeon Yoon and Wonje Jeung and Yoonjun Cho and Dongjae Jeon and Albert No},
  year = {2026},
  month = may,
  eprint = {2605.10998},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2605.10998}
}