August 2026Unreviewed
Balancing Privacy, Utility, and Safety in LLM Alignment through Preference Optimization
Di-Shu Yang, Jing-Jing Liu, Ji-Ze Li
Abstract
Preference optimization is widely used to align large language models with human preferences, but preference-data composition may also influence privacy-relevant memorization. We examine whether adding synthetic privacy-preference pairs to Direct Preference Optimization (DPO) is associated with lower canary-based memorization signals without modifying the objective or introducing a formal privacy mechanism. We propose Privacy-Pressure Preference Mixing (P3M), a data-composition protocol that var
Categories
Cite
@misc{yang2026balancing,
title = {{Balancing Privacy, Utility, and Safety in LLM Alignment through Preference Optimization}},
author = {Di-Shu Yang and Jing-Jing Liu and Ji-Ze Li},
year = {2026},
month = aug,
eprint = {2608.30141},
archivePrefix = {arXiv},
url = {https://www.semanticscholar.org/paper/242428093c491f916ba347c255fbd3ae086e8fb5}
}