Skip to content
Search
paperAugust 2026Unreviewed

Balancing Privacy, Utility, and Safety in LLM Alignment through Preference Optimization

Di-Shu Yang, Jing-Jing Liu, Ji-Ze Li

Abstract

Preference optimization is widely used to align large language models with human preferences, but preference-data composition may also influence privacy-relevant memorization. We examine whether adding synthetic privacy-preference pairs to Direct Preference Optimization (DPO) is associated with lower canary-based memorization signals without modifying the objective or introducing a formal privacy mechanism. We propose Privacy-Pressure Preference Mixing (P3M), a data-composition protocol that var

Categories

Cite

@misc{yang2026balancing,
  title = {{Balancing Privacy, Utility, and Safety in LLM Alignment through Preference Optimization}},
  author = {Di-Shu Yang and Jing-Jing Liu and Ji-Ze Li},
  year = {2026},
  month = aug,
  eprint = {2608.30141},
  archivePrefix = {arXiv},
  url = {https://www.semanticscholar.org/paper/242428093c491f916ba347c255fbd3ae086e8fb5}
}