Skip to content
Search
paperSeptember 2026Unreviewed

AlignDiff: Exploiting Model-Intrinsic Information for Better Preference Data Selection

Peng Lai, He Zhu, Zhiwen Ruan, Dongdong Zhang, Yun Chen, Peng Li, Furu Wei, Yang Liu, Guanhua Chen

Abstract

Aligning large language models with human preferences remains a challenge, primarily due to the critical role of preference data quality in effective alignment. Existing datasets are frequently plagued by inherent noise and distribution shifts, which inherently limit model performance. To bridge this gap, we propose AlignDiff, a preference data filtering framework driven by intrinsic model signals. AlignDiff first identifies samples with clear preferences using both positive and inverse signals,

Categories

Cite

@misc{lai2026aligndiff,
  title = {{AlignDiff: Exploiting Model-Intrinsic Information for Better Preference Data Selection}},
  author = {Peng Lai and He Zhu and Zhiwen Ruan and Dongdong Zhang and Yun Chen and Peng Li and Furu Wei and Yang Liu and Guanhua Chen},
  year = {2026},
  month = sep,
  eprint = {2609.05899},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2609.05899}
}