LLMs1 min read
AlignDiff: Filtering Preference Data with Model Signals
AlignDiff is a new framework that improves LLM alignment by prioritizing challenging preference samples based on model-intrinsic signals like negative log-likelihood gaps. Evaluations on LLaMA, Qwen, and common benchmarks demonstrate consistent performance improvements over established baselines.
From arXiv cs.CL