Skip to content
arXiv stat.ML · Papers

Mind the Gap: Structure-Aware Consistency in Preference Learning

arXiv:2604.27733v2 Announce Type: replace-cross Abstract: Aligning Large Language Models (LLMs) with human intent, whether through explicit reward modeling or direct methods such as DPO, fundamentally relies on minimizing a surrogate loss as a proxy for the true pairwise ranking objective. We prove that this reliance i