arXiv stat.ML
· Papers
Mind the Gap: Structure-Aware Consistency in Preference Learning
arXiv:2604.27733v2 Announce Type: replace-cross Abstract: Aligning Large Language Models (LLMs) with human intent, whether through explicit reward modeling or direct methods such as DPO, fundamentally relies on minimizing a surrogate loss as a proxy for the true pairwise ranking objective. We prove that this reliance i