Skip to content
arXiv cs.CL · Papers

Preference Tuning as Spectral Update Reorganization

arXiv:2607.20438v1 Announce Type: new Abstract: Preference-based post-training is usually understood through endpoint behavior, yet the learned update that produces this behavior remains largely opaque. We study RLHF and related preference optimization through the spectral structure of their induced parameter updates.