Skip to content
arXiv cs.CL · Papers

Rewarding Better Thinking for LLM Preference Alignment

arXiv:2607.19824v1 Announce Type: cross Abstract: LLM preference alignment aims to optimize models toward human preferences across diverse user instructions. Reinforcement learning has become a major post-training approach for this goal, but existing proxy rewards are often outcome-level, mainly evaluating the final re