arXiv cs.CL
· Papers
Rewarding Better Thinking for LLM Preference Alignment
arXiv:2607.19824v1 Announce Type: cross Abstract: LLM preference alignment aims to optimize models toward human preferences across diverse user instructions. Reinforcement learning has become a major post-training approach for this goal, but existing proxy rewards are often outcome-level, mainly evaluating the final re