Skip to content
arXiv cs.LG · Papers

Meta-Learned Reward Shaping for Reinforcement Learning from Human Feedback

arXiv:2607.26094v1 Announce Type: new Abstract: Reinforcement Learning from Human Feedback (RLHF) is the standard approach for aligning large language models with human preferences, but its quality is limited by static, task-agnostic reward models. This mismatch leads to sparse learning signals and suboptimal alignment