Skip to content
arXiv cs.CL · Papers

Wasserstein Distributionally Robust Regret Optimization for Reinforcement Learning from Human Feedback

arXiv:2605.00155v3 Announce Type: replace-cross Abstract: Reinforcement learning from human feedback (RLHF) is a central post-training tool for aligning large language models, but its training reward is only a learned proxy for true human utility. This creates a decision problem under objective misspecification: the po