arXiv cs.CL
· Papers
Wasserstein Distributionally Robust Regret Optimization for Reinforcement Learning from Human Feedback
arXiv:2605.00155v3 Announce Type: replace-cross Abstract: Reinforcement learning from human feedback (RLHF) is a central post-training tool for aligning large language models, but its training reward is only a learned proxy for true human utility. This creates a decision problem under objective misspecification: the po