LessWrong AI
· Communities
Tie training can make DPO/RLHF-trained AIs generalize better
This post covers our recent ICML paper: Spurious Correlation Learning in Preference Optimization: Mechanisms, Consequences, and Mitigation via Tie Training.TL;DROur theorems and experiments suggest that DPO and RLHF have an unwelcome consequence: they make AIs care about every feature of actions that correlates with tr