arXiv cs.LG
· Papers
Warp RL: Reshaping Base Policy Distributions for Dynamics Adaptation
arXiv:2606.31043v2 Announce Type: replace Abstract: Residual reinforcement learning adapts a pretrained robot policy by learning an additive correction to its actions. While effective when adaptation amounts to shifting the base policy's action distribution, additive corrections cannot change the distribution's shape,