Skip to content
arXiv cs.LG · Papers

Warp RL: Reshaping Base Policy Distributions for Dynamics Adaptation

arXiv:2606.31043v2 Announce Type: replace Abstract: Residual reinforcement learning adapts a pretrained robot policy by learning an additive correction to its actions. While effective when adaptation amounts to shifting the base policy's action distribution, additive corrections cannot change the distribution's shape,