Skip to content
arXiv cs.LG · Papers

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning?

arXiv:2607.27203v2 Announce Type: replace Abstract: Pre-training followed by fine-tuning has become the dominant recipe for learning performant policies, and in value-based reinforcement learning (RL) this raises a natural question: given a pretrained policy, should the Q-function be pretrained on offline data too? Con