Skip to content
arXiv cs.CV · Papers

LC-GRPO: Bridging Train-Inference Gap for Flow-Based GRPO with Langevin Correction

arXiv:2608.05600v1 Announce Type: cross Abstract: Flow-based generative models are typically sampled by solving a deterministic ordinary differential equation (ODE), whereas online reinforcement learning requires stochastic rollouts for policy exploration and optimization. Existing GRPO methods for flow models therefor