Skip to content
arXiv cs.NE · Papers

Finite Difference Flow Optimization for RL Post-Training of Text-to-Image Models

arXiv:2603.12893v2 Announce Type: replace-cross Abstract: Reinforcement learning (RL) has become a standard technique for post-training diffusion-based image synthesis models, as it enables learning from reward signals to explicitly improve desirable aspects such as image quality and prompt alignment. In this paper, we