Post-Training with Policy Gradients: Optimality and the Base Model Barrier
arXiv:2603.06957v2 Announce Type: replace Abstract: We study post-training linear autoregressive models with outcome and process rewards. Given a context $boldsymbol{x}$, the model must predict the response…