arXiv stat.ML
· Papers
Post-Training with Policy Gradients: Optimality and the Base Model Barrier
arXiv:2603.06957v2 Announce Type: replace Abstract: We study post-training linear autoregressive models with outcome and process rewards. Given a context $boldsymbol{x}$, the model must predict the response $boldsymbol{y} in Y^N$, a sequence of length $N$ that satisfies a $gamma$ margin condition, an extension of t