arXiv stat.ML
· Papers
Training for the Model You Return: Improving Optimization for Iterate-Averaged Language Models
arXiv:2606.25086v1 Announce Type: cross Abstract: Many modern Language Model (LM) pipelines return an averaged model, such as an exponential moving average of the training iterates, rather than the final iterate itself. This raises a fundamental question: given that we will return an iterate average, how should we chan