Skip to content
arXiv stat.ML · Papers

How to Allocate Your Tokens? Scaling Laws with Training Steps and Batch Size

arXiv:2607.01487v1 Announce Type: cross Abstract: We propose a scaling law that takes into account model size and training data while explicitly splitting the latter into training steps and batch size (called three-term law). Fitting the proposed law on a large set of training runs, we find that it correctly recovers t