$mu$pscaling small models: Principled warm starts and hyperparameter transfer
arXiv:2602.10545v2 Announce Type: replace Abstract: Modern large-scale neural networks are often trained and released in multiple sizes to accommodate diverse inference budgets. To improve efficiency, recent…