Skip to content
arXiv cs.LG · Papers

Overcoming the Communication-Performance Tradeoff in LLM Pretraining

arXiv:2508.15706v3 Announce Type: replace Abstract: Communication-efficient distributed training algorithms (e.g., DiLoCo) have received considerable interest due to their benefits for training large language models (LLMs) in bandwidth-constrained settings, such as across datacenters and over the internet. While these