arXiv cs.LG
· Papers
Overcoming the Communication-Performance Tradeoff in LLM Pretraining
arXiv:2508.15706v3 Announce Type: replace Abstract: Communication-efficient distributed training algorithms (e.g., DiLoCo) have received considerable interest due to their benefits for training large language models (LLMs) in bandwidth-constrained settings, such as across datacenters and over the internet. While these