Skip to content
r/LocalLLaMA · Communities

Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour | Lit. Review Distributed Training

So, I finished reading this paper: Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour The What? This paper, takes on the problem of scaling experiments done on 1 GPU to like thousands, keeping in mind that the experiments performed on N GPU can be explained by linear extrapolation of the same done on 1 GPU! The