arXiv stat.ML
· Papers
Dropout and Random Gradient Masking Are Asymptotically Equivalent in Large ResNets
arXiv:2607.16761v1 Announce Type: new Abstract: Dropout and Random Gradient Masking (RaM) are two training techniques used to improve performance in deep learning. Both techniques inject randomness into the training dynamics, but in significantly different ways: dropout applies random masks to the activations in the fo