arXiv cs.LG
· Papers
Gradient Smoothing: Coupling Layer-wise Updates for Improved Optimization
arXiv:2606.30813v1 Announce Type: new Abstract: Deep neural networks with repeated architectural blocks, such as transformers, often exhibit structured relationships across layers that emerge during training. Motivated by this observation, we introduce emph{Depth-wise Gradient Augmentation}, a general optimization par