Skip to content
arXiv stat.ML · Papers

Backpropagation-Free Trunk Training via the Split Forward Gradients

arXiv:2607.16612v1 Announce Type: new Abstract: Backpropagation makes training deep networks memory intensive because it must store intermediate activations. Forward-mode methods avoid this cost, but their gradient estimates become increasingly noisy as the number of trained parameters grows. We introduce Split Forward