Backpropagation-Free Trunk Training via the Split Forward Gradients
arXiv:2607.16612v1 Announce Type: new Abstract: Backpropagation makes training deep networks memory intensive because it must store intermediate activations. Forward-mode methods avoid this cost, but their gradient…