Skip to content
arXiv cs.LG · Papers

On the global convergence of gradient flow for wide shallow models beyond homogeneous nonlinearities

arXiv:2605.10775v2 Announce Type: replace-cross Abstract: A surprising phenomenon in the training of neural networks is the ability of gradient descent to find global minimizers of the training loss despite its non-convexity. Following earlier work, we investigate this behavior for wide shallow models. Existing global