Skip to content
arXiv cs.LG · Papers

A convergence result of a continuous model of deep learning via a L{}ojasiewicz–Simon inequality

arXiv:2311.15365v3 Announce Type: replace Abstract: We study an idealized training process for deep neural networks in a continuous-depth, mean-field model in which each layer is parameterized by a probability measure on a Euclidean parameter space. The training dynamics are formulated as a Wasserstein-type gradient fl