Skip to content
arXiv cs.LG · Papers

Minimax Optimal Early-Stopped Gradient Descent for Gaussian Mixture Classification

arXiv:2608.06250v1 Announce Type: cross Abstract: In overparameterised classification, training data can be linearly separable even when the underlying distribution is not. In this setting, gradient descent (GD) on the logistic loss diverges in norm while converging in direction to a max-margin interpolating classifier