arXiv cs.LG
· Papers
Minimax Optimal Early-Stopped Gradient Descent for Gaussian Mixture Classification
arXiv:2608.06250v1 Announce Type: cross Abstract: In overparameterised classification, training data can be linearly separable even when the underlying distribution is not. In this setting, gradient descent (GD) on the logistic loss diverges in norm while converging in direction to a max-margin interpolating classifier