Minimax Optimal Early-Stopped Gradient Descent for Gaussian Mixture Classification

By Alex Buna · Paper · stat.ML

In overparameterised classification, training data can be linearly separable even when the underlying distribution is not. In this setting, gradient descent (GD) on the logistic loss diverges in norm while converging in direction to a max-margin interpolating classifier, whose im

Open Models · Stat.ml

View original

HomeResourceLoading…