Early-stopped negative-shifted gradient descent escapes fundamental limits of negative ridge regularization by creating adaptive, mixed-sign filters that improve risk by polynomial factors while recovering multiple scales simultaneously.
This paper studies how early stopping in gradient descent with a shifted learning rate can improve regularization in overparameterized linear regression. Unlike standard negative ridge penalties that have structural limits, the proposed approach creates mixed-sign filters that can better handle weak spectral directions, achieving polynomial improvements in risk under certain conditions.