212 citations · 382 across the 21 of their papers we have counts for
6 papers · 1 filter
Last Iterate Risk Bounds of SGD with Decaying Stepsize for Overparameterized Linear Regression
Jingfeng Wu, Difan Zou, Vladimir Braverman +2
Stochastic gradient descent (SGD) has been shown to generalize well in many deep learning applications. In practice, one often runs SGD with a geometrically decaying stepsize, i.e.…
Understanding the Generalization of Adam in Learning Neural Networks with Proper Regularization
Difan Zou, Yuan Cao, Yuanzhi Li +1
Adaptive gradient methods such as Adam have gained increasing popularity in deep learning optimization. However, it has been observed that compared with (stochastic) gradient desce…
The Benefits of Implicit Regularization from SGD in Least Squares Problems
Difan Zou, Jingfeng Wu, Vladimir Braverman +3
Stochastic gradient descent (SGD) exhibits strong algorithmic regularization effects in practice, which has been hypothesized to play an important role in the generalization of mod…
Self-training Converts Weak Learners to Strong Learners in Mixture Models
Spencer Frei, Difan Zou, Zixiang Chen +1
We consider a binary classification problem when the data comes from a mixture of two rotationally symmetric distributions satisfying concentration and anti-concentration propertie…
Provable Robustness of Adversarial Training for Learning Halfspaces with Noise
Difan Zou, Spencer Frei, Quanquan Gu
We analyze the properties of adversarial training for learning adversarially robust halfspaces in the presence of agnostic label noise. Denoting as the best ro…
Benign Overfitting of Constant-Stepsize SGD for Linear Regression
Difan Zou, Jingfeng Wu, Vladimir Braverman +2
There is an increasing realization that algorithmic inductive biases are central in preventing overfitting; empirically, we often see a benign overfitting phenomenon in overparamet…