activity
20162024
most citedStochastic Gradient Descent Optimizes Over-parameterized Deep ReLU Networks

212 citations · 382 across the 21 of their papers we have counts for

collaborators
Showing 2021Show all

6 papers · 1 filter

cs.LG2021

Last Iterate Risk Bounds of SGD with Decaying Stepsize for Overparameterized Linear Regression

Jingfeng Wu, Difan Zou, Vladimir Braverman +2

Stochastic gradient descent (SGD) has been shown to generalize well in many deep learning applications. In practice, one often runs SGD with a geometrically decaying stepsize, i.e.…

cs.LG2021★ 6 cited

Understanding the Generalization of Adam in Learning Neural Networks with Proper Regularization

Difan Zou, Yuan Cao, Yuanzhi Li +1

Adaptive gradient methods such as Adam have gained increasing popularity in deep learning optimization. However, it has been observed that compared with (stochastic) gradient desce…

cs.LG2021

The Benefits of Implicit Regularization from SGD in Least Squares Problems

Difan Zou, Jingfeng Wu, Vladimir Braverman +3

Stochastic gradient descent (SGD) exhibits strong algorithmic regularization effects in practice, which has been hypothesized to play an important role in the generalization of mod…

cs.LG2021★ 3 cited

Self-training Converts Weak Learners to Strong Learners in Mixture Models

Spencer Frei, Difan Zou, Zixiang Chen +1

We consider a binary classification problem when the data comes from a mixture of two rotationally symmetric distributions satisfying concentration and anti-concentration propertie…

cs.LG2021★ 1 cited

Provable Robustness of Adversarial Training for Learning Halfspaces with Noise

Difan Zou, Spencer Frei, Quanquan Gu

We analyze the properties of adversarial training for learning adversarially robust halfspaces in the presence of agnostic label noise. Denoting as the best ro…

cs.LG2021

Benign Overfitting of Constant-Stepsize SGD for Linear Regression

Difan Zou, Jingfeng Wu, Vladimir Braverman +2

There is an increasing realization that algorithmic inductive biases are central in preventing overfitting; empirically, we often see a benign overfitting phenomenon in overparamet…