activity
20172022
most citedStochastic Gradient Descent Optimizes Over-parameterized Deep ReLU Networks

212 citations · 374 across the 13 of their papers we have counts for

collaborators

16 papers

cs.LG20216 cited

Understanding the Generalization of Adam in Learning Neural Networks with Proper Regularization

Difan Zou, Yuan Cao, Yuanzhi Li +1

Adaptive gradient methods such as Adam have gained increasing popularity in deep learning optimization. However, it has been observed that compared with (stochastic) gradient desce…

cs.LG20213 cited

Self-training Converts Weak Learners to Strong Learners in Mixture Models

Spencer Frei, Difan Zou, Zixiang Chen +1

We consider a binary classification problem when the data comes from a mixture of two rotationally symmetric distributions satisfying concentration and anti-concentration propertie…

cs.LG20211 cited

Provable Robustness of Adversarial Training for Learning Halfspaces with Noise

Difan Zou, Spencer Frei, Quanquan Gu

We analyze the properties of adversarial training for learning adversarially robust halfspaces in the presence of agnostic label noise. Denoting as the best ro…

cs.LG2021

Benign Overfitting of Constant-Stepsize SGD for Linear Regression

Difan Zou, Jingfeng Wu, Vladimir Braverman +2

There is an increasing realization that algorithmic inductive biases are central in preventing overfitting; empirically, we often see a benign overfitting phenomenon in overparamet…

cs.LG20201 cited

Direction Matters: On the Implicit Bias of Stochastic Gradient Descent with Moderate Learning Rate

Jingfeng Wu, Difan Zou, Vladimir Braverman +1

Understanding the algorithmic bias of \emph{stochastic gradient descent} (SGD) is one of the key challenges in modern machine learning and deep learning theory. Most of the existin…

cs.LG2020

Faster Convergence of Stochastic Gradient Langevin Dynamics for Non-Log-Concave Sampling

Difan Zou, Pan Xu, Quanquan Gu

We provide a new convergence analysis of stochastic gradient Langevin dynamics (SGLD) for sampling from a class of distributions that can be non-log-concave. At the core of our app…