212 citations · 374 across the 13 of their papers we have counts for
16 papers
Understanding the Generalization of Adam in Learning Neural Networks with Proper Regularization
Difan Zou, Yuan Cao, Yuanzhi Li +1
Adaptive gradient methods such as Adam have gained increasing popularity in deep learning optimization. However, it has been observed that compared with (stochastic) gradient desce…
Self-training Converts Weak Learners to Strong Learners in Mixture Models
Spencer Frei, Difan Zou, Zixiang Chen +1
We consider a binary classification problem when the data comes from a mixture of two rotationally symmetric distributions satisfying concentration and anti-concentration propertie…
Provable Robustness of Adversarial Training for Learning Halfspaces with Noise
Difan Zou, Spencer Frei, Quanquan Gu
We analyze the properties of adversarial training for learning adversarially robust halfspaces in the presence of agnostic label noise. Denoting as the best ro…
Benign Overfitting of Constant-Stepsize SGD for Linear Regression
Difan Zou, Jingfeng Wu, Vladimir Braverman +2
There is an increasing realization that algorithmic inductive biases are central in preventing overfitting; empirically, we often see a benign overfitting phenomenon in overparamet…
Direction Matters: On the Implicit Bias of Stochastic Gradient Descent with Moderate Learning Rate
Jingfeng Wu, Difan Zou, Vladimir Braverman +1
Understanding the algorithmic bias of \emph{stochastic gradient descent} (SGD) is one of the key challenges in modern machine learning and deep learning theory. Most of the existin…
Faster Convergence of Stochastic Gradient Langevin Dynamics for Non-Log-Concave Sampling
Difan Zou, Pan Xu, Quanquan Gu
We provide a new convergence analysis of stochastic gradient Langevin dynamics (SGLD) for sampling from a class of distributions that can be non-log-concave. At the core of our app…