activity
20182021
most citedStochastic Gradient Descent Optimizes Over-parameterized Deep ReLU Networks

212 citations · 257 across the 6 of their papers we have counts for

collaborators

13 papers

cs.LG20216 cited

Understanding the Generalization of Adam in Learning Neural Networks with Proper Regularization

Difan Zou, Yuan Cao, Yuanzhi Li +1

Adaptive gradient methods such as Adam have gained increasing popularity in deep learning optimization. However, it has been observed that compared with (stochastic) gradient desce…

cs.LG2021

Provable Generalization of SGD-trained Neural Networks of Any Width in the Presence of Adversarial Label Noise

Spencer Frei, Yuan Cao, Quanquan Gu

We consider a one-hidden-layer leaky ReLU network of arbitrary width trained by stochastic gradient descent (SGD) following an arbitrary initialization. We prove that SGD produces…

cs.LG2020

Agnostic Learning of Halfspaces with Gradient Descent via Soft Margins

Spencer Frei, Yuan Cao, Quanquan Gu

We analyze the properties of gradient descent on convex surrogates for the zero-one loss for the agnostic learning of linear halfspaces. If is the best classificatio…

cs.LG202014 cited

Agnostic Learning of a Single Neuron with Gradient Descent

Spencer Frei, Yuan Cao, Quanquan Gu

We consider the problem of learning the best-fitting single neuron as measured by the expected square loss over some unknown…

cs.LG2020

A Generalized Neural Tangent Kernel Analysis for Two-layer Neural Networks

Zixiang Chen, Yuan Cao, Quanquan Gu +1

A recent breakthrough in deep learning theory shows that the training of over-parameterized deep neural networks can be characterized by a kernel function called \textit{neural tan…

cs.LG2019

Towards Understanding the Spectral Bias of Deep Learning

Yuan Cao, Zhiying Fang, Yue Wu +2

An intriguing phenomenon observed during training neural networks is the spectral bias, which states that neural networks are biased towards learning less complex functions. The pr…