4 papers
Implicit Bias of SGD in Multivariate ReLU Networks: Effective Width Collapse
Shuang Liang, Tom Jacobs, Guido Montúfar
We study the implicit bias of noisy stochastic gradient descent in training wide two-layer ReLU networks for multivariate regression. In a mean-field regime, the training dynamics…
Gradient Descent with Large Step Sizes: Chaos and Fractal Convergence Region
Shuang Liang, Guido Montúfar
We examine gradient descent in matrix factorization and show that under large step sizes the parameter space develops a fractal structure. We derive the exact critical step size fo…
Depth-induced NTK: Bridging Over-parameterized Neural Networks and Deep Neural Kernels
Yong-Ming Tian, Shuang Liang, Shao-Qun Zhang +1
While deep learning has achieved remarkable success across a wide range of applications, its theoretical understanding of representation learning remains limited. Deep neural kerne…
Implicit Bias of Mirror Flow for Shallow Neural Networks in Univariate Regression
Shuang Liang, Guido Montúfar
We examine the implicit bias of mirror flow in univariate least squares error regression with wide and shallow neural networks. For a broad class of potential functions, we show th…