7 citations · 7 across the 2 of their papers we have counts for
2 papers
cs.LG2023
Depth Dependence of P Learning Rates in ReLU MLPs
Samy Jelassi, Boris Hanin, Ziwei Ji +3
In this short note we consider random fully connected ReLU networks of width and depth equipped with a mean-field weight initialization. Our purpose is to study the depende…
cs.LG2022★ 7 cited
Towards understanding how momentum improves generalization in deep learning
Samy Jelassi, Yuanzhi Li
Stochastic gradient descent (SGD) with momentum is widely used for training modern deep learning architectures. While it is well-understood that using momentum can lead to faster c…