7 citations · 7 across the 2 of their papers we have counts for
1 paper · 1 filter
Samy Jelassi, Yuanzhi Li
Stochastic gradient descent (SGD) with momentum is widely used for training modern deep learning architectures. While it is well-understood that using momentum can lead to faster c…