1 paper
Qianli Liao, Liu Ziyin, Yulu Gan +3
Over the last four decades, the amazing success of deep learning has been driven by the use of Stochastic Gradient Descent (SGD) as the main optimization technique. The default imp…