1 paper
Zhipeng Yao, Rui Yu, Guisong Chang +3
Stochastic Gradient Descent (SGD) and its momentum variants form the backbone of deep learning optimization, yet the underlying dynamics of their gradient behavior remain insuffici…