1 paper
Wenjun Xiong, Juan Ding, Xinlei Zuo +1
Stochastic Gradient Descent (SGD) is fundamental for training deep neural networks, especially in non-convex settings. Understanding SGD's generalization properties is crucial for…