Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
Noise-Driven Exploration and Transient Freezing Select Flat Minima in Stochastic Gradient Descent
Ning Yang, Yikuan Zhang, Qi Ouyang +2
Stochastic gradient descent (SGD) is central to deep learning, yet the dynamical origin of its preference for flatter, more generalizable solutions remains unclear. Here, by analyz…
cs.LG2026
On the Superlinear Relationship between SGD Noise Covariance and Loss Landscape Curvature
Yikuan Zhang, Ning Yang, Yuhai Tu
Stochastic Gradient Descent (SGD) introduces anisotropic noise that is correlated with the local curvature of the loss landscape, thereby biasing optimization toward flat minima. P…
cs.LG2026
Manifold-Constrained Adversarial Training for Long-Tailed Robustness via Geometric Alignment
Guanmeng Xian, Ning Yang, Philip S. Yu
Adversarial training is effective on balanced datasets, but its robustness degrades under longtailed class distributions, where tail classes suffer high robust error and unstable d…