3 papers
cs.LG2026
Explaining Near-Zero Hessian Eigenvalues Through Approximate Symmetries in Neural Networks
Marcel Kühn, Bernd Rosenow
The Hessian of the training loss governs the local geometry of the loss landscape, yet despite existing explanations for its largest eigenvalues, the origin of the vast multitude o…
cs.LG2026
A Boundary-Layer Mechanism for One-Third Scaling in Online Softmax Classification
Marcel Kühn, Yoon Thelge, Bernd Rosenow
Hard-label classification is usually trained with smooth surrogate losses, most prominently softmax cross-entropy. We isolate an asymptotic mechanism by which this mismatch between…
cs.LG2023
Anti-Correlated Noise in Epoch-Based Stochastic Gradient Descent: Implications for Weight Variances in Flat Directions
Marcel Kühn, Bernd Rosenow
Stochastic Gradient Descent (SGD) has become a cornerstone of neural network optimization due to its computational efficiency and generalization capabilities. However, the gradient…