4 papers
Explaining Near-Zero Hessian Eigenvalues Through Approximate Symmetries in Neural Networks
Marcel Kühn, Bernd Rosenow
The Hessian of the training loss governs the local geometry of the loss landscape, yet despite existing explanations for its largest eigenvalues, the origin of the vast multitude o…
A Boundary-Layer Mechanism for One-Third Scaling in Online Softmax Classification
Marcel Kühn, Yoon Thelge, Bernd Rosenow
Hard-label classification is usually trained with smooth surrogate losses, most prominently softmax cross-entropy. We isolate an asymptotic mechanism by which this mismatch between…
Continuous Specialization Transition in the Soft Committee Machine with ReLU Activation
Assem Afanah, Bernd Rosenow
We analyze the soft committee machine with Rectified Linear Unit (ReLU) activation by means of the replica method. In a realizable teacher--student setting, we compute the quenched…
Analyzing Neural Scaling Laws in Two-Layer Networks with Power-Law Data Spectra
Roman Worschech, Bernd Rosenow
Neural scaling laws describe how the performance of deep neural networks scales with key factors such as training data size, model complexity, and training time, often following po…