Subaging in underparametrized Deep Neural Networks
arXiv:2209.02517 · doi:10.1088/2632-2153/ac8f1b
Abstract
We consider a simple classification problem to show that the dynamics of finite-width Deep Neural Networks in the underparametrized regime gives rise to effects similar to those associated with glassy systems, namely a slow evolution of the loss function and aging. Remarkably, the aging is sublinear in the waiting time (subaging) and the power-law exponent characterizing it is robust to different architectures under the constraint of a constant total number of parameters. Our results are maintained in the more complex scenario of the MNIST database. We find that for this database there is a unique exponent ruling the subaging behavior in the whole phase.
16 pages, 10 figures. Manuscript accepted in"Machine Learning: Science and Technology"
References in corpus (5)
- The Loss Surfaces of Multilayer Networks
- Rejuvenation and Memory Effects in a Structural Glass
- Critical scaling and aging in cooling systems near the jamming transition
- How noise affects the Hessian spectrum in overparameterized neural networks
- Phases of learning dynamics in artificial neural networks: with or without mislabeled data