Stochastic Resetting Mitigates Latent Gradient Bias of SGD from Label Noise
arXiv:2406.00396 · doi:10.1088/2632-2153/adbc46
Abstract
Giving up and starting over may seem wasteful in many situations such as searching for a target or training deep neural networks (DNNs). Our study, though, demonstrates that resetting from a checkpoint can significantly improve generalization performance when training DNNs with noisy labels. In the presence of noisy labels, DNNs initially learn the general patterns of the data but then gradually memorize the corrupted data, leading to overfitting. By deconstructing the dynamics of stochastic gradient descent (SGD), we identify the behavior of a latent gradient bias induced by noisy labels, which harms generalization. To mitigate this negative effect, we apply the stochastic resetting method to SGD, inspired by recent developments in the field of statistical physics achieving efficient target searches. We first theoretically identify the conditions where resetting becomes beneficial, and then we empirically validate our theory, confirming the significant improvements achieved by resetting. We further demonstrate that our method is both easy to implement and compatible with other methods for handling noisy labels. Additionally, this work offers insights into the learning dynamics of DNNs from an interpretability perspective, expanding the potential to analyze training methods through the lens of statistical physics.
30 pages, 14 figures
References in corpus (17)
- Stochastic Resetting and Applications
- First Passage Under Restart
- Experimental realization of diffusion with stochastic resetting
- Diffusion in a potential landscape with stochastic resetting
- Diffusion with resetting in arbitrary spatial dimension
- Diffusion with resetting in a logarithmic potential
- Stochastic resetting in underdamped Brownian motion
- Stochastic resetting in interacting particle systems: A review
- Random acceleration process under stochastic resetting
- Dynamical Regimes of Diffusion Models
- Optimal non-Markovian search strategies with n-step memory
- Stochastic Resetting for Enhanced Sampling
- Stochastic gradient descent introduces an effective landscape-dependent regularization favoring flat solutions
- Mitigating long queues and waiting times with service resetting
- Sampling with flows, diffusion and autoregressive neural networks: A spin-glass perspective
- Self-trapping self-repelling random walks
- Unexpected advantages of exploitation for target searches in complex networks