First-Passage Approach to Optimizing Perturbations for Improved Training of Machine Learning Models
arXiv:2502.04121 · doi:10.1088/2632-2153/add8df
Abstract
Machine learning models have become indispensable tools in applications across the physical sciences. Their training is often time-consuming, vastly exceeding the inference timescales. Several protocols have been developed to perturb the learning process and improve the training, such as shrink and perturb, warm restarts, and stochastic resetting. For classifiers, these perturbations have been shown to result in enhanced speedups or improved generalization. However, the design of such perturbations is usually done ad hoc by intuition and trial and error. To rationally optimize training protocols, we frame them as first-passage processes and consider their response to perturbations. We show that if the unperturbed learning process reaches a quasi-steady state, the response at a single perturbation frequency can predict the behavior at a wide range of frequencies. We employ this approach to a CIFAR-10 classifier using the ResNet-18 model and identify a useful perturbation and frequency among several possibilities. We demonstrate the transferability of the approach to other datasets, architectures, optimizers and even tasks (regression instead of classification). Our work allows optimization of perturbations for improving the training of machine learning models using a first-passage approach.
References in corpus (17)
- Machine learning and the physical sciences
- Fast and Accurate Modeling of Molecular Atomization Energies with Machine Learning
- Solving the Quantum Many-Body Problem with Artificial Neural Networks
- Skillful Precipitation Nowcasting using Deep Generative Models of Radar
- Stochastic Resetting and Applications
- Persistence and First-Passage Properties in Non-equilibrium Systems
- First Passage Under Restart
- A Variational Approach to Enhanced Sampling and Free Energy Calculations
- Optimal stochastic restart renders fluctuations in first passage times universal
- Stochastic Gradient Descent as Approximate Bayesian Inference
- Learning Molecular Dynamics with Simple Language Model built upon Long Short-Term Memory Neural Network
- The role of water and steric constraints in the kinetics of cavity-ligand unbinding
- Stochastic Resetting for Enhanced Sampling
- Machine-learning Iterative Calculation of Entropy for Physical Systems
- Mean-performance of sharp restart I: Statistical roadmap
- Neural Thermodynamic Integration: Free Energies from Energy-based Diffusion Models
- Stochastic Resetting Mitigates Latent Gradient Bias of SGD from Label Noise