Scaling of hardware-compatible perturbative training algorithms
arXiv:2501.15403 · doi:10.1063/5.0258271
Abstract
In this work, we explore the capabilities of multiplexed gradient descent (MGD), a scalable and efficient perturbative zeroth-order training method for estimating the gradient of a loss function in hardware and training it via stochastic gradient descent. We extend the framework to include both weight and node perturbation, and discuss the advantages and disadvantages of each approach. We investigate the time to train networks using MGD as a function of network size and task complexity. Previous research has suggested that perturbative training methods do not scale well to large problems, since in these methods the time to estimate the gradient scales linearly with the number of network parameters. However, in this work we show that the time to reach a target accuracy--that is, actually solve the problem of interest--does not follow this undesirable linear scaling, and in fact often decreases with network size. Furthermore, we demonstrate that MGD can be used to calculate a drop-in replacement for the gradient in stochastic gradient descent, and therefore optimization accelerators such as momentum can be used alongside MGD, ensuring compatibility with existing machine learning practices. Our results indicate that MGD can efficiently train large networks on hardware, achieving accuracy comparable to backpropagation, thus presenting a practical solution for future neuromorphic computing systems.
References in corpus (18)
- Adam: A Method for Stochastic Optimization
- Deep physical neural networks enabled by a backpropagation algorithm for arbitrary physical systems
- Experimentally realized in situ backpropagation for deep learning in nanophotonic neural networks
- Eligibility Traces and Plasticity on Behavioral Time Scales: Experimental Support of neoHebbian Three-Factor Learning Rules
- Single chip photonic deep neural network with accelerated training
- Direct Feedback Alignment Provides Learning in Deep Neural Networks
- Gradient learning in spiking neural networks by dynamic perturbation of conductances
- Mixed-precision deep learning based on computational memory
- The Forward-Forward Algorithm: Some Preliminary Investigations
- Assessing the Scalability of Biologically-Motivated Deep Learning Algorithms and Architectures
- Fine-Tuning Language Models with Just Forward Passes
- Machine Learning Without a Processor: Emergent Learning in a Nonlinear Electronic Metamaterial
- Multiplexed gradient descent: Fast online training of modern datasets on hardware neural networks without backpropagation
- Scaling Forward Gradient With Local Losses
- Tensor-Compressed Back-Propagation-Free Training for (Physics-Informed) Neural Networks
- Node Perturbation Can Effectively Train Multi-Layer Neural Networks
- DeepZero: Scaling up Zeroth-Order Optimization for Deep Model Training
- Online Pseudo-Zeroth-Order Training of Neuromorphic Spiking Neural Networks