Decoupled Neural Interfaces using Synthetic Gradients
arXiv:1608.05343
Abstract
Training directed neural networks typically requires forward-propagating data through a computation graph, followed by backpropagating error signal, to produce weight updates. All layers, or more generally, modules, of the network are therefore locked, in the sense that they must wait for the remainder of the network to execute forwards and propagate error backwards before they can be updated. In this work we break this constraint by decoupling modules by introducing a model of the future computation of the network graph. These models predict what the result of the modelled subgraph will produce using only local information. In particular we focus on modelling error gradients: by using the modelled synthetic gradient in place of true backpropagated error gradients we decouple subgraphs, and can update them independently and asynchronously i.e. we realise decoupled neural interfaces. We show results for feed-forward models, where every layer is trained asynchronously, recurrent neural networks (RNNs) where predicting one's future gradient extends the time over which the RNN can effectively model, and also a hierarchical RNN system with ticking at different timescales. Finally, we demonstrate that in addition to predicting gradients, the same framework can be used to predict inputs, resulting in models which are decoupled in both the forward and backwards pass -- amounting to independent networks which co-learn such that they can be composed into a single functioning corporation.
References in corpus (5)
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
- Direct Feedback Alignment Provides Learning in Deep Neural Networks
- Training Neural Networks Without Gradients: A Scalable ADMM Approach
- How Auto-Encoders Could Provide Credit Assignment in Deep Networks via Target Propagation
- Distributed optimization of deeply nested systems
Cited by in corpus (29)
- Synaptic Plasticity Dynamics for Deep Continuous Local Learning (DECOLLE)
- Var-CNN: A Data-Efficient Website Fingerprinting Attack Based on Deep Learning
- The Reversible Residual Network: Backpropagation Without Storing Activations
- Biologically inspired alternatives to backpropagation through time for learning in recurrent neural nets
- Sobolev Training for Neural Networks
- AMPNet: Asynchronous Model-Parallel Training for Dynamic Neural Networks
- Understanding Synthetic Gradients and Decoupled Neural Interfaces
- Learning to solve the credit assignment problem
- Principled Training of Neural Networks with Direct Feedback Alignment
- Greedy Layerwise Learning Can Scale to ImageNet
- Conducting Credit Assignment by Aligning Local Representations
- TF-Replicator: Distributed Machine Learning for Researchers
- Fully Decoupled Neural Network Learning Using Delayed Gradients
- Learning to Adapt by Minimizing Discrepancy
- Sparse Attentive Backtracking: Long-Range Credit Assignment in Recurrent Networks
- Pipelined Backpropagation at Scale: Training Large Models without Batches
- Embodied Neuromorphic Vision with Event-Driven Random Backpropagation
- Biologically Plausible Training Mechanisms for Self-Supervised Learning in Deep Networks
- DANTE: Deep AlterNations for Training nEural networks
- Large-Scale Gradient-Free Deep Learning with Recursive Local Representation Alignment
- Backprop-Q: Generalized Backpropagation for Stochastic Computation Graphs
- Ghost Units Yield Biologically Plausible Backprop in Deep Neural Networks
- Feed-Forward Optimization With Delayed Feedback for Neural Network Training
- Substitute Teacher Networks: Learning with Almost No Supervision
- Alternating Synthetic and Real Gradients for Neural Language Modeling
- Accumulated Decoupled Learning: Mitigating Gradient Staleness in Inter-Layer Model Parallelization
- Associated Learning: Decomposing End-to-end Backpropagation based on Auto-encoders and Target Propagation
- Long Timescale Credit Assignment in NeuralNetworks with External Memory
- On-Chip Error-triggered Learning of Multi-layer Memristive Spiking Neural Networks