How Important is Weight Symmetry in Backpropagation?
arXiv:1510.05067
Abstract
Gradient backpropagation (BP) requires symmetric feedforward and feedback connections -- the same weights must be used for forward and backward passes. This "weight transport problem" (Grossberg 1987) is thought to be one of the main reasons to doubt BP's biologically plausibility. Using 15 different classification datasets, we systematically investigate to what extent BP really depends on weight symmetry. In a study that turned out to be surprisingly similar in spirit to Lillicrap et al.'s demonstration (Lillicrap et al. 2014) but orthogonal in its results, our experiments indicate that: (1) the magnitudes of feedback weights do not matter to performance (2) the signs of feedback weights do matter -- the more concordant signs between feedforward and their corresponding feedback connections, the better (3) with feedback weights having random magnitudes and 100% concordant signs, we were able to achieve the same or even better performance than SGD. (4) some normalizations/stabilizations are indispensable for such asymmetric BP to work, namely Batch Normalization (BN) (Ioffe and Szegedy 2015) and/or a "Batch Manhattan" (BM) update rule.
References in corpus (6)
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
- MatConvNet - Convolutional Neural Networks for MATLAB
- Towards Biologically Plausible Deep Learning
- How Auto-Encoders Could Provide Credit Assignment in Deep Networks via Target Propagation
- Neural Turing Machines
- Random feedback weights support learning in deep neural networks
Cited by in corpus (25)
- Bridging the Gaps Between Residual Learning, Recurrent Neural Networks and Visual Cortex
- Learning without feedback: Fixed random learning signals allow for feedforward training of deep neural networks
- Biologically-plausible learning algorithms can scale to large datasets
- A Theoretical Framework for Target Propagation
- Principled Training of Neural Networks with Direct Feedback Alignment
- Two Routes to Scalable Credit Assignment without Weight Symmetry
- Streaming Normalization: Towards Simpler and More Biologically-plausible Normalizations for Online and Recurrent Learning
- Conducting Credit Assignment by Aligning Local Representations
- Invariant recognition drives neural representations of action sequences
- Learning to Adapt by Minimizing Discrepancy
- Learning in the Machine: Random Backpropagation and the Deep Learning Channel
- Biologically Plausible Training Mechanisms for Self-Supervised Learning in Deep Networks
- Identifying Learning Rules From Neural Network Observables
- Large-Scale Gradient-Free Deep Learning with Recursive Local Representation Alignment
- Differentially Private Deep Learning with Direct Feedback Alignment
- Direct Feedback Alignment Scales to Modern Deep Learning Tasks and Architectures
- Dictionary Learning by Dynamical Neural Networks
- Credit Assignment Through Broadcasting a Global Error Vector
- Benchmarking the Accuracy and Robustness of Feedback Alignment Algorithms
- Questions to Guide the Future of Artificial Intelligence Research
- Tourbillon: a Physically Plausible Neural Architecture
- MAP Propagation Algorithm: Faster Learning with a Team of Reinforcement Learning Agents
- Training DNNs in O(1) memory with MEM-DFA using Random Matrices
- Deep learning with asymmetric connections and Hebbian updates
- Efficient Training Convolutional Neural Networks on Edge Devices with Gradient-pruned Sign-symmetric Feedback Alignment