A Theory of Local Learning, the Learning Channel, and the Optimality of Backpropagation
arXiv:1506.06472 · doi:10.1016/j.neunet.2016.07.006
Abstract
In a physical neural system, where storage and processing are intimately intertwined, the rules for adjusting the synaptic weights can only depend on variables that are available locally, such as the activity of the pre- and post-synaptic neurons, resulting in local learning rules. A systematic framework for studying the space of local learning rules is obtained by first specifying the nature of the local variables, and then the functional form that ties them together into each learning rule. Such a framework enables also the systematic discovery of new learning rules and exploration of relationships between learning rules and group symmetries. We study polynomial local learning rules stratified by their degree and analyze their behavior and capabilities in both linear and non-linear units and networks. Stacking local learning rules in deep feedforward networks leads to deep local learning. While deep local learning can learn interesting representations, it cannot learn complex input-output functions, even when targets are available for the top layer. Learning complex input-output functions requires local deep learning where target information is communicated to the deep layers through a backward learning channel. The nature of the communicated information about the targets and the structure of the learning channel partition the space of learning algorithms. We estimate the learning channel capacity associated with several algorithms and show that backpropagation outperforms them by simultaneously maximizing the information rate and minimizing the computational cost, even in recurrent networks. The theory clarifies the concept of Hebbian learning, establishes the power and limitations of local learning rules, introduces the learning channel which enables a formal analysis of the optimality of backpropagation, and explains the sparsity of the space of learning rules discovered so far.
References in corpus (2)
Cited by in corpus (16)
- Surrogate Gradient Learning in Spiking Neural Networks
- Supervised learning in physical networks: From machine learning to learning machines
- Asymmetric Residual Neural Network for Accurate Human Activity Recognition
- Predictive Coding Can Do Exact Backpropagation on Convolutional and Recurrent Neural Networks
- Relaxing the Constraints on Predictive Coding Models
- SpikeGrad: An ANN-equivalent Computation Model for Implementing Backpropagation with Spikes
- Neural Network Gradient Hamiltonian Monte Carlo
- Activation Relaxation: A Local Dynamical Approximation to Backpropagation in the Brain
- Learning in the Machine: Random Backpropagation and the Deep Learning Channel
- Gradient target propagation
- Investigating the Scalability and Biological Plausibility of the Activation Relaxation Algorithm
- Tourbillon: a Physically Plausible Neural Architecture
- Learning in the Machine: the Symmetries of the Deep Learning Channel
- Questions to Guide the Future of Artificial Intelligence Research
- PR Product: A Substitute for Inner Product in Neural Networks
- Few-shot learning using pre-training and shots, enriched by pre-trained samples