Learning may need only a few bits of synaptic precision
arXiv:1602.04129 · doi:10.1103/PhysRevE.93.052313
Abstract
Learning in neural networks poses peculiar challenges when using discretized rather then continuous synaptic states. The choice of discrete synapses is motivated by biological reasoning and experiments, and possibly by hardware implementation considerations as well. In this paper we extend a previous large deviations analysis which unveiled the existence of peculiar dense regions in the space of synaptic states which accounts for the possibility of learning efficiently in networks with binary synapses. We extend the analysis to synapses with multiple states and generally more plausible biological features. The results clearly indicate that the overall qualitative picture is unchanged with respect to the binary case, and very robust to variation of the details of the model. We also provide quantitative results which suggest that the advantages of increasing the synaptic precision (i.e.~the number of internal synaptic states) rapidly vanish after the first few bits, and therefore that, for practical applications, only few bits may be needed for near-optimal performance, consistently with recent biological findings. Finally, we demonstrate how the theoretical analysis can be exploited to design efficient algorithmic search strategies.
38 pages (main text: 16 pages), 5 figures; http://link.aps.org/doi/10.1103/PhysRevE.93.052313
References in corpus (5)
- Gibbs States and the Set of Solutions of Random Constraint Satisfaction Problems
- Efficient supervised learning in networks with binary synapses
- Origin of the computational hardness for learning with binary synapses
- Local entropy as a measure for sampling solutions in Constraint Satisfaction Problems
- Generalization learning in a perceptron with binary synapses
Cited by in corpus (14)
- Unreasonable Effectiveness of Learning Neural Networks: From Accessible States and Robust Ensembles to Basic Algorithmic Schemes
- Entropy-SGD: Biasing Gradient Descent Into Wide Valleys
- A neuromorphic systems approach to in-memory computing with non-ideal memristive devices: From mitigation to exploitation
- Shaping the learning landscape in neural networks around wide flat minima
- Efficiency of quantum versus classical annealing in non-convex learning problems
- An Optimal Control Approach to Deep Learning and Applications to Discrete-Weight Neural Networks
- On the role of synaptic stochasticity in training low-precision neural networks
- Parle: parallelizing stochastic gradient descent
- Deep learning via message passing algorithms based on belief propagation
- Vector Symbolic Finite State Machines in Attractor Neural Networks
- A characterization of the Edge of Criticality in Binary Echo State Networks
- Understanding the computational difficulty of a binary-weight perceptron and the advantage of input sparseness
- Solvable Model for Inheriting the Regularization through Knowledge Distillation
- SGB: Stochastic Gradient Bound Method for Optimizing Partition Functions