On the role of synaptic stochasticity in training low-precision neural networks
arXiv:1710.09825 · doi:10.1103/PhysRevLett.120.268103
Abstract
Stochasticity and limited precision of synaptic weights in neural network models are key aspects of both biological and hardware modeling of learning processes. Here we show that a neural network model with stochastic binary weights naturally gives prominence to exponentially rare dense regions of solutions with a number of desirable properties such as robustness and good generalization performance, while typical solutions are isolated and hard to find. Binary solutions of the standard perceptron problem are obtained from a simple gradient descent procedure on a set of real values parametrizing a probability distribution over the binary synapses. Both analytical and numerical results are presented. An algorithmic extension aimed at training discrete deep neural networks is also investigated.
7 pages + 14 pages of supplementary material
References in corpus (7)
- Proceedings of the 29th International Conference on Machine Learning (ICML-12)
- Probabilistic Backpropagation for Scalable Learning of Bayesian Neural Networks
- Efficient supervised learning in networks with binary synapses
- Origin of the computational hardness for learning with binary synapses
- Efficiency of quantum versus classical annealing in non-convex learning problems
- Constraint satisfaction problems with isolated solutions are hard
- Generalization learning in a perceptron with binary synapses
Cited by in corpus (13)
- Shaping the learning landscape in neural networks around wide flat minima
- Mean-field inference methods for neural networks
- HEMP: High-order Entropy Minimization for neural network comPression
- Choose your tools carefully: A Comparative Evaluation of Deterministic vs. Stochastic and Binary vs. Analog Neuron models for Implementing Emerging Computing Paradigms
- Statistical mechanics of continual learning: variational principle and mean-field potential
- Variational mean-field theory for training restricted Boltzmann machines with binary synapses
- Critical initialisation in continuous approximations of binary neural networks
- Equivalence between algorithmic instability and transition to replica symmetry breaking in perceptron learning systems
- Understanding the computational difficulty of a binary-weight perceptron and the advantage of input sparseness
- Recognition Capabilities of a Hopfield Model with Auxiliary Hidden Neurons
- Solvable Model for Inheriting the Regularization through Knowledge Distillation
- Training Restricted Boltzmann Machines with Binary Synapses using the Bayesian Learning Rule
- Deep Networks on Toroids: Removing Symmetries Reveals the Structure of Flat Regions in the Landscape Geometry