Adding Gradient Noise Improves Learning for Very Deep Networks
arXiv:1511.06807
Abstract
Deep feedforward and recurrent networks have achieved impressive results in many perception and language processing applications. This success is partially attributed to architectural innovations such as convolutional and long short-term memory networks. The main motivation for these architectural innovations is that they capture better domain knowledge, and importantly are easier to optimize than more basic architectures. Recently, more complex architectures such as Neural Turing Machines and Memory Networks have been proposed for tasks including question answering and general computation, creating a new set of optimization challenges. In this paper, we discuss a low-overhead and easy-to-implement technique of adding gradient noise which we find to be surprisingly effective when training these very deep architectures. The technique not only helps to avoid overfitting, but also can result in lower training loss. This method alone allows a fully-connected 20-layer deep network to be trained with standard gradient descent, even starting from a poor initialization. We see consistent improvements for many complex models, including a 72% relative reduction in error rate over a carefully-tuned baseline on a challenging question-answering task, and a doubling of the number of accurate binary multiplication models learned across 7,000 random restarts. We encourage further application of this technique to additional complex modern architectures.
References in corpus (7)
Cited by in corpus (112)
- An overview of gradient descent optimization algorithms
- DoReFa-Net: Training Low Bitwidth Convolutional Neural Networks with Low Bitwidth Gradients
- Quantum machine learning: a classical perspective
- Accurate deep neural network inference using computational phase-change memory
- Convolutional Neural Networks using Logarithmic Data Representation
- Shake-Shake regularization
- Over-the-Air Federated Learning from Heterogeneous Data
- Don't Use Large Mini-Batches, Use Local SGD
- Neural Optimizer Search with Reinforcement Learning
- Regularization for Deep Learning: A Taxonomy
- A Comprehensive Study of Deep Bidirectional LSTM RNNs for Acoustic Modeling in Speech Recognition
- An Empirical Model of Large-Batch Training
- Cyclical Stochastic Gradient MCMC for Bayesian Deep Learning
- Review: Deep Learning in Electron Microscopy
- Silicon Photonic Architecture for Training Deep Neural Networks with Direct Feedback Alignment
- TerpreT: A Probabilistic Programming Language for Program Induction
- Noisy Activation Functions
- Deep Neural Network Approximation for Custom Hardware: Where We've Been, Where We're Going
- Differential Privacy Has Disparate Impact on Model Accuracy
- Deep neural networks are robust to weight binarization and other non-linear distortions
- Neural Programmer: Inducing Latent Programs with Gradient Descent
- Training Recurrent Answering Units with Joint Loss Minimization for VQA
- A Hitting Time Analysis of Stochastic Gradient Langevin Dynamics
- Towards an Intelligent Edge: Wireless Communication Meets Machine Learning
- Strong error analysis for stochastic gradient descent optimization algorithms
- CNN Architecture Comparison for Radio Galaxy Classification
- Three Mechanisms of Weight Decay Regularization
- Approximate Gradient Coding via Sparse Random Graphs
- A Little Is Enough: Circumventing Defenses For Distributed Learning
- On the Protection of Private Information in Machine Learning Systems: Two Recent Approaches
- Noisy Differentiable Architecture Search
- Noisy Softmax: Improving the Generalization Ability of DCNN via Postponing the Early Softmax Saturation
- The Impact of the Mini-batch Size on the Variance of Gradients in Stochastic Gradient Descent
- Adaptive Neural Compilation
- Learning to solve the credit assignment problem
- Training Recurrent Neural Networks by Diffusion
- Improving the Neural GPU Architecture for Algorithm Learning
- Word2Bits - Quantized Word Vectors
- ProbAct: A Probabilistic Activation Function for Deep Neural Networks
- GradAug: A New Regularization Method for Deep Neural Networks
- Time Matters in Regularizing Deep Networks: Weight Decay and Data Augmentation Affect Early Learning Dynamics, Matter Little Near Convergence
- A Universally Optimal Multistage Accelerated Stochastic Gradient Method
- Gradient Diversity: a Key Ingredient for Scalable Distributed Learning
- Semi-Supervised Deep Learning for Multi-Tissue Segmentation from Multi-Contrast MRI
- Shape Matters: Understanding the Implicit Bias of the Noise Covariance
- Accelerated Linear Convergence of Stochastic Momentum Methods in Wasserstein Distances
- LEASGD: an Efficient and Privacy-Preserving Decentralized Algorithm for Distributed Learning
- Visual Question Answering with Memory-Augmented Networks
- On the energy landscape of deep networks
- VR-SGD: A Simple Stochastic Variance Reduction Method for Machine Learning
- Lie Access Neural Turing Machine
- F2A2: Flexible Fully-decentralized Approximate Actor-critic for Cooperative Multi-agent Reinforcement Learning
- An Empirical Study of Large-Batch Stochastic Gradient Descent with Structured Covariance Noise
- CPT: Efficient Deep Neural Network Training via Cyclic Precision
- Positive-Negative Momentum: Manipulating Stochastic Gradient Noise to Improve Generalization
- On Stationary-Point Hitting Time and Ergodicity of Stochastic Gradient Langevin Dynamics
- Context-Dependent Acoustic Modeling without Explicit Phone Clustering
- Boosting Binary Masks for Multi-Domain Learning through Affine Transformations
- Data Augmentation for Bayesian Deep Learning
- On the Importance of Consistency in Training Deep Neural Networks
- Comparison of Lattice-Free and Lattice-Based Sequence Discriminative Training Criteria for LVCSR
- Artificial Neural Variability for Deep Learning: On Overfitting, Noise Memorization, and Catastrophic Forgetting
- Distributed Bayesian Learning with Stochastic Natural-gradient Expectation Propagation and the Posterior Server
- Faster Meta Update Strategy for Noise-Robust Deep Learning
- Improving Generalization by Controlling Label-Noise Information in Neural Network Weights
- Scalable Bayesian Learning of Recurrent Neural Networks for Language Modeling
- Local Propagation in Constraint-based Neural Network
- Neural Taylor Approximations: Convergence and Exploration in Rectifier Networks
- Nested Dithered Quantization for Communication Reduction in Distributed Training
- Unsupervised Deep-Learning Based Deformable Image Registration: A Bayesian Framework
- Differentiable Functional Program Interpreters
- Robustness Threats of Differential Privacy
- Energy Confused Adversarial Metric Learning for Zero-Shot Image Retrieval and Clustering
- Towards Consistent Hybrid HMM Acoustic Modeling
- Dynamic Stale Synchronous Parallel Distributed Training for Deep Learning
- What Colour is Neural Noise?
- Reinforced stochastic gradient descent for deep neural network learning
- Analytically Tractable Inference in Deep Neural Networks
- Towards Deep Physical Reservoir Computing Through Automatic Task Decomposition And Mapping
- Modeling the spatio-temporal dynamics of land use change with recurrent neural networks
- Statistically Robust Neural Network Classification
- Stochastic gradient descent with random learning rate
- Learning Neural Network Classifiers with Low Model Complexity
- Differentially Private Dropout
- GDRQ: Group-based Distribution Reshaping for Quantization
- Fine-tuning Handwriting Recognition systems with Temporal Dropout
- Communication-efficient Decentralized Machine Learning over Heterogeneous Networks
- Noise Optimization for Artificial Neural Networks
- Cross-Lingual Dependency Parsing with Late Decoding for Truly Low-Resource Languages
- q-Neurons: Neuron Activations based on Stochastic Jackson's Derivative Operators
- Binary Stochastic Filtering: feature selection and beyond
- Deep Learning in Memristive Nanowire Networks
- PAC-Bayesian Transportation Bound
- Larger is Better: The Effect of Learning Rates Enjoyed by Stochastic Optimization with Progressive Variance Reduction
- Polygonal Unadjusted Langevin Algorithms: Creating stable and efficient adaptive algorithms for neural networks
- Deep Learning for Prostate Pathology
- Enhancement of land-use change modeling using convolutional neural networks and convolutional denoising autoencoders
- Towards Recognizing New Semantic Concepts in New Visual Domains
- Noise Is Useful: Exploiting Data Diversity for Edge Intelligence
- Stabilizing Training of Generative Adversarial Nets via Langevin Stein Variational Gradient Descent
- Introducing Noise in Decentralized Training of Neural Networks
- Text Classification and Clustering with Annealing Soft Nearest Neighbor Loss
- Training Deep Neural Networks by optimizing over nonlocal paths in hyperparameter space
- Adaptive Stochastic Gradient Langevin Dynamics: Taming Convergence and Saddle Point Escape Time
- Diagnostic Visualization for Deep Neural Networks Using Stochastic Gradient Langevin Dynamics
- Discontinuous Constituency Parsing with a Stack-Free Transition System and a Dynamic Oracle
- Semi-Relaxed Quantization with DropBits: Training Low-Bit Neural Networks via Bit-wise Regularization
- Enhance Convolutional Neural Networks with Noise Incentive Block
- XConv: Low-memory stochastic backpropagation for convolutional layers
- Regularization in ResNet with Stochastic Depth
- Binary Search and First Order Gradient Based Method for Stochastic Optimization
- SVGD: A Virtual Gradients Descent Method for Stochastic Optimization