Training Recurrent Neural Networks by Diffusion
arXiv:1601.04114
Abstract
This work presents a new algorithm for training recurrent neural networks (although ideas are applicable to feedforward networks as well). The algorithm is derived from a theory in nonconvex optimization related to the diffusion equation. The contributions made in this work are two fold. First, we show how some seemingly disconnected mechanisms used in deep learning such as smart initialization, annealed learning rate, layerwise pretraining, and noise injection (as done in dropout and SGD) arise naturally and automatically from this framework, without manually crafting them into the algorithms. Second, we present some preliminary results on comparing the proposed method against SGD. It turns out that the new algorithm can achieve similar level of generalization accuracy of SGD in much fewer number of epochs.
References in corpus (7)
- Sequence to Sequence Learning with Neural Networks
- Improving neural networks by preventing co-adaptation of feature detectors
- Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
- Learning with Pseudo-Ensembles
- Adding Gradient Noise Improves Learning for Very Deep Networks
- On the saddle point problem for non-convex optimization
- Closed Form for Some Gaussian Convolutions
Cited by in corpus (16)
- On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima
- Noisy Networks for Exploration
- Empirical Analysis of the Hessian of Over-Parametrized Neural Networks
- Sharpness-Aware Minimization for Efficiently Improving Generalization
- Entropy-SGD: Biasing Gradient Descent Into Wide Valleys
- Noisy Activation Functions
- Kickstarting Deep Reinforcement Learning
- Improving Local Identifiability in Probabilistic Box Embeddings
- ASAM: Adaptive Sharpness-Aware Minimization for Scale-Invariant Learning of Deep Neural Networks
- On the energy landscape of deep networks
- Approximating meta-heuristics with homotopic recurrent neural networks
- Adversarial Training Makes Weight Loss Landscape Sharper in Logistic Regression
- Deep Contrastive Graph Representation via Adaptive Homotopy Learning
- StackSeq2Seq: Dual Encoder Seq2Seq Recurrent Networks
- Detecting Synapse Location and Connectivity by Signed Proximity Estimation and Pruning with Deep Nets
- Convergence Analysis of Homotopy-SGD for non-convex optimization