A Simple Way to Initialize Recurrent Networks of Rectified Linear Units
arXiv:1504.00941
Abstract
Learning long term dependencies in recurrent networks is difficult due to vanishing and exploding gradients. To overcome this difficulty, researchers have developed sophisticated optimization techniques and network architectures. In this paper, we propose a simpler solution that use recurrent neural networks composed of rectified linear units. Key to our solution is the use of the identity matrix or its scaled version to initialize the recurrent weight matrix. We find that our solution is comparable to LSTM on our four benchmarks: two toy problems involving long-range temporal structures, a large language modeling problem and a benchmark speech recognition problem.
References in corpus (5)
Cited by in corpus (56)
- Artificial neural networks for neuroscientists: A primer
- Real-valued (Medical) Time Series Generation with Recurrent Conditional GANs
- RETAIN: An Interpretable Predictive Model for Healthcare using Reverse Time Attention Mechanism
- Data Noising as Smoothing in Neural Network Language Models
- What-and-Where to Match: Deep Spatially Multiplicative Integration Networks for Person Re-identification
- Full-Capacity Unitary Recurrent Neural Networks
- Compressing Recurrent Neural Network with Tensor Train
- FastGRNN: A Fast, Accurate, Stable and Tiny Kilobyte Sized Gated Recurrent Neural Network
- R-Transformer: Recurrent Neural Network Enhanced Transformer
- Characterizing Driving Styles with Deep Learning
- Implicit Regularization in Deep Learning
- Towards End-to-End Speech Recognition with Deep Convolutional Neural Networks
- Revisiting Activation Regularization for Language RNNs
- Trivializations for Gradient-Based Optimization on Manifolds
- Dynamic Neural Turing Machine with Soft and Hard Addressing Schemes
- Provable Benefit of Orthogonal Initialization in Optimizing Deep Linear Networks
- Wider and Deeper, Cheaper and Faster: Tensorized LSTMs for Sequence Learning
- On Generalization Bounds of a Family of Recurrent Neural Networks
- The Statistical Recurrent Unit
- Learning to Remember More with Less Memorization
- Calibration, Entropy Rates, and Memory in Language Models
- On the Initialization of Long Short-Term Memory Networks
- RRA: Recurrent Residual Attention for Sequence Learning
- Demystifying Deep Learning in Predictive Spatio-Temporal Analytics: An Information-Theoretic Framework
- Effective and Efficient Computation with Multiple-timescale Spiking Recurrent Neural Networks
- State-Regularized Recurrent Neural Networks
- Kaleidoscope: An Efficient, Learnable Representation For All Structured Linear Maps
- A Simple LSTM model for Transition-based Dependency Parsing
- Cached Long Short-Term Memory Neural Networks for Document-Level Sentiment Classification
- On the Variance of Unbiased Online Recurrent Optimization
- How Large a Vocabulary Does Text Classification Need? A Variational Approach to Vocabulary Selection
- Contracting Implicit Recurrent Neural Networks: Stable Models with Improved Trainability
- A Basic Recurrent Neural Network Model
- Recognizing Video Events with Varying Rhythms
- S++: A Fast and Deployable Secure-Computation Framework for Privacy-Preserving Neural Network Training
- Encoding-based Memory Modules for Recurrent Neural Networks
- Eigenvalue Normalized Recurrent Neural Networks for Short Term Memory
- The relationship between Biological and Artificial Intelligence
- Learning to Adaptively Scale Recurrent Neural Networks
- Character-level Deep Conflation for Business Data Analytics
- A Fully Trainable Network with RNN-based Pooling
- Bayesian Sparsification of Gated Recurrent Neural Networks
- Predicting Patient State-of-Health using Sliding Window and Recurrent Classifiers
- Empirical Evaluation of A New Approach to Simplifying Long Short-term Memory (LSTM)
- Learning Various Length Dependence by Dual Recurrent Neural Networks
- Discrete Function Bases and Convolutional Neural Networks
- Multi-Task Learning for Argumentation Mining
- Graphical RNN Models
- Network of Recurrent Neural Networks
- An empirical analysis of phrase-based and neural machine translation
- Back to Square One: Superhuman Performance in Chutes and Ladders Through Deep Neural Networks and Tree Search
- Thick-Net: Parallel Network Structure for Sequential Modeling
- Relational Weight Priors in Neural Networks for Abstract Pattern Learning and Language Modelling
- Deep Symbolic Representation Learning for Heterogeneous Time-series Classification
- Learning Longer-term Dependencies via Grouped Distributor Unit
- Online Spatiotemporal Action Detection and Prediction via Causal Representations