A Simple Way to Initialize Recurrent Networks of Rectified Linear Units
arXiv:1504.00941
Abstract
Learning long term dependencies in recurrent networks is difficult due to vanishing and exploding gradients. To overcome this difficulty, researchers have developed sophisticated optimization techniques and network architectures. In this paper, we propose a simpler solution that use recurrent neural networks composed of rectified linear units. Key to our solution is the use of the identity matrix or its scaled version to initialize the recurrent weight matrix. We find that our solution is comparable to LSTM on our four benchmarks: two toy problems involving long-range temporal structures, a large language modeling problem and a benchmark speech recognition problem.
References in corpus (5)
Cited by in corpus (66)
- Artificial neural networks for neuroscientists: A primer
- Working Memory Connections for LSTM
- Real-valued (Medical) Time Series Generation with Recurrent Conditional GANs
- RETAIN: An Interpretable Predictive Model for Healthcare using Reverse Time Attention Mechanism
- Data Noising as Smoothing in Neural Network Language Models
- What-and-Where to Match: Deep Spatially Multiplicative Integration Networks for Person Re-identification
- Full-Capacity Unitary Recurrent Neural Networks
- Compressing Recurrent Neural Network with Tensor Train
- FastGRNN: A Fast, Accurate, Stable and Tiny Kilobyte Sized Gated Recurrent Neural Network
- Characterizing Driving Styles with Deep Learning
- R-Transformer: Recurrent Neural Network Enhanced Transformer
- Implicit Regularization in Deep Learning
- Towards End-to-End Speech Recognition with Deep Convolutional Neural Networks
- Revisiting Activation Regularization for Language RNNs
- Dirichlet Energy Constrained Learning for Deep Graph Neural Networks
- Trivializations for Gradient-Based Optimization on Manifolds
- Dynamic Neural Turing Machine with Soft and Hard Addressing Schemes
- Provable Benefit of Orthogonal Initialization in Optimizing Deep Linear Networks
- Physics Informed LSTM Network for Flexibility Identification in Evaporative Cooling Systems
- Wider and Deeper, Cheaper and Faster: Tensorized LSTMs for Sequence Learning
- On Generalization Bounds of a Family of Recurrent Neural Networks
- The Statistical Recurrent Unit
- Learning to Remember More with Less Memorization
- On the Initialization of Long Short-Term Memory Networks
- Calibration, Entropy Rates, and Memory in Language Models
- Demystifying Deep Learning in Predictive Spatio-Temporal Analytics: An Information-Theoretic Framework
- RRA: Recurrent Residual Attention for Sequence Learning
- State-Regularized Recurrent Neural Networks
- Effective and Efficient Computation with Multiple-timescale Spiking Recurrent Neural Networks
- A Simple LSTM model for Transition-based Dependency Parsing
- Kaleidoscope: An Efficient, Learnable Representation For All Structured Linear Maps
- Cached Long Short-Term Memory Neural Networks for Document-Level Sentiment Classification
- On the Variance of Unbiased Online Recurrent Optimization
- Efficient LSTM Training with Eligibility Traces
- How Large a Vocabulary Does Text Classification Need? A Variational Approach to Vocabulary Selection
- Contracting Implicit Recurrent Neural Networks: Stable Models with Improved Trainability
- A Basic Recurrent Neural Network Model
- Encoding-based Memory Modules for Recurrent Neural Networks
- Recognizing Video Events with Varying Rhythms
- S++: A Fast and Deployable Secure-Computation Framework for Privacy-Preserving Neural Network Training
- Eigenvalue Normalized Recurrent Neural Networks for Short Term Memory
- Cortico-cerebellar networks as decoupling neural interfaces
- The relationship between Biological and Artificial Intelligence
- Learning to Adaptively Scale Recurrent Neural Networks
- A Fully Trainable Network with RNN-based Pooling
- Predicting Patient State-of-Health using Sliding Window and Recurrent Classifiers
- Empirical Evaluation of A New Approach to Simplifying Long Short-term Memory (LSTM)
- Learning Various Length Dependence by Dual Recurrent Neural Networks
- Discrete Function Bases and Convolutional Neural Networks
- Multi-Task Learning for Argumentation Mining
- Memory and attention in deep learning
- Recurrent Neural Network from Adder's Perspective: Carry-lookahead RNN
- Character-level Deep Conflation for Business Data Analytics
- Bayesian Sparsification of Gated Recurrent Neural Networks
- Learning Longer-term Dependencies via Grouped Distributor Unit
- Thick-Net: Parallel Network Structure for Sequential Modeling
- Short-Term Memory Optimization in Recurrent Neural Networks by Autoencoder-based Initialization
- Network of Recurrent Neural Networks
- Online Spatiotemporal Action Detection and Prediction via Causal Representations
- Deep Symbolic Representation Learning for Heterogeneous Time-series Classification
- Graphical RNN Models
- Solving hybrid machine learning tasks by traversing weight space geodesics
- Back to Square One: Superhuman Performance in Chutes and Ladders Through Deep Neural Networks and Tree Search
- Relational Weight Priors in Neural Networks for Abstract Pattern Learning and Language Modelling
- Target Propagation via Regularized Inversion
- An empirical analysis of phrase-based and neural machine translation