How to Construct Deep Recurrent Neural Networks
arXiv:1312.6026
Abstract
In this paper, we explore different ways to extend a recurrent neural network (RNN) to a \textit{deep} RNN. We start by arguing that the concept of depth in an RNN is not as clear as it is in feedforward neural networks. By carefully analyzing and understanding the architecture of an RNN, however, we find three points of an RNN which may be made deeper; (1) input-to-hidden function, (2) hidden-to-hidden transition and (3) hidden-to-output function. Based on this observation, we propose two novel architectures of a deep RNN which are orthogonal to an earlier attempt of stacking multiple recurrent layers to build a deep RNN (Schmidhuber, 1992; El Hihi and Bengio, 1996). We provide an alternative interpretation of these deep RNNs using a novel framework based on neural operators. The proposed deep RNNs are empirically evaluated on the tasks of polyphonic music prediction and language modeling. The experimental result supports our claim that the proposed deep RNNs benefit from the depth and outperform the conventional, shallow RNNs.
Accepted at ICLR 2014 (Conference Track). 10-page text + 3-page references
References in corpus (7)
- Improving neural networks by preventing co-adaptation of feature detectors
- Modeling Temporal Dependencies in High-Dimensional Sequences: Application to Polyphonic Music Generation and Transcription
- Better Mixing via Deep Representations
- On the number of response regions of deep feed forward networks with piece-wise linear activations
- On Fast Dropout and its Applicability to Recurrent Networks
- A Primal-Dual Method for Training Recurrent Neural Networks Constrained by the Echo-State Property
- Learned-Norm Pooling for Deep Feedforward and Recurrent Neural Networks
Cited by in corpus (87)
- Deep Learning in Neural Networks: An Overview
- Neural Architecture Search with Reinforcement Learning
- Recurrent Neural Network Regularization
- LSTM Fully Convolutional Networks for Time Series Classification
- Multivariate LSTM-FCNs for Time Series Classification
- Character-Aware Neural Language Models
- Recent Advances in Recurrent Neural Networks
- Pointer Sentinel Mixture Models
- Learning to Execute
- Neural Paraphrase Generation with Stacked Residual LSTM Networks
- Learning Stochastic Recurrent Networks
- Pooling Methods in Deep Neural Networks, a Review
- Bridging the Gaps Between Residual Learning, Recurrent Neural Networks and Visual Cortex
- On Extended Long Short-term Memory and Dependent Bidirectional Recurrent Neural Network
- Zoneout: Regularizing RNNs by Randomly Preserving Hidden Activations
- Deep learning models for price forecasting of financial time series: A review of recent advancements: 2020-2022
- Independently Recurrent Neural Network (IndRNN): Building A Longer and Deeper RNN
- Architectural Complexity Measures of Recurrent Neural Networks
- Protein Secondary Structure Prediction with Long Short Term Memory Networks
- On the Origin of Deep Learning
- DR-RNN: A deep residual recurrent neural network for model reduction
- Feedforward Sequential Memory Networks: A New Structure to Learn Long-term Dependency
- Deep Echo State Network (DeepESN): A Brief Survey
- Greenhouse Gas Emission Prediction on Road Network using Deep Sequence Learning
- Sequential Weakly Labeled Multi-Activity Localization and Recognition on Wearable Sensors using Recurrent Attention Networks
- Imposing higher-level Structure in Polyphonic Music Generation using Convolutional Restricted Boltzmann Machines and Constraints
- On Fast Dropout and its Applicability to Recurrent Networks
- DNA-Level Splice Junction Prediction using Deep Recurrent Neural Networks
- Chinese Lexical Analysis with Deep Bi-GRU-CRF Network
- Fast-Slow Recurrent Neural Networks
- Mental Task Classification Using Electroencephalogram Signal
- Interpretable Recurrent Neural Networks Using Sequential Sparse Recovery
- Deep-ESN: A Multiple Projection-encoding Hierarchical Reservoir Computing Framework
- Gated Recurrent Neural Tensor Network
- Image Captioning with Deep Bidirectional LSTMs
- Data Efficient Direct Speech-to-Text Translation with Modality Agnostic Meta-Learning
- Relational recurrent neural networks
- Recurrent Gaussian Processes
- A Survey: Time Travel in Deep Learning Space: An Introduction to Deep Learning Models and How Deep Learning Models Evolved from the Initial Ideas
- Memory-enhanced Decoder for Neural Machine Translation
- deepTarget: End-to-end Learning Framework for microRNA Target Prediction using Deep Recurrent Neural Networks
- On the Evaluation of Sequential Machine Learning for Network Intrusion Detection
- Context-aware Captions from Context-agnostic Supervision
- Producing radiologist-quality reports for interpretable artificial intelligence
- Classification of Alzheimers Disease with Deep Learning on Eye-tracking Data
- Building a Neural Machine Translation System Using Only Synthetic Parallel Data
- Reservoir Topology in Deep Echo State Networks
- Recomposition vs. Prediction: A Novel Anomaly Detection for Discrete Events Based On Autoencoder
- Grow and Prune Compact, Fast, and Accurate LSTMs
- Strongly-Typed Recurrent Neural Networks
- Exploring the Landscape of Ubiquitous In-home Health Monitoring: A Comprehensive Survey
- Fully Convolutional Recurrent Network for Handwritten Chinese Text Recognition
- Deep Neural Machine Translation with Linear Associative Unit
- CKConv: Continuous Kernel Convolution For Sequential Data
- Constructing Long Short-Term Memory based Deep Recurrent Neural Networks for Large Vocabulary Speech Recognition
- Recurrent Neural Networks for Fuzz Testing Web Browsers
- CORAL8: Concurrent Object Regression for Area Localization in Medical Image Panels
- Short-term Memory of Deep RNN
- Modernizing Historical Documents: a User Study
- Syntactically Informed Text Compression with Recurrent Neural Networks
- Understanding Recurrent Neural Networks Using Nonequilibrium Response Theory
- Batch-normalized Recurrent Highway Networks
- Scalable Balanced Training of Conditional Generative Adversarial Neural Networks on Image Data
- Managing travel demand: Location recommendation for system efficiency based on mobile phone data
- Learned-Norm Pooling for Deep Feedforward and Recurrent Neural Networks
- A Hierarchical Recurrent Encoder-Decoder For Generative Context-Aware Query Suggestion
- Teaching Machines to Code: Neural Markup Generation with Visual Attention
- Stacked LSTM Based Deep Recurrent Neural Network with Kalman Smoothing for Blood Glucose Prediction
- RCNet: Incorporating Structural Information into Deep RNN for MIMO-OFDM Symbol Detection with Limited Training
- DeepHoops: Evaluating Micro-Actions in Basketball Using Deep Feature Representations of Spatio-Temporal Data
- Learning Less-Overlapping Representations
- Financial Table Extraction in Image Documents
- Deep Recurrent Gaussian Process with Variational Sparse Spectrum Approximation
- Interpreting and Disentangling Feature Components of Various Complexity from DNNs
- Long Short-Term Memory with Dynamic Skip Connections
- Learning Spatial-Semantic Context with Fully Convolutional Recurrent Network for Online Handwritten Chinese Text Recognition
- Variational inference of latent state sequences using Recurrent Networks
- Regularizing Recurrent Networks - On Injected Noise and Norm-based Methods
- Recurrent Neural Network from Adder's Perspective: Carry-lookahead RNN
- A Differentially Private Multi-Output Deep Generative Networks Approach For Activity Diary Synthesis
- Optimizing and Contrasting Recurrent Neural Network Architectures
- Bi-directional LSTM Recurrent Neural Network for Chinese Word Segmentation
- Language Modeling with Highway LSTM
- Residual Memory Networks: Feed-forward approach to learn long temporal dependencies
- Temporally Folded Convolutional Neural Networks for Sequence Forecasting
- RNN Training along Locally Optimal Trajectories via Frank-Wolfe Algorithm
- Deformable Stacked Structure for Named Entity Recognition