Architectural Complexity Measures of Recurrent Neural Networks
arXiv:1602.08210
Abstract
In this paper, we systematically analyze the connecting architectures of recurrent neural networks (RNNs). Our main contribution is twofold: first, we present a rigorous graph-theoretic framework describing the connecting architectures of RNNs in general. Second, we propose three architecture complexity measures of RNNs: (a) the recurrent depth, which captures the RNN's over-time nonlinear complexity, (b) the feedforward depth, which captures the local input-output nonlinearity (similar to the "depth" in feedforward neural networks (FNNs)), and (c) the recurrent skip coefficient which captures how rapidly the information propagates over time. We rigorously prove each measure's existence and computability. Our experimental results show that RNNs might benefit from larger recurrent depth and feedforward depth. We further demonstrate that increasing recurrent skip coefficient offers performance boosts on long term dependency problems.
17 pages, 8 figures; To appear in NIPS2016
References in corpus (4)
Cited by in corpus (22)
- An overview and comparative analysis of Recurrent Neural Networks for Short Term Load Forecasting
- Hierarchical Multiscale Recurrent Neural Networks
- AdaNet: Adaptive Structural Learning of Artificial Neural Networks
- Multiplicative LSTM for sequence modelling
- Convolutional Tensor-Train LSTM for Spatio-temporal Learning
- The unreasonable effectiveness of the forget gate
- On Multiplicative Integration with Recurrent Neural Networks
- Recurrent Batch Normalization
- Image Super-Resolution via Dual-State Recurrent Networks
- Efficient Visual Recognition with Deep Neural Networks: A Survey on Recent Advances and New Directions
- Path-Normalized Optimization of Recurrent Neural Networks with ReLU Activations
- Bayesian LSTMs in medicine
- Techniques for visualizing LSTMs applied to electrocardiograms
- Low-rank passthrough neural networks
- Making Good on LSTMs' Unfulfilled Promise
- Long Short-Term Memory with Dynamic Skip Connections
- Multi-Zone Unit for Recurrent Neural Networks
- On the Long-Term Memory of Deep Recurrent Networks
- Learning to Adaptively Scale Recurrent Neural Networks
- Recurrent Neural Network from Adder's Perspective: Carry-lookahead RNN
- Optimizing Recurrent Neural Networks Architectures under Time Constraints
- Temporally Folded Convolutional Neural Networks for Sequence Forecasting