On Generalization Bounds of a Family of Recurrent Neural Networks
arXiv:1910.12947
Abstract
Recurrent Neural Networks (RNNs) have been widely applied to sequential data analysis. Due to their complicated modeling structures, however, the theory behind is still largely missing. To connect theory and practice, we study the generalization properties of vanilla RNNs as well as their variants, including Minimal Gated Unit (MGU), Long Short Term Memory (LSTM), and Convolutional (Conv) RNNs. Specifically, our theory is established under the PAC-Learning framework. The generalization bound is presented in terms of the spectral norms of the weight matrices and the total number of parameters. We also establish refined generalization bounds with additional norm assumptions, and draw a comparison among these bounds. We remark: (1) Our generalization bound for vanilla RNNs is significantly tighter than the best of existing results; (2) We are not aware of any other generalization bounds for MGU, LSTM, and Conv RNNs in the exiting literature; (3) We demonstrate the advantages of these variants in generalization.
30 pages, 5 figures
References in corpus (7)
- Sequence to Sequence Learning with Neural Networks
- Sequence Transduction with Recurrent Neural Networks
- Understanding deep learning requires rethinking generalization
- DRAW: A Recurrent Neural Network For Image Generation
- A Simple Way to Initialize Recurrent Networks of Rectified Linear Units
- Pointer Sentinel Mixture Models
- Path-SGD: Path-Normalized Optimization in Deep Neural Networks
Cited by in corpus (7)
- Generalization and Representational Limits of Graph Neural Networks
- Recent advances in deep learning theory
- Inductive Biases and Variable Creation in Self-Attention Mechanisms
- Can SGD Learn Recurrent Neural Networks with Provable Generalization?
- PAC-Bayes Generalisation Bounds for Dynamical Systems Including Stable RNNs
- Understanding Deep Architectures with Reasoning Layer
- Spectral Pruning for Recurrent Neural Networks