Training Input-Output Recurrent Neural Networks through Spectral Methods
arXiv:1603.00954
Abstract
We consider the problem of training input-output recurrent neural networks (RNN) for sequence labeling tasks. We propose a novel spectral approach for learning the network parameters. It is based on decomposition of the cross-moment tensor between the output and a non-linear transformation of the input, based on score functions. We guarantee consistent learning with polynomial sample and computational complexity under transparent conditions such as non-degeneracy of model parameters, polynomial activations for the neurons, and a Markovian evolution of the input sequence. We also extend our results to Bidirectional RNN which uses both previous and future information to output the label at each time point, and is employed in many NLP tasks such as POS tagging.
References in corpus (5)
- A Critical Review of Recurrent Neural Networks for Sequence Learning
- A Spectral Algorithm for Latent Dirichlet Allocation
- Beating the Perils of Non-Convexity: Guaranteed Training of Neural Networks using Tensor Methods
- Concentration inequalities for dependent Random variables via the martingale method
- Sample Complexity Analysis for Learning Overcomplete Latent Variable Models through Tensor Methods
Cited by in corpus (8)
- Non-convex Optimization for Machine Learning
- On the Optimization of Deep Networks: Implicit Acceleration by Overparameterization
- Stable Recurrent Models
- Tensor Regression Networks
- Deep Learning and Quantum Entanglement: Fundamental Connections with Implications to Network Design
- Connecting Weighted Automata and Recurrent Neural Networks through Spectral Learning
- Tensor Contraction Layers for Parsimonious Deep Nets
- Prediction with a Short Memory