Unitary Evolution Recurrent Neural Networks
arXiv:1511.06464
Abstract
Recurrent neural networks (RNNs) are notoriously difficult to train. When the eigenvalues of the hidden to hidden weight matrix deviate from absolute value 1, optimization becomes difficult due to the well studied issue of vanishing and exploding gradients, especially when trying to learn long-term dependencies. To circumvent this problem, we propose a new architecture that learns a unitary weight matrix, with eigenvalues of absolute value exactly 1. The challenge we address is that of parametrizing unitary matrices in a way that does not require expensive computations (such as eigendecomposition) after each weight update. We construct an expressive unitary weight matrix by composing several structured matrices that act as building blocks with parameters to be learned. Optimization with this parameterization becomes feasible only when considering hidden states in the complex domain. We demonstrate the potential of this architecture by achieving state of the art results in several hard tasks involving very long-term dependencies.
References in corpus (3)
Cited by in corpus (48)
- Deep Learning with Coherent Nanophotonic Circuits
- Circuit-centric quantum classifiers
- Regularizing and Optimizing LSTM Language Models
- Feed-Forward Networks with Attention Can Solve Some Long-Term Memory Problems
- Machine Learning for Microcontroller-Class Hardware: A Review
- Hardware error correction for programmable photonics
- Data Noising as Smoothing in Neural Network Language Models
- Deep Complex Networks
- On orthogonality and learning recurrent networks with long term dependencies
- Independently Recurrent Neural Network (IndRNN): Building A Longer and Deeper RNN
- Architectural Complexity Measures of Recurrent Neural Networks
- Tunable Efficient Unitary Neural Networks (EUNN) and their application to RNNs
- Universal discriminative quantum neural networks
- Multiplicative LSTM for sequence modelling
- Implicit Regularization in Deep Learning
- Recurrent Batch Normalization
- OSLNet: Deep Small-Sample Classification with an Orthogonal Softmax Layer
- Training Optimization for Gate-Model Quantum Neural Networks
- Parseval Proximal Neural Networks
- Geoopt: Riemannian Optimization in PyTorch
- Photonic convolutional neural networks using integrated diffractive optics
- Compression of Recurrent Neural Networks for Efficient Language Modeling
- Neural Networks Compression for Language Modeling
- Compressing RNNs for IoT devices by 15-38x using Kronecker Products
- DizzyRNN: Reparameterizing Recurrent Neural Networks for Norm-Preserving Backpropagation
- Classification of Periodic Variable Stars with Novel Cyclic-Permutation Invariant Neural Networks
- Stabilizing Gradients for Deep Neural Networks via Efficient SVD Parameterization
- Path-Normalized Optimization of Recurrent Neural Networks with ReLU Activations
- Parameterized Hypercomplex Graph Neural Networks for Graph Classification
- Reservoir Computing meets Recurrent Kernels and Structured Transforms
- Analysis of Deep Complex-Valued Convolutional Neural Networks for MRI Reconstruction
- Norm-preserving Orthogonal Permutation Linear Unit Activation Functions (OPLU)
- Orthogonal Deep Neural Networks
- Co-VeGAN: Complex-Valued Generative Adversarial Network for Compressive Sensing MR Image Reconstruction
- Information Geometry of Orthogonal Initializations and Training
- A Variance-Reduced Stochastic Gradient Tracking Algorithm for Decentralized Optimization with Orthogonality Constraints
- Low-rank passthrough neural networks
- Efficient LSTM Training with Eligibility Traces
- Feedback Gradient Descent: Efficient and Stable Optimization with Orthogonality for DNNs
- Improving training of deep neural networks via Singular Value Bounding
- Encoding-based Memory Modules for Recurrent Neural Networks
- On the biological plausibility of orthogonal initialisation for solving gradient instability in deep neural networks
- MCRM: Mother Compact Recurrent Memory
- Orthogonal Directions Constrained Gradient Method: from non-linear equality constraints to Stiefel manifold
- Differentiate Everything with a Reversible Embeded Domain-Specific Language
- Learning Longer-term Dependencies via Grouped Distributor Unit
- Utilizing Complex-valued Network for Learning to Compare Image Patches
- Alternating Synthetic and Real Gradients for Neural Language Modeling