Tunable Efficient Unitary Neural Networks (EUNN) and their application to RNNs
arXiv:1612.05231
Abstract
Using unitary (instead of general) matrices in artificial neural networks (ANNs) is a promising way to solve the gradient explosion/vanishing problem, as well as to enable ANNs to learn long-term correlations in the data. This approach appears particularly promising for Recurrent Neural Networks (RNNs). In this work, we present a new architecture for implementing an Efficient Unitary Neural Network (EUNNs); its main advantages can be summarized as follows. Firstly, the representation capacity of the unitary space in an EUNN is fully tunable, ranging from a subspace of SU(N) to the entire unitary space. Secondly, the computational complexity for training an EUNN is merely per parameter. Finally, we test the performance of EUNNs on the standard copying task, the pixel-permuted MNIST digit recognition benchmark as well as the Speech Prediction Test (TIMIT). We find that our architecture significantly outperforms both other state-of-the-art unitary RNNs and the LSTM architecture, in terms of the final performance and/or the wall-clock training speed. EUNNs are thus promising alternatives to RNNs and LSTMs for a wide variety of applications.
9 pages, 4 figures
References in corpus (5)
Cited by in corpus (44)
- Circuit-centric quantum classifiers
- Wave Physics as an Analog Recurrent Neural Network
- Experimentally realized in situ backpropagation for deep learning in nanophotonic neural networks
- Differentiable Learning of Quantum Circuit Born Machine
- Barren Plateaus in Variational Quantum Computing
- Single chip photonic deep neural network with accelerated training
- FastGRNN: A Fast, Accurate, Stable and Tiny Kilobyte Sized Gated Recurrent Neural Network
- Stable Recurrent Models
- Cheap Orthogonal Constraints in Neural Networks: A Simple Parametrization of the Orthogonal and Unitary Group
- Learning quantum data with the quantum Earth Mover's distance
- Parseval Proximal Neural Networks
- How to Start Training: The Effect of Initialization and Architecture
- Lipschitz Recurrent Neural Networks
- Trivializations for Gradient-Based Optimization on Manifolds
- Continual Learning in Low-rank Orthogonal Subspaces
- Compressing RNNs for IoT devices by 15-38x using Kronecker Products
- Non-normal Recurrent Neural Network (nnRNN): learning long time dependencies while improving expressivity with transient dynamics
- Heavy Ball Neural Ordinary Differential Equations
- Estimating the randomness of quantum circuit ensembles up to 50 qubits
- Slower is Better: Revisiting the Forgetting Mechanism in LSTM for Slower Information Decay
- Kaleidoscope: An Efficient, Learnable Representation For All Structured Linear Maps
- Pixelated Butterfly: Simple and Efficient Sparse training for Neural Network Models
- The Recurrent Neural Tangent Kernel
- MomentumRNN: Integrating Momentum into Recurrent Neural Networks
- Equilibrated Recurrent Neural Network: Neuronal Time-Delayed Self-Feedback Improves Accuracy and Stability
- Orthogonal Over-Parameterized Training
- Feedback Gradient Descent: Efficient and Stable Optimization with Orthogonality for DNNs
- RNNs Evolving on an Equilibrium Manifold: A Panacea for Vanishing and Exploding Gradients?
- A Differential Geometry Perspective on Orthogonal Recurrent Models
- Eigenvalue Normalized Recurrent Neural Networks for Short Term Memory
- Scaling-up Diverse Orthogonal Convolutional Networks with a Paraunitary Framework
- Deep Unitary Convolutional Neural Networks
- Differentiate Everything with a Reversible Embeded Domain-Specific Language
- Recurrent Neural Network from Adder's Perspective: Carry-lookahead RNN
- Building Compact and Robust Deep Neural Networks with Toeplitz Matrices
- Implicit Bias of Linear RNNs
- Multi-Decoder RNN Autoencoder Based on Variational Bayes Method
- RotLSTM: Rotating Memories in Recurrent Neural Networks
- Learning with Hyperspherical Uniformity
- RNN Training along Locally Optimal Trajectories via Frank-Wolfe Algorithm
- Depth Enables Long-Term Memory for Recurrent Neural Networks
- Complex Evolution Recurrent Neural Networks (ceRNNs)
- Parallelized Computation and Backpropagation Under Angle-Parametrized Orthogonal Matrices
- CWY Parametrization: a Solution for Parallelized Optimization of Orthogonal and Stiefel Matrices