Single stream parallelization of generalized LSTM-like RNNs on a GPU
arXiv:1503.02852 · doi:10.1109/ICASSP.2015.7178129
Abstract
Recurrent neural networks (RNNs) have shown outstanding performance on processing sequence data. However, they suffer from long training time, which demands parallel implementations of the training procedure. Parallelization of the training algorithms for RNNs are very challenging because internal recurrent paths form dependencies between two different time frames. In this paper, we first propose a generalized graph-based RNN structure that covers the most popular long short-term memory (LSTM) network. Then, we present a parallelization approach that automatically explores parallelisms of arbitrary RNNs by analyzing the graph structure. The experimental results show that the proposed approach shows great speed-up even with a single training stream, and further accelerates the training when combined with multiple parallel training streams.
Accepted by the 40th IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) 2015
Cited by in corpus (13)
- Pix2Vox: Context-aware 3D Reconstruction from Single and Multi-view Images
- Pix2Vox++: Multi-scale Context-aware 3D Object Reconstruction from Single and Multiple Images
- Character-Level Incremental Speech Recognition with Recurrent Neural Networks
- Online Keyword Spotting with a Character-Level Recurrent Neural Network
- ELSA: A Throughput-Optimized Design of an LSTM Accelerator for Energy-Constrained Devices
- DeepTurbo: Deep Turbo Decoder
- ConcealNet: An End-to-end Neural Network for Packet Loss Concealment in Deep Speech Emotion Recognition
- Character-Level Language Modeling with Hierarchical Recurrent Neural Networks
- Generative Artificial Intelligence for Literature Reviews
- FPGA-Based Low-Power Speech Recognition with Recurrent Neural Networks
- Online Sequence Training of Recurrent Neural Networks with Connectionist Temporal Classification
- Fixed-Point Performance Analysis of Recurrent Neural Networks
- Refine3DNet: Scaling Precision in 3D Object Reconstruction from Multi-View RGB Images using Attention