Dual-path RNN: efficient long sequence modeling for time-domain single-channel speech separation
arXiv:1910.06379
Abstract
Recent studies in deep learning-based speech separation have proven the superiority of time-domain approaches to conventional time-frequency-based methods. Unlike the time-frequency domain approaches, the time-domain separation systems often receive input sequences consisting of a huge number of time steps, which introduces challenges for modeling extremely long sequences. Conventional recurrent neural networks (RNNs) are not effective for modeling such long sequences due to optimization difficulties, while one-dimensional convolutional neural networks (1-D CNNs) cannot perform utterance-level sequence modeling when its receptive field is smaller than the sequence length. In this paper, we propose dual-path recurrent neural network (DPRNN), a simple yet effective method for organizing RNN layers in a deep structure to model extremely long sequences. DPRNN splits the long sequential input into smaller chunks and applies intra- and inter-chunk operations iteratively, where the input length can be made proportional to the square root of the original sequence length in each operation. Experiments show that by replacing 1-D CNN with DPRNN and apply sample-level modeling in the time-domain audio separation network (TasNet), a new state-of-the-art performance on WSJ0-2mix is achieved with a 20 times smaller model than the previous best system.
ICASSP 2020
References in corpus (7)
- An Empirical Evaluation of Generic Convolutional and Recurrent Networks for Sequence Modeling
- Conv-TasNet: Surpassing Ideal Time-Frequency Magnitude Masking for Speech Separation
- SampleRNN: An Unconditional End-to-End Neural Audio Generation Model
- Hierarchical Multiscale Recurrent Neural Networks
- A Hierarchical Neural Autoencoder for Paragraphs and Documents
- FaSNet: Low-latency Adaptive Beamforming for Multi-microphone Audio Processing
- Meeting Transcription Using Virtual Microphone Arrays
Cited by in corpus (12)
- DCCRN: Deep Complex Convolution Recurrent Network for Phase-Aware Speech Enhancement
- Dual-Path Transformer Network: Direct Context-Aware Modeling for End-to-End Monaural Speech Separation
- LEAF: A Learnable Frontend for Audio Classification
- Phase-Aware Deep Speech Enhancement: It's All About The Frame Length
- Continuous speech separation: dataset and analysis
- Effective Low-Cost Time-Domain Audio Separation Using Globally Attentive Locally Recurrent Networks
- End-to-end Microphone Permutation and Number Invariant Multi-channel Speech Separation
- LaFurca: Iterative Refined Speech Separation Based on Context-Aware Dual-Path Parallel Bi-LSTM
- Rethinking the Separation Layers in Speech Separation Networks
- Speech Separation Based on Multi-Stage Elaborated Dual-Path Deep BiLSTM with Auxiliary Identity Loss
- Toward Speech Separation in The Pre-Cocktail Party Problem with TasTas
- Contrastive Separative Coding for Self-supervised Representation Learning