Dual-Path Transformer Network: Direct Context-Aware Modeling for End-to-End Monaural Speech Separation
arXiv:2007.13975
Abstract
The dominant speech separation models are based on complex recurrent or convolution neural network that model speech sequences indirectly conditioning on context, such as passing information through many intermediate states in recurrent neural network, leading to suboptimal separation performance. In this paper, we propose a dual-path transformer network (DPTNet) for end-to-end speech separation, which introduces direct context-awareness in the modeling for speech sequences. By introduces a improved transformer, elements in speech sequences can interact directly, which enables DPTNet can model for the speech sequences with direct context-awareness. The improved transformer in our approach learns the order information of the speech sequences without positional encodings by incorporating a recurrent neural network into the original transformer. In addition, the structure of dual paths makes our model efficient for extremely long speech sequence modeling. Extensive experiments on benchmark datasets show that our approach outperforms the current state-of-the-arts (20.6 dB SDR on the public WSj0-2mix data corpus).
5 pages. Accepted by INTERSPEECH 2020
Cited by in corpus (9)
- TSTNN: Two-stage Transformer based Neural Network for Speech Enhancement in the Time Domain
- Effective Low-Cost Time-Domain Audio Separation Using Globally Attentive Locally Recurrent Networks
- Distributed speech separation in spatially unconstrained microphone arrays
- Sandglasset: A Light Multi-Granularity Self-attentive Network For Time-Domain Speech Separation
- North America Bixby Speaker Diarization System for the VoxCeleb Speaker Recognition Challenge 2021
- TransMask: A Compact and Fast Speech Separation Model Based on Transformer
- Multi-Task Audio Source Separation
- MIMO Self-attentive RNN Beamformer for Multi-speaker Speech Separation
- Tune-In: Training Under Negative Environments with Interference for Attention Networks Simulating Cocktail Party Effect