Supervised Speech Separation Based on Deep Learning: An Overview
arXiv:1708.07524
Abstract
Speech separation is the task of separating target speech from background interference. Traditionally, speech separation is studied as a signal processing problem. A more recent approach formulates speech separation as a supervised learning problem, where the discriminative patterns of speech, speakers, and background noise are learned from training data. Over the past decade, many supervised separation algorithms have been put forward. In particular, the recent introduction of deep learning to supervised speech separation has dramatically accelerated progress and boosted separation performance. This article provides a comprehensive overview of the research on deep learning based supervised speech separation in the last several years. We first introduce the background of speech separation and the formulation of supervised separation. Then we discuss three main components of supervised separation: learning machines, training targets, and acoustic features. Much of the overview is on separation algorithms where we review monaural methods, including speech enhancement (speech-nonspeech separation), speaker separation (multi-talker separation), and speech dereverberation, as well as multi-microphone techniques. The important issue of generalization, unique to supervised learning, is discussed. This overview provides a historical perspective on how advances are made. In addition, we discuss a number of conceptual issues, including what constitutes the target source.
27 pages, 17 figures
References in corpus (5)
- Deep Learning in Neural Networks: An Overview
- On the difficulty of training Recurrent Neural Networks
- Conditional Generative Adversarial Networks for Speech Enhancement and Noise-Robust Speaker Verification
- Deep Long Short-Term Memory Adaptive Beamforming Networks For Multichannel Robust Speech Recognition
- Raw Waveform-based Speech Enhancement by Fully Convolutional Networks
Cited by in corpus (6)
- Deep Learning for Audio Signal Processing
- End-to-End Speech Separation with Unfolded Iterative Phase Reconstruction
- End-to-End Waveform Utterance Enhancement for Direct Evaluation Metrics Optimization by Fully Convolutional Neural Networks
- Fréchet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms
- The CHiME-7 UDASE task: Unsupervised domain adaptation for conversational speech enhancement
- Complex Spectral Mapping With Attention Based Convolution Recurrent Neural Network for Speech Enhancement