Deep attractor network for single-microphone speaker separation
arXiv:1611.08930 · doi:10.1109/ICASSP.2017.7952155
Abstract
Despite the overwhelming success of deep learning in various speech processing tasks, the problem of separating simultaneous speakers in a mixture remains challenging. Two major difficulties in such systems are the arbitrary source permutation and unknown number of sources in the mixture. We propose a novel deep learning framework for single channel speech separation by creating attractor points in high dimensional embedding space of the acoustic signals which pull together the time-frequency bins corresponding to each source. Attractor points in this study are created by finding the centroids of the sources in the embedding space, which are subsequently used to determine the similarity of each bin in the mixture to each source. The network is then trained to minimize the reconstruction error of each source by optimizing the embeddings. The proposed model is different from prior works in that it implements an end-to-end training, and it does not depend on the number of sources in the mixture. Two strategies are explored in the test time, K-means and fixed attractor points, where the latter requires no post-processing and can be implemented in real-time. We evaluated our system on Wall Street Journal dataset and show 5.49\% improvement over the previous state-of-the-art methods.
2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
References in corpus (1)
Cited by in corpus (97)
- Conv-TasNet: Surpassing Ideal Time-Frequency Magnitude Masking for Speech Separation
- Deep Learning for Audio Signal Processing
- Speaker-independent Speech Separation with Deep Attractor Network
- SpEx: Multi-Scale Time Domain Speaker Extraction Network
- Voice Separation with an Unknown Number of Multiple Speakers
- Deep Filtering: Signal Extraction and Reconstruction Using Complex Time-Frequency Filters
- A Multi-Phase Gammatone Filterbank for Speech Separation via TasNet
- Encoder-Decoder Based Attractors for End-to-End Neural Diarization
- Face Landmark-based Speaker-Independent Audio-Visual Speech Enhancement in Multi-Talker Environments
- Asteroid: the PyTorch-based audio source separation toolkit for researchers
- VoiceFilter: Targeted Voice Separation by Speaker-Conditioned Spectrogram Masking
- On End-to-end Multi-channel Time Domain Speech Separation in Reverberant Environments
- Analysis of Deep Clustering as Preprocessing for Automatic Speech Recognition of Sparsely Overlapping Speech
- TasNet: time-domain audio separation network for real-time, single-channel speech separation
- Time-domain speaker extraction network
- Dual-Path Transformer Network: Direct Context-Aware Modeling for End-to-End Monaural Speech Separation
- Progressive Joint Modeling in Unsupervised Single-channel Overlapped Speech Recognition
- BW-EDA-EEND: Streaming End-to-End Neural Speaker Diarization for a Variable Number of Speakers
- End-to-End Speech Separation with Unfolded Iterative Phase Reconstruction
- Psychoacoustic Calibration of Loss Functions for Efficient End-to-End Neural Audio Coding
- Event-Independent Network for Polyphonic Sound Event Localization and Detection
- SAGRNN: Self-Attentive Gated RNN for Binaural Speaker Separation with Interaural Cue Preservation
- SpEx+: A Complete Time Domain Speaker Extraction Network
- Signal-Aware Direction-of-Arrival Estimation Using Attention Mechanisms
- FurcaNet: An end-to-end deep gated convolutional, long short-term memory, deep neural networks for single channel speech separation
- End-to-End Speaker Diarization for an Unknown Number of Speakers with Encoder-Decoder Based Attractors
- Improving speaker discrimination of target speech extraction with time-domain SpeakerBeam
- Recognizing Multi-talker Speech with Permutation Invariant Training
- Audio-visual Recognition of Overlapped speech for the LRS2 dataset
- VoiceFilter-Lite: Streaming Targeted Voice Separation for On-Device Speech Recognition
- X-TaSNet: Robust and Accurate Time-Domain Speaker Extraction Network
- A comparison of handcrafted, parameterized, and learnable features for speech separation
- LaFurca: Iterative Refined Speech Separation Based on Context-Aware Dual-Path Parallel Bi-LSTM
- Joint Separation and Denoising of Noisy Multi-talker Speech using Recurrent Neural Networks and Permutation Invariant Training
- DBNET: DOA-driven beamforming network for end-to-end farfield sound source separation
- Simultaneous Speech Recognition and Speaker Diarization for Monaural Dialogue Recordings with Target-Speaker Acoustic Models
- Improving Source Separation via Multi-Speaker Representations
- Building Corpora for Single-Channel Speech Separation Across Multiple Domains
- Target Speaker Extraction for Overlapped Multi-Talker Speaker Verification
- Deep Attention Fusion Feature for Speech Separation with End-to-End Post-filter Method
- Joint Speaker Counting, Speech Recognition, and Speaker Identification for Overlapped Speech of Any Number of Speakers
- Atss-Net: Target Speaker Separation via Attention-based Neural Network
- Speaker-Conditional Chain Model for Speech Separation and Extraction
- Two-stage model and optimal SI-SNR for monaural multi-speaker speech separation in noisy environment
- Single-Channel Speech Separation with Auxiliary Speaker Embeddings
- FurcaNeXt: End-to-end monaural speech separation with dynamic gated dilated temporal convolutional networks
- Auxiliary Interference Speaker Loss for Target-Speaker Speech Recognition
- Improved Speech Separation with Time-and-Frequency Cross-domain Joint Embedding and Clustering
- Unsupervised training of a deep clustering model for multichannel blind source separation
- Closing the Training/Inference Gap for Deep Attractor Networks
- Investigation of End-To-End Speaker-Attributed ASR for Continuous Multi-Talker Recordings
- Neural Spatio-Temporal Beamformer for Target Speech Separation
- Unsupervised Speaker Diarization in Distributed IoT Networks Using Federated Learning
- A Study of Transfer Learning in Music Source Separation
- Cracking the cocktail party problem by multi-beam deep attractor network
- Supervised Speaker Embedding De-Mixing in Two-Speaker Environment
- Mixup-breakdown: a consistency training method for improving generalization of speech separation models
- The sound of my voice: speaker representation loss for target voice separation
- Multimodal Target Speech Separation with Voice and Face References
- Multi-channel target speech extraction with channel decorrelation and target speaker adaptation
- MITAS: A Compressed Time-Domain Audio Separation Network with Parameter Sharing
- Improving Neural Diarization through Speaker Attribute Attractors and Local Dependency Modeling
- Discriminative Learning for Monaural Speech Separation Using Deep Embedding Features
- Low-Latency Speaker-Independent Continuous Speech Separation
- Auditory Separation of a Conversation from Background via Attentional Gating
- Speakerfilter-Pro: an improved target speaker extractor combines the time domain and frequency domain
- On TasNet for Low-Latency Single-Speaker Speech Enhancement
- Distortion-controlled Training for End-to-end Reverberant Speech Separation with Auxiliary Autoencoding Loss
- Multi-channel Speech Separation Using Deep Embedding Model with Multilayer Bootstrap Networks
- Optimization of Speaker Extraction Neural Network with Magnitude and Temporal Spectrum Approximation Loss
- Unsupervised Training for Deep Speech Source Separation with Kullback-Leibler Divergence Based Probabilistic Loss Function
- Distortionless Multi-Channel Target Speech Enhancement for Overlapped Speech Recognition
- MIMO Self-attentive RNN Beamformer for Multi-speaker Speech Separation
- Speech Separation Based on Multi-Stage Elaborated Dual-Path Deep BiLSTM with Auxiliary Identity Loss
- Towards Neural Diarization for Unlimited Numbers of Speakers Using Global and Local Attractors
- Fact-aware Sentence Split and Rephrase with Permutation Invariant Training
- Orthonormal Embedding-based Deep Clustering for Single-channel Speech Separation
- Far-Field Automatic Speech Recognition
- Lightweight Dual-channel Target Speaker Separation for Mobile Voice Communication
- Tune-In: Training Under Negative Environments with Interference for Attention Networks Simulating Cocktail Party Effect
- Manifold-Aware Deep Clustering: Maximizing Angles between Embedding Vectors Based on Regular Simplex
- Multichannel Loss Function for Supervised Speech Source Separation by Mask-based Beamforming
- Toward Speech Separation in The Pre-Cocktail Party Problem with TasTas
- End-to-End Multi-Look Keyword Spotting
- Attention-based scaling adaptation for target speech extraction
- Stabilizing Label Assignment for Speech Separation by Self-supervised Pre-training
- X-DC: Explainable Deep Clustering based on Learnable Spectrogram Templates
- Do We Need Sound for Sound Source Localization?
- Online Self-Attentive Gated RNNs for Real-Time Speaker Separation
- Region segmentation via deep learning and convex optimization
- Time Domain Audio Visual Speech Separation
- Practical applicability of deep neural networks for overlapping speaker separation
- Low-Latency Deep Clustering For Speech Separation
- Deep Speech Denoising with Vector Space Projections
- Generalization Challenges for Neural Architectures in Audio Source Separation
- Speaker and Direction Inferred Dual-channel Speech Separation
- Dual-Path Modeling for Long Recording Speech Separation in Meetings