Speaker-independent Speech Separation with Deep Attractor Network
arXiv:1707.03634 · doi:10.1109/TASLP.2018.2795749
Abstract
Despite the recent success of deep learning for many speech processing tasks, single-microphone, speaker-independent speech separation remains challenging for two main reasons. The first reason is the arbitrary order of the target and masker speakers in the mixture permutation problem, and the second is the unknown number of speakers in the mixture output dimension problem. We propose a novel deep learning framework for speech separation that addresses both of these issues. We use a neural network to project the time-frequency representation of the mixture signal into a high-dimensional embedding space. A reference point attractor is created in the embedding space to represent each speaker which is defined as the centroid of the speaker in the embedding space. The time-frequency embeddings of each speaker are then forced to cluster around the corresponding attractor point which is used to determine the time-frequency assignment of the speaker. We propose three methods for finding the attractors for each source in the embedding space and compare their advantages and limitations. The objective function for the network is standard signal reconstruction error which enables end-to-end operation during both training and test phases. We evaluated our system using the Wall Street Journal dataset WSJ0 on two and three speaker mixtures and report comparable or better performance than other state-of-the-art deep learning methods for speech separation.
IEEE/ACM Transactions on Audio, Speech and Language Processing (TASLP), Volume 26 Issue 4, April 2018, Page 787-796
References in corpus (3)
Cited by in corpus (63)
- Conv-TasNet: Surpassing Ideal Time-Frequency Magnitude Masking for Speech Separation
- Ablation Studies in Artificial Neural Networks
- SpEx: Multi-Scale Time Domain Speaker Extraction Network
- Voice Separation with an Unknown Number of Multiple Speakers
- End-to-End Multi-Channel Speech Separation
- Improving Universal Sound Separation Using Sound Classification
- FaceFilter: Audio-visual speech separation using still images
- Audio-Visual Speech Separation and Dereverberation with a Two-Stage Multimodal Network
- Encoder-Decoder Based Attractors for End-to-End Neural Diarization
- Dual-path RNN: efficient long sequence modeling for time-domain single-channel speech separation
- Neural Speaker Diarization with Speaker-Wise Chain Rule
- TasNet: time-domain audio separation network for real-time, single-channel speech separation
- Dual-Path Transformer Network: Direct Context-Aware Modeling for End-to-End Monaural Speech Separation
- End-to-End Speech Separation with Unfolded Iterative Phase Reconstruction
- FaSNet: Low-latency Adaptive Beamforming for Multi-microphone Audio Processing
- Group Communication with Context Codec for Lightweight Source Separation
- Event-Independent Network for Polyphonic Sound Event Localization and Detection
- SpEx+: A Complete Time Domain Speaker Extraction Network
- Deep Audio-Visual Learning: A Survey
- FurcaNet: An end-to-end deep gated convolutional, long short-term memory, deep neural networks for single channel speech separation
- End-to-End Speaker Diarization for an Unknown Number of Speakers with Encoder-Decoder Based Attractors
- Stepwise-Refining Speech Separation Network via Fine-Grained Encoding in High-order Latent Domain
- Exploring the time-domain deep attractor network with two-stream architectures in a reverberant environment
- End-to-end Networks for Supervised Single-channel Speech Separation
- Deep Learning Based Phase Reconstruction for Speaker Separation: A Trigonometric Perspective
- LaFurca: Iterative Refined Speech Separation Based on Context-Aware Dual-Path Parallel Bi-LSTM
- Recursive speech separation for unknown number of speakers
- Continuous Speech Separation Using Speaker Inventory for Long Multi-talker Recording
- Improved Source Counting and Separation for Monaural Mixture
- Deep Attention Fusion Feature for Speech Separation with End-to-End Post-filter Method
- Divide and Conquer: A Deep CASA Approach to Talker-independent Monaural Speaker Separation
- Building Corpora for Single-Channel Speech Separation Across Multiple Domains
- Speaker-Conditional Chain Model for Speech Separation and Extraction
- Sequence to Multi-Sequence Learning via Conditional Chain Mapping for Mixture Signals
- Efficient Integration of Multi-channel Information for Speaker-independent Speech Separation
- Cracking the cocktail party problem by multi-beam deep attractor network
- Multi-path RNN for hierarchical modeling of long sequential data and its application to speaker stream separation
- Unsupervised Speaker Diarization in Distributed IoT Networks Using Federated Learning
- Ultra-Lightweight Speech Separation via Group Communication
- Closing the Training/Inference Gap for Deep Attractor Networks
- Improved Speech Separation with Time-and-Frequency Cross-domain Joint Embedding and Clustering
- Mixup-breakdown: a consistency training method for improving generalization of speech separation models
- MITAS: A Compressed Time-Domain Audio Separation Network with Parameter Sharing
- CNN-LSTM models for Multi-Speaker Source Separation using Bayesian Hyper Parameter Optimization
- SADDEL: Joint Speech Separation and Denoising Model based on Multitask Learning
- Improving Neural Diarization through Speaker Attribute Attractors and Local Dependency Modeling
- Speaker Separation Using Speaker Inventories and Estimated Speech
- Supervised Speaker Embedding De-Mixing in Two-Speaker Environment
- Unsupervised Training for Deep Speech Source Separation with Kullback-Leibler Divergence Based Probabilistic Loss Function
- Speech Separation Based on Multi-Stage Elaborated Dual-Path Deep BiLSTM with Auxiliary Identity Loss
- Multi-channel Speech Separation Using Deep Embedding Model with Multilayer Bootstrap Networks
- Distortionless Multi-Channel Target Speech Enhancement for Overlapped Speech Recognition
- Guided Training: A Simple Method for Single-channel Speaker Separation
- TriBERT: Full-body Human-centric Audio-visual Representation Learning for Visual Sound Separation
- Investigation of Practical Aspects of Single Channel Speech Separation for ASR
- Lightweight Dual-channel Target Speaker Separation for Mobile Voice Communication
- A Study of the Scale Invariant Signal to Distortion Ratio in Speech Separation with Noisy References
- Manifold-Aware Deep Clustering: Maximizing Angles between Embedding Vectors Based on Regular Simplex
- Toward Speech Separation in The Pre-Cocktail Party Problem with TasTas
- Stabilizing Label Assignment for Speech Separation by Self-supervised Pre-training
- Overlapped speech recognition from a jointly learned multi-channel neural speech extraction and representation
- Improved Speaker-Dependent Separation for CHiME-5 Challenge
- Dual-Path Modeling for Long Recording Speech Separation in Meetings