Deep Clustering and Conventional Networks for Music Separation: Stronger Together
arXiv:1611.06265 · doi:10.1109/ICASSP.2017.7952118
Abstract
Deep clustering is the first method to handle general audio separation scenarios with multiple sources of the same type and an arbitrary number of sources, performing impressively in speaker-independent speech separation tasks. However, little is known about its effectiveness in other challenging situations such as music source separation. Contrary to conventional networks that directly estimate the source signals, deep clustering generates an embedding for each time-frequency bin, and separates sources by clustering the bins in the embedding space. We show that deep clustering outperforms conventional networks on a singing voice separation task, in both matched and mismatched conditions, even though conventional networks have the advantage of end-to-end training for best signal approximation, presumably because its more flexible objective engenders better regularization. Since the strengths of deep clustering and conventional network architectures appear complementary, we explore combining them in a single hybrid network trained via an approach akin to multi-task learning. Remarkably, the combination significantly outperforms either of its components.
Published in ICASSP 2017
References in corpus (1)
Cited by in corpus (37)
- Conv-TasNet: Surpassing Ideal Time-Frequency Magnitude Masking for Speech Separation
- Speaker-independent Speech Separation with Deep Attractor Network
- Deep Filtering: Signal Extraction and Reconstruction Using Complex Time-Frequency Filters
- Encoder-Decoder Based Attractors for End-to-End Neural Diarization
- Blind Monaural Source Separation on Heart and Lung Sounds Based on Periodic-Coded Deep Autoencoder
- Singing voice separation: a study on training data
- End-to-End Speech Separation with Unfolded Iterative Phase Reconstruction
- The Cone of Silence: Speech Separation by Localization
- Finding Strength in Weakness: Learning to Separate Sounds with Weak Supervision
- Signal-Aware Direction-of-Arrival Estimation Using Attention Mechanisms
- Bootstrapping deep music separation from primitive auditory grouping principles
- SVSGAN: Singing Voice Separation via Generative Adversarial Network
- Audio Source Separation via Multi-Scale Learning with Dilated Dense U-Nets
- Onssen: an open-source speech separation and enhancement library
- Deep Attention Fusion Feature for Speech Separation with End-to-End Post-filter Method
- Closing the Training/Inference Gap for Deep Attractor Networks
- Improved Speech Separation with Time-and-Frequency Cross-domain Joint Embedding and Clustering
- Unsupervised training of a deep clustering model for multichannel blind source separation
- Model selection for deep audio source separation via clustering analysis
- A Study of Transfer Learning in Music Source Separation
- MITAS: A Compressed Time-Domain Audio Separation Network with Parameter Sharing
- Discriminative Learning for Monaural Speech Separation Using Deep Embedding Features
- SADDEL: Joint Speech Separation and Denoising Model based on Multitask Learning
- Evolving Multi-Resolution Pooling CNN for Monaural Singing Voice Separation
- Sparse Mixture of Local Experts for Efficient Speech Enhancement
- OtoWorld: Towards Learning to Separate by Learning to Move
- Distortion-controlled Training for End-to-end Reverberant Speech Separation with Auxiliary Autoencoding Loss
- Multichannel Singing Voice Separation by Deep Neural Network Informed DOA Constrained CNMF
- An Overview of Lead and Accompaniment Separation in Music
- Boosted Locality Sensitive Hashing: Discriminative Binary Codes for Source Separation
- Manifold-Aware Deep Clustering: Maximizing Angles between Embedding Vectors Based on Regular Simplex
- Deep Bayesian Unsupervised Source Separation Based on a Complex Gaussian Mixture Model
- Dual-Path Modeling for Long Recording Speech Separation in Meetings
- Audiovisual Singing Voice Separation
- Remixing Music with Visual Conditioning
- X-DC: Explainable Deep Clustering based on Learnable Spectrogram Templates
- Mask-dependent Phase Estimation for Monaural Speaker Separation