Joint Optimization of Masks and Deep Recurrent Neural Networks for Monaural Source Separation
arXiv:1502.04149 · doi:10.1109/TASLP.2015.2468583
Abstract
Monaural source separation is important for many real world applications. It is challenging because, with only a single channel of information available, without any constraints, an infinite number of solutions are possible. In this paper, we explore joint optimization of masking functions and deep recurrent neural networks for monaural source separation tasks, including monaural speech separation, monaural singing voice separation, and speech denoising. The joint optimization of the deep recurrent neural networks with an extra masking layer enforces a reconstruction constraint. Moreover, we explore a discriminative criterion for training neural networks to further enhance the separation performance. We evaluate the proposed system on the TSP, MIR-1K, and TIMIT datasets for speech separation, singing voice separation, and speech denoising tasks, respectively. Our approaches achieve 2.30--4.98 dB SDR gain compared to NMF models in the speech separation task, 2.30--2.48 dB GNSDR gain and 4.32--5.42 dB GSIR gain compared to existing models in the singing voice separation task, and outperform NMF and DNN baselines in the speech denoising task.
Cited by in corpus (50)
- Deep Learning for Audio Signal Processing
- Speaker-independent Speech Separation with Deep Attractor Network
- SpEx: Multi-Scale Time Domain Speaker Extraction Network
- Two-Step Sound Source Separation: Training on Learned Latent Targets
- Deep Learning for Environmentally Robust Speech Recognition: An Overview of Recent Developments
- Supervised Speech Separation Based on Deep Learning: An Overview
- Blind Monaural Source Separation on Heart and Lung Sounds Based on Periodic-Coded Deep Autoencoder
- Audio Source Separation Using Variational Autoencoders and Weak Class Supervision
- Efficient Personalized Speech Enhancement through Self-Supervised Learning
- End-to-End Waveform Utterance Enhancement for Direct Evaluation Metrics Optimization by Fully Convolutional Neural Networks
- Guided Variational Autoencoder for Speech Enhancement With a Supervised Classifier
- Incremental Binarization On Recurrent Neural Networks For Single-Channel Source Separation
- Deep Extractor Network for Target Speaker Recovery From Single Channel Speech Mixtures
- RF Challenge: The Data-Driven Radio Frequency Signal Separation Challenge
- Co-Separating Sounds of Visual Objects
- Source Separation with Deep Generative Priors
- Modeling the Comb Filter Effect and Interaural Coherence for Binaural Source Separation
- End-to-end Networks for Supervised Single-channel Speech Separation
- CatNet: music source separation system with mix-audio augmentation
- Data-Driven Blind Synchronization and Interference Rejection for Digital Communication Signals
- Exploiting Temporal Structures of Cyclostationary Signals for Data-Driven Single-Channel Source Separation
- WildMix Dataset and Spectro-Temporal Transformer Model for Monoaural Audio Source Separation
- Joint Separation and Denoising of Noisy Multi-talker Speech using Recurrent Neural Networks and Permutation Invariant Training
- Sparse Pursuit and Dictionary Learning for Blind Source Separation in Polyphonic Music Recordings
- Deep RNN Framework for Visual Sequential Applications
- Improving Source Separation via Multi-Speaker Representations
- Feature Binding with Category-Dependant MixUp for Semantic Segmentation and Adversarial Robustness
- Cyclic Co-Learning of Sounding Object Visual Grounding and Sound Separation
- Investigating Cross-Domain Losses for Speech Enhancement
- Does Phase Matter For Monaural Source Separation?
- On Psychoacoustically Weighted Cost Functions Towards Resource-Efficient Deep Neural Networks for Speech Denoising
- Recent Progresses in Deep Learning based Acoustic Models (Updated)
- Optimisation and Performance Computation of a Phase Frequency Detector Module for IoT Devices
- Personalized Speech Enhancement through Self-Supervised Data Augmentation and Purification
- MITAS: A Compressed Time-Domain Audio Separation Network with Parameter Sharing
- Evolving Multi-Resolution Pooling CNN for Monaural Singing Voice Separation
- An Overview of Lead and Accompaniment Separation in Music
- Dereverberation using joint estimation of dry speech signal and acoustic system
- A Speech Enhancement Algorithm based on Non-negative Hidden Markov Model and Kullback-Leibler Divergence
- The Performance Evaluation of Attention-Based Neural ASR under Mixed Speech Input
- Deep Ad-hoc Beamforming
- Towards Automated Single Channel Source Separation using Neural Networks
- Backpropagation with N-D Vector-Valued Neurons Using Arbitrary Bilinear Products
- Discriminative Enhancement for Single Channel Audio Source Separation using Deep Neural Networks
- Identify Speakers in Cocktail Parties with End-to-End Attention
- Source separation with weakly labelled data: An approach to computational auditory scene analysis
- Collaborative Deep Learning for Speech Enhancement: A Run-Time Model Selection Method Using Autoencoders
- Music Signal Processing Using Vector Product Neural Networks
- Bitwise Source Separation on Hashed Spectra: An Efficient Posterior Estimation Scheme Using Partial Rank Order Metrics
- Multi-talker Speech Separation with Utterance-level Permutation Invariant Training of Deep Recurrent Neural Networks