Music Source Separation in the Waveform Domain
arXiv:1911.13254
Abstract
Source separation for music is the task of isolating contributions, or stems, from different instruments recorded individually and arranged together to form a song. Such components include voice, bass, drums and any other accompaniments.Contrarily to many audio synthesis tasks where the best performances are achieved by models that directly generate the waveform, the state-of-the-art in source separation for music is to compute masks on the magnitude spectrum. In this paper, we compare two waveform domain architectures. We first adapt Conv-Tasnet, initially developed for speech source separation,to the task of music source separation. While Conv-Tasnet beats many existing spectrogram-domain methods, it suffersfrom significant artifacts, as shown by human evaluations. We propose instead Demucs, a novel waveform-to-waveform model,with a U-Net structure and bidirectional LSTM.Experiments on the MusDB dataset show that, with proper data augmentation, Demucs beats allexisting state-of-the-art architectures, including Conv-Tasnet, with 6.3 SDR on average, (and up to 6.8 with 150 extra training songs, even surpassing the IRM oracle for the bass source).Using recent development in model quantization, Demucs can be compressed down to 120MBwithout any loss of accuracy.We also provide human evaluations, showing that Demucs benefit from a large advantagein terms of the naturalness of the audio. However, it suffers from some bleeding,especially between the vocals and other source.
References in corpus (5)
Cited by in corpus (38)
- Voice Separation with an Unknown Number of Multiple Speakers
- Hybrid Spectrogram and Waveform Source Separation
- Music Demixing Challenge 2021
- D3Net: Densely connected multidilated DenseNet for music source separation
- Compute and memory efficient universal sound source separation
- KUIELab-MDX-Net: A Two-Stream Neural Network for Music Demixing
- RMVPE: A Robust Model for Vocal Pitch Estimation in Polyphonic Music
- Differentiable Model Compression via Pseudo Quantization Noise
- Music Source Separation Based on a Lightweight Deep Learning Framework (DTTNET: DUAL-PATH TFC-TDF UNET)
- Source Separation with Deep Generative Priors
- Points2Sound: From mono to binaural audio using 3D point cloud scenes
- Music Source Separation with Generative Flow
- Multichannel-based learning for audio object extraction
- Deep Audio Waveform Prior
- Densely connected multidilated convolutional networks for dense prediction tasks
- A Generalized Bandsplit Neural Network for Cinematic Audio Source Separation
- SynthSOD: Developing an Heterogeneous Dataset for Orchestra Music Source Separation
- Multitask learning for instrument activation aware music source separation
- Exploring Aligned Lyrics-Informed Singing Voice Separation
- Parallel and Flexible Sampling from Autoregressive Models via Langevin Dynamics
- Revisiting Representation Learning for Singing Voice Separation with Sinkhorn Distances
- Adversarial attacks on audio source separation
- Sams-Net: A Sliced Attention-based Neural Network for Music Source Separation
- Unified Gradient Reweighting for Model Biasing with Applications to Source Separation
- Unsupervised Source Separation via Bayesian Inference in the Latent Domain
- Continuous Monitoring of Blood Pressure with Evidential Regression
- Semi-Supervised Singing Voice Separation with Noisy Self-Training
- Complex ratio masking for singing voice separation
- Multi-Task Audio Source Separation
- DJCM: A Deep Joint Cascade Model for Singing Voice Separation and Vocal Pitch Estimation
- LightSAFT: Lightweight Latent Source Aware Frequency Transform for Source Separation
- Source Mixing and Separation Robust Audio Steganography
- Toward Expressive Singing Voice Correction: On Perceptual Validity of Evaluation Metrics for Vocal Melody Extraction
- End-to-end Music-mixed Speech Recognition
- Learning source-aware representations of music in a discrete latent space
- How much to Dereverberate? Low-Latency Single-Channel Speech Enhancement in Distant Microphone Scenarios
- Unsupervised Source Separation By Steering Pretrained Music Models
- Deep Learning Tools for Audacity: Helping Researchers Expand the Artist's Toolkit