Hybrid Spectrogram and Waveform Source Separation
arXiv:2111.03600
Abstract
Source separation models either work on the spectrogram or waveform domain. In this work, we show how to perform end-to-end hybrid source separation, letting the model decide which domain is best suited for each source, and even combining both. The proposed hybrid version of the Demucs architecture won the Music Demixing Challenge 2021 organized by Sony. This architecture also comes with additional improvements, such as compressed residual branches, local attention or singular value regularization. Overall, a 1.4 dB improvement of the Signal-To-Distortion (SDR) was observed across all sources as measured on the MusDB HQ dataset, an improvement confirmed by human subjective evaluation, with an overall quality rated at 2.83 out of 5 (2.36 for the non hybrid Demucs), and absence of contamination at 3.04 (against 2.37 for the non hybrid Demucs and 2.44 for the second ranking model submitted at the competition).
ISMIR 2021 MDX Workshop, 11 pages, 2 figures
References in corpus (4)
Cited by in corpus (10)
- Music Demixing Challenge 2021
- The Sound Demixing Challenge 2023 $\unicode{x2013}$ Music Demixing Track
- Music Source Separation Based on a Lightweight Deep Learning Framework (DTTNET: DUAL-PATH TFC-TDF UNET)
- JaCappella Corpus: A Japanese a Cappella Vocal Ensemble Corpus
- Music Source Separation with Generative Flow
- A Generalized Bandsplit Neural Network for Cinematic Audio Source Separation
- The ICASSP SP Cadenza Challenge: Music Demixing/Remixing for Hearing Aids
- Quantifying Spatial Audio Quality Impairment
- SynthSOD: Developing an Heterogeneous Dataset for Orchestra Music Source Separation
- DJCM: A Deep Joint Cascade Model for Singing Voice Separation and Vocal Pitch Estimation