Sudo rm -rf: Efficient Networks for Universal Audio Source Separation
arXiv:2007.06833 · doi:10.1109/MLSP49062.2020.9231900
Abstract
In this paper, we present an efficient neural network for end-to-end general purpose audio source separation. Specifically, the backbone structure of this convolutional network is the SUccessive DOwnsampling and Resampling of Multi-Resolution Features (SuDoRMRF) as well as their aggregation which is performed through simple one-dimensional convolutions. In this way, we are able to obtain high quality audio source separation with limited number of floating point operations, memory requirements, number of parameters and latency. Our experiments on both speech and environmental sound separation datasets show that SuDoRMRF performs comparably and even surpasses various state-of-the-art approaches with significantly higher computational resource requirements.
accepted to MLSP 2020
Cited by in corpus (15)
- Group Communication with Context Codec for Lightweight Source Separation
- Objective and subjective evaluation of speech enhancement methods in the UDASE task of the 7th CHiME challenge
- Separate but Together: Unsupervised Federated Learning for Speech Enhancement from Non-IID Data
- Continual self-training with bootstrapped remixing for speech enhancement
- DeFTAN-II: Efficient Multichannel Speech Enhancement with Subgroup Processing
- RF Challenge: The Data-Driven Radio Frequency Signal Separation Challenge
- Stepwise-Refining Speech Separation Network via Fine-Grained Encoding in High-order Latent Domain
- Time-Domain Speech Extraction with Spatial Information and Multi Speaker Conditioning Mechanism
- MixCycle: Unsupervised Speech Separation via Cyclic Mixture Permutation Invariant Training
- RemixIT: Continual self-training of speech enhancement models via bootstrapped remixing
- Exploiting Temporal Structures of Cyclostationary Signals for Data-Driven Single-Channel Source Separation
- On Neural Architectures for Deep Learning-based Source Separation of Co-Channel OFDM Signals
- TDFNet: An Efficient Audio-Visual Speech Separation Model with Top-down Fusion
- Unified Gradient Reweighting for Model Biasing with Applications to Source Separation
- A Study of the Scale Invariant Signal to Distortion Ratio in Speech Separation with Noisy References