SlimDiffuSE: Towards Efficient Diffusion-Based Speech Enhancement using Slimmable Networks
arXiv:2608.21188
Abstract
Diffusion-based models are emerging in the speech enhancement domain and are achieving state-of-the-art performance across various benchmark datasets. A major downside of diffusion models is that data generation requires many evaluations of a typically large neural network, which results in high overall complexity. In this work, we propose a slimmable diffusion model that employs adaptive network widths throughout the data generation process to reduce computational cost. By using a greedy search algorithm to optimize the network width schedule, our method achieves performance comparable to baseline diffusion models with significantly reduced computational complexity. Notably, our approach reduces the computational complexity by up to without a significant drop in objective metrics, such as perceptual evaluation of speech quality (PESQ) and SI-SDR.