4 papers
Towards Balanced Spectral Reconstruction: Spectrally Adaptive Loss for Streaming Speech Enhancement
Haixin Zhao, Nilesh Madhu
This paper proposes two spectrally weighted STFT loss functions for lightweight streaming speech enhancement, addressing the magnitude over-attenuation in mid-to-high frequency reg…
Towards Robust Generative Speech Enhancement Using Vector Quantisation-Based Neural Audio Codec
Haixin Zhao, Nilesh Madhu
This work investigates modelling strategies in continuous and discrete latent spaces in the vector quantisation (VQ)-based neural audio codec (NAC) speech enhancement (SE), along w…
Dynamically Slimmable Speech Enhancement Network with Metric-Guided Training
Haixin Zhao, Kaixuan Yang, Nilesh Madhu
To further reduce the complexity of lightweight speech enhancement models, we introduce a gating-based Dynamically Slimmable Network (DSN). The DSN comprises static and dynamic com…
Study of Lightweight Transformer Architectures for Single-Channel Speech Enhancement
Haixin Zhao, Nilesh Madhu
In speech enhancement, achieving state-of-the-art (SotA) performance while adhering to the computational constraints on edge devices remains a formidable challenge. Networks integr…