13 papers
SlimDiffuSE: Towards Efficient Diffusion-Based Speech Enhancement using Slimmable Networks
Nagashree K. S. Rao, Shrishti Saha Shetu, Mohamed Elminshawi +2
Diffusion-based models are emerging in the speech enhancement domain and are achieving state-of-the-art performance across various benchmark datasets. A major downside of diffusion…
μNet: Ultra-Low-Memory and Low-Complexity Speech Enhancement for Embedded Digital Signal Processors
Shrishti Saha Shetu, Jose Miguel Martinez Aponte, Nagashree K. S. Rao +3
Speech enhancement on embedded digital signal processors (DSPs) imposes strict constraints on memory footprint, computational complexity, latency, and support for integer operation…
Revisiting Vocos: That Phasiness Business in Time-Frequency Neural Vocoding
Ãnal Ege GaznepoÄlu, Frank Zalkow, Mohammad Joshaghani +3
Recently, time-frequency neural vocoders have been approaching the state-of-the-art quality of time-domain neural vocoders. Vocos is a notable example due to its efficiency, but it…
Quantifying the Uncertainty of Blindly Estimated Room Embeddings Using a Dispersion-Calibrated Score
Yang Xiang, Philipp Götz, Emanuël A. P. Habets +3
Room embeddings derived from reverberant speech are often unreliable: speech content and recording degradation can alter the representation even when speaker, room, and source-rece…
Beyond Cross-Reconstruction: Probing-Based Disentanglement Evaluation for Acoustic Teleportation Codecs
Philipp Grundhuber, Emanuël A. P. Habets
Some neural audio codecs disentangle speech into latent subspaces encoding content, speaker identity, and acoustics, enabling acoustic teleportation and voice conversion. Existing…
Adaptive Regularization for Sparsity Control in Bregman-Based Optimizers
Ahmad Aloradi, Tim Roith, Emanuël A. P. Habets +1
Sparse training reduces the memory and computational costs of deep neural networks. However, sparse optimization methods, e.g., those adding an penalty, often control spar…