3 papers
cs.SD2025
Shortcut Flow Matching for Speech Enhancement: Step-Invariant flows via single stage training
Naisong Zhou, Saisamarth Rajesh Phaye, Milos Cernak +4
Diffusion-based generative models have achieved state-of-the-art performance for perceptual quality in speech enhancement (SE). However, their iterative nature requires numerous Ne…
eess.AS2025
DeepFilterGAN: A Full-band Real-time Speech Enhancement System with GAN-based Stochastic Regeneration
Sanberk Serbest, Tijana Stojkovic, Milos Cernak +1
In this work, we propose a full-band real-time speech enhancement system with GAN-based stochastic regeneration. Predictive models focus on estimating the mean of the target distri…
cs.SD2025
Model as Loss: A Self-Consistent Training Paradigm
Saisamarth Rajesh Phaye, Milos Cernak, Andrew Harper
Conventional methods for speech enhancement rely on handcrafted loss functions (e.g., time or frequency domain losses) or deep feature losses (e.g., using WavLM or wav2vec), which…