8 papers
Rethinking Language Model-Based Generative Speech Enhancement in the Latent Space of a Neural Audio Codec
Yihui Fu, Zhengyang Li, Tim Fingscheidt
Language model (LM)-based speech enhancement (SE) has recently emerged rapidly using latent space features of neural audio codecs (NACs). In this paper, first, we present a unified…
DisContSE: Single-Step Diffusion Speech Enhancement Based on Joint Discrete and Continuous Embeddings
Yihui Fu, Tim Fingscheidt
Diffusion speech enhancement on discrete audio codec features gain immense attention due to their improved speech component reconstruction capability. However, they usually suffer…
UrgentMOS: Unified Multi-Metric and Preference Learning for Robust Speech Quality Assessment
Wei Wang, Wangyou Zhang, Chenda Li +12
Automatic speech quality assessment has become increasingly important as modern speech generation systems continue to advance, while human listening tests remain costly, time-consu…
ICASSP 2026 URGENT Speech Enhancement Challenge
Chenda Li, Wei Wang, Marvin Sach +8
The ICASSP 2026 URGENT Challenge advances the series by focusing on universal speech enhancement (SE) systems that handle diverse distortions, domains, and input conditions. This o…
P.808 Multilingual Speech Enhancement Testing: Approach and Results of URGENT 2025 Challenge
Marvin Sach, Yihui Fu, Kohei Saijo +9
In speech quality estimation for speech enhancement (SE) systems, subjective listening tests so far are considered as the gold standard. This should be even more true considering t…
URGENT-PK: Perceptually-Aligned Ranking Model Designed for Speech Enhancement Competition
Jiahe Wang, Chenda Li, Wei Wang +11
The Mean Opinion Score (MOS) is fundamental to speech quality assessment. However, its acquisition requires significant human annotation. Although deep neural network approaches, s…