5 papers
Schrödinger Bridge Consistency Trajectory Models for Speech Enhancement
Shuichiro Nishigori, Koichi Saito, Naoki Murata +3
Speech enhancement (SE) utilizing diffusion models is a promising technology that improves speech quality in noisy speech data. Furthermore, the Schrödinger bridge (SB) has recentl…
The Sound Demixing Challenge 2023 $\unicode{x2013}$ Cinematic Demixing Track
Stefan Uhlich, Giorgio Fabbro, Masato Hirano +14
This paper summarizes the cinematic demixing (CDX) track of the Sound Demixing Challenge 2023 (SDX'23). We provide a comprehensive summary of the challenge setup, detailing the str…
Diffusion-Based Speech Enhancement with Joint Generative and Predictive Decoders
Hao Shi, Kazuki Shimada, Masato Hirano +6
Diffusion-based generative speech enhancement (SE) has recently received attention, but reverse diffusion remains time-consuming. One solution is to initialize the reverse diffusio…
Extending Audio Masked Autoencoders Toward Audio Restoration
Zhi Zhong, Hao Shi, Masato Hirano +5
Audio classification and restoration are among major downstream tasks in audio signal processing. However, restoration derives less of a benefit from pretrained models compared to…
Diffusion-based Signal Refiner for Speech Enhancement and Separation
Masato Hirano, Ryosuke Sawata, Naoki Murata +2
Although recent speech processing technologies have achieved significant improvements in objective metrics, there still remains a gap in human perceptual quality. This paper propos…