9 papers
SURF: Separation via Unsupervised Remixing Flow
Henry Li, Robin Scheibler, Efthymios Tzinis +3
The goal of single-channel source separation is to reconstruct sources given their mixture. In supervised settings where vast amounts of clean source data are available, this c…
SequenceLayers: Sequence Processing and Streaming Neural Networks Made Easy
RJ Skerry-Ryan, Julian Salazar, Soroosh Mariooryad +8
We introduce a neural network layer API and library for sequence modeling, designed for easy creation of sequence models that can be executed both layer-by-layer (e.g., teacher-for…
P.808 Multilingual Speech Enhancement Testing: Approach and Results of URGENT 2025 Challenge
Marvin Sach, Yihui Fu, Kohei Saijo +9
In speech quality estimation for speech enhancement (SE) systems, subjective listening tests so far are considered as the gold standard. This should be even more true considering t…
Miipher-2: A Universal Speech Restoration Model for Million-Hour Scale Data Restoration
Shigeki Karita, Yuma Koizumi, Heiga Zen +3
Training data cleaning is a new application for generative model-based speech restoration (SR). This paper introduces Miipher-2, an SR model designed for million-hour scale data, f…
Source Separation by Flow Matching
Robin Scheibler, John R. Hershey, Arnaud Doucet +1
We consider the problem of single-channel audio source separation with the goal of reconstructing sources from their mixture. We address this ill-posed problem with FLOSS (FLOw…
ReverbMiipher: Generative Speech Restoration meets Reverberation Characteristics Controllability
Wataru Nakata, Yuma Koizumi, Shigeki Karita +5
Reverberation encodes spatial information regarding the acoustic source environment, yet traditional Speech Restoration (SR) usually completely removes reverberation. We propose Re…