most citedScenario-Aware Audio-Visual TF-GridNet for Target Speech Extraction

1 citations · 1 across the 3 of their papers we have counts for

collaborators

9 papers

eess.AS2025

Factorized RVQ-GAN For Disentangled Speech Tokenization

Sameer Khurana, Dominik Klement, Antoine Laurent +13

We propose Hierarchical Audio Codec (HAC), a unified neural speech codec that factorizes its bottleneck into three linguistic levels-acoustic, phonetic, and lexical-within a single…

eess.AS2024

Task-Aware Unified Source Separation

Kohei Saijo, Janek Ebbers, François G. Germain +2

Several attempts have been made to handle multiple source separation tasks such as speech enhancement, speech separation, sound event separation, music source separation (MSS), or…

eess.AS2024

TF-Locoformer: Transformer with Local Modeling by Convolution for Speech Separation and Enhancement

Kohei Saijo, Gordon Wichern, François G. Germain +2

Time-frequency (TF) domain dual-path models achieve high-fidelity speech separation. While some previous state-of-the-art (SoTA) models rely on RNNs, this reliance means they lack…

eess.AS2024

Enhanced Reverberation as Supervision for Unsupervised Speech Separation

Kohei Saijo, Gordon Wichern, François G. Germain +2

Reverberation as supervision (RAS) is a framework that allows for training monaural speech separation models from multi-channel mixtures in an unsupervised manner. In RAS, models a…

eess.AS2024

Sound Event Bounding Boxes

Janek Ebbers, Francois G. Germain, Gordon Wichern +1

Sound event detection is the task of recognizing sounds and determining their extent (onset/offset times) within an audio clip. Existing systems commonly predict sound presence con…

eess.AS2024

Why does music source separation benefit from cacophony?

Chang-Bin Jeon, Gordon Wichern, François G. Germain +1

In music source separation, a standard training data augmentation procedure is to create new training samples by randomly combining instrument stems from different songs. These ran…