most citedScenario-Aware Audio-Visual TF-GridNet for Target Speech Extraction

1 citations · 1 across the 8 of their papers we have counts for

collaborators
Showing eess.ASShow all

7 papers · 1 filter

eess.AS2024

Task-Aware Unified Source Separation

Kohei Saijo, Janek Ebbers, François G. Germain +2

Several attempts have been made to handle multiple source separation tasks such as speech enhancement, speech separation, sound event separation, music source separation (MSS), or…

eess.AS2024

TF-Locoformer: Transformer with Local Modeling by Convolution for Speech Separation and Enhancement

Kohei Saijo, Gordon Wichern, François G. Germain +2

Time-frequency (TF) domain dual-path models achieve high-fidelity speech separation. While some previous state-of-the-art (SoTA) models rely on RNNs, this reliance means they lack…

eess.AS2024

Enhanced Reverberation as Supervision for Unsupervised Speech Separation

Kohei Saijo, Gordon Wichern, François G. Germain +2

Reverberation as supervision (RAS) is a framework that allows for training monaural speech separation models from multi-channel mixtures in an unsupervised manner. In RAS, models a…

eess.AS2024

Sound Event Bounding Boxes

Janek Ebbers, Francois G. Germain, Gordon Wichern +1

Sound event detection is the task of recognizing sounds and determining their extent (onset/offset times) within an audio clip. Existing systems commonly predict sound presence con…

eess.AS2024

Why does music source separation benefit from cacophony?

Chang-Bin Jeon, Gordon Wichern, François G. Germain +1

In music source separation, a standard training data augmentation procedure is to create new training samples by randomly combining instrument stems from different songs. These ran…

eess.AS20231 cited

Scenario-Aware Audio-Visual TF-GridNet for Target Speech Extraction

Zexu Pan, Gordon Wichern, Yoshiki Masuyama +4

Target speech extraction aims to extract, based on a given conditioning cue, a target speech signal that is corrupted by interfering sources, such as noise or competing speakers. B…