most citedScenario-Aware Audio-Visual TF-GridNet for Target Speech Extraction

1 citations · 1 across the 4 of their papers we have counts for

collaborators

6 papers

eess.AS2024

TF-Locoformer: Transformer with Local Modeling by Convolution for Speech Separation and Enhancement

Kohei Saijo, Gordon Wichern, François G. Germain +2

Time-frequency (TF) domain dual-path models achieve high-fidelity speech separation. While some previous state-of-the-art (SoTA) models rely on RNNs, this reliance means they lack…

eess.AS2024

Enhanced Reverberation as Supervision for Unsupervised Speech Separation

Kohei Saijo, Gordon Wichern, François G. Germain +2

Reverberation as supervision (RAS) is a framework that allows for training monaural speech separation models from multi-channel mixtures in an unsupervised manner. In RAS, models a…

eess.AS2024

NIIRF: Neural IIR Filter Field for HRTF Upsampling and Personalization

Yoshiki Masuyama, Gordon Wichern, François G. Germain +4

Head-related transfer functions (HRTFs) are important for immersive audio, and their spatial interpolation has been studied to upsample finite measurements. Recently, neural fields…

eess.AS2023

NeuroHeed+: Improving Neuro-steered Speaker Extraction with Joint Auditory Attention Detection

Zexu Pan, Gordon Wichern, Francois G. Germain +2

Neuro-steered speaker extraction aims to extract the listener's brain-attended speech signal from a multi-talker speech signal, in which the attention is derived from the cortical…

eess.AS20231 cited

Scenario-Aware Audio-Visual TF-GridNet for Target Speech Extraction

Zexu Pan, Gordon Wichern, Yoshiki Masuyama +4

Target speech extraction aims to extract, based on a given conditioning cue, a target speech signal that is corrupted by interfering sources, such as noise or competing speakers. B…

eess.AS2023

Generation or Replication: Auscultating Audio Latent Diffusion Models

Dimitrios Bralios, Gordon Wichern, François G. Germain +4

The introduction of audio latent diffusion models possessing the ability to generate realistic sound clips on demand from a text description has the potential to revolutionize how…