1 citations · 1 across the 4 of their papers we have counts for
6 papers
TF-Locoformer: Transformer with Local Modeling by Convolution for Speech Separation and Enhancement
Kohei Saijo, Gordon Wichern, François G. Germain +2
Time-frequency (TF) domain dual-path models achieve high-fidelity speech separation. While some previous state-of-the-art (SoTA) models rely on RNNs, this reliance means they lack…
Enhanced Reverberation as Supervision for Unsupervised Speech Separation
Kohei Saijo, Gordon Wichern, François G. Germain +2
Reverberation as supervision (RAS) is a framework that allows for training monaural speech separation models from multi-channel mixtures in an unsupervised manner. In RAS, models a…
NIIRF: Neural IIR Filter Field for HRTF Upsampling and Personalization
Yoshiki Masuyama, Gordon Wichern, François G. Germain +4
Head-related transfer functions (HRTFs) are important for immersive audio, and their spatial interpolation has been studied to upsample finite measurements. Recently, neural fields…
NeuroHeed+: Improving Neuro-steered Speaker Extraction with Joint Auditory Attention Detection
Zexu Pan, Gordon Wichern, Francois G. Germain +2
Neuro-steered speaker extraction aims to extract the listener's brain-attended speech signal from a multi-talker speech signal, in which the attention is derived from the cortical…
Scenario-Aware Audio-Visual TF-GridNet for Target Speech Extraction
Zexu Pan, Gordon Wichern, Yoshiki Masuyama +4
Target speech extraction aims to extract, based on a given conditioning cue, a target speech signal that is corrupted by interfering sources, such as noise or competing speakers. B…
Generation or Replication: Auscultating Audio Latent Diffusion Models
Dimitrios Bralios, Gordon Wichern, François G. Germain +4
The introduction of audio latent diffusion models possessing the ability to generate realistic sound clips on demand from a text description has the potential to revolutionize how…