6 papers · 1 filter
TF-Locoformer: Transformer with Local Modeling by Convolution for Speech Separation and Enhancement
Kohei Saijo, Gordon Wichern, François G. Germain +2
Time-frequency (TF) domain dual-path models achieve high-fidelity speech separation. While some previous state-of-the-art (SoTA) models rely on RNNs, this reliance means they lack…
Enhanced Reverberation as Supervision for Unsupervised Speech Separation
Kohei Saijo, Gordon Wichern, François G. Germain +2
Reverberation as supervision (RAS) is a framework that allows for training monaural speech separation models from multi-channel mixtures in an unsupervised manner. In RAS, models a…
Scenario-Aware Audio-Visual TF-GridNet for Target Speech Extraction
Zexu Pan, Gordon Wichern, Yoshiki Masuyama +4
Target speech extraction aims to extract, based on a given conditioning cue, a target speech signal that is corrupted by interfering sources, such as noise or competing speakers. B…
Generation or Replication: Auscultating Audio Latent Diffusion Models
Dimitrios Bralios, Gordon Wichern, François G. Germain +4
The introduction of audio latent diffusion models possessing the ability to generate realistic sound clips on demand from a text description has the potential to revolutionize how…
NeuroHeed: Neuro-Steered Speaker Extraction using EEG Signals
Zexu Pan, Marvin Borsdorf, Siqi Cai +2
Humans possess the remarkable ability to selectively attend to a single speaker amidst competing voices and background noise, known as selective auditory attention. Recent studies…
Target Active Speaker Detection with Audio-visual Cues
Yidi Jiang, Ruijie Tao, Zexu Pan +1
In active speaker detection (ASD), we would like to detect whether an on-screen person is speaking based on audio-visual cues. Previous studies have primarily focused on modeling a…