1 citations · 1 across the 6 of their papers we have counts for
6 papers
Scenario-Aware Audio-Visual TF-GridNet for Target Speech Extraction
Zexu Pan, Gordon Wichern, Yoshiki Masuyama +4
Target speech extraction aims to extract, based on a given conditioning cue, a target speech signal that is corrupted by interfering sources, such as noise or competing speakers. B…
LocSelect: Target Speaker Localization with an Auditory Selective Hearing Mechanism
Yu Chen, Xinyuan Qian, Zexu Pan +2
The prevailing noise-resistant and reverberation-resistant localization algorithms primarily emphasize separating and providing directional output for each speaker in multi-speaker…
Generation or Replication: Auscultating Audio Latent Diffusion Models
Dimitrios Bralios, Gordon Wichern, François G. Germain +4
The introduction of audio latent diffusion models possessing the ability to generate realistic sound clips on demand from a text description has the potential to revolutionize how…
Audio-Visual Active Speaker Extraction for Sparsely Overlapped Multi-talker Speech
Junjie Li, Ruijie Tao, Zexu Pan +3
Target speaker extraction aims to extract the speech of a specific speaker from a multi-talker mixture as specified by an auxiliary reference. Most studies focus on the scenario wh…
NeuroHeed: Neuro-Steered Speaker Extraction using EEG Signals
Zexu Pan, Marvin Borsdorf, Siqi Cai +2
Humans possess the remarkable ability to selectively attend to a single speaker amidst competing voices and background noise, known as selective auditory attention. Recent studies…
Target Active Speaker Detection with Audio-visual Cues
Yidi Jiang, Ruijie Tao, Zexu Pan +1
In active speaker detection (ASD), we would like to detect whether an on-screen person is speaking based on audio-visual cues. Previous studies have primarily focused on modeling a…