most citedScenario-Aware Audio-Visual TF-GridNet for Target Speech Extraction

1 citations · 1 across the 6 of their papers we have counts for

collaborators

6 papers

eess.AS20231 cited

Scenario-Aware Audio-Visual TF-GridNet for Target Speech Extraction

Zexu Pan, Gordon Wichern, Yoshiki Masuyama +4

Target speech extraction aims to extract, based on a given conditioning cue, a target speech signal that is corrupted by interfering sources, such as noise or competing speakers. B…

cs.SD2023

LocSelect: Target Speaker Localization with an Auditory Selective Hearing Mechanism

Yu Chen, Xinyuan Qian, Zexu Pan +2

The prevailing noise-resistant and reverberation-resistant localization algorithms primarily emphasize separating and providing directional output for each speaker in multi-speaker…

eess.AS2023

Generation or Replication: Auscultating Audio Latent Diffusion Models

Dimitrios Bralios, Gordon Wichern, François G. Germain +4

The introduction of audio latent diffusion models possessing the ability to generate realistic sound clips on demand from a text description has the potential to revolutionize how…

cs.SD2023

Audio-Visual Active Speaker Extraction for Sparsely Overlapped Multi-talker Speech

Junjie Li, Ruijie Tao, Zexu Pan +3

Target speaker extraction aims to extract the speech of a specific speaker from a multi-talker mixture as specified by an auxiliary reference. Most studies focus on the scenario wh…

eess.AS2023

NeuroHeed: Neuro-Steered Speaker Extraction using EEG Signals

Zexu Pan, Marvin Borsdorf, Siqi Cai +2

Humans possess the remarkable ability to selectively attend to a single speaker amidst competing voices and background noise, known as selective auditory attention. Recent studies…

eess.AS2023

Target Active Speaker Detection with Audio-visual Cues

Yidi Jiang, Ruijie Tao, Zexu Pan +1

In active speaker detection (ASD), we would like to detect whether an on-screen person is speaking based on audio-visual cues. Previous studies have primarily focused on modeling a…