3 papers
eess.AS2024
Enhancing Real-World Active Speaker Detection with Multi-Modal Extraction Pre-Training
Ruijie Tao, Xinyuan Qian, Rohan Kumar Das +3
Audio-visual active speaker detection (AV-ASD) aims to identify which visible face is speaking in a scene with one or more persons. Most existing AV-ASD methods prioritize capturin…
cs.SD2023
Audio-Visual Active Speaker Extraction for Sparsely Overlapped Multi-talker Speech
Junjie Li, Ruijie Tao, Zexu Pan +3
Target speaker extraction aims to extract the speech of a specific speaker from a multi-talker mixture as specified by an auxiliary reference. Most studies focus on the scenario wh…
eess.AS2023
Target Active Speaker Detection with Audio-visual Cues
Yidi Jiang, Ruijie Tao, Zexu Pan +1
In active speaker detection (ASD), we would like to detect whether an on-screen person is speaking based on audio-visual cues. Previous studies have primarily focused on modeling a…