132 citations · 139 across the 3 of their papers we have counts for
5 papers
VCSE: Time-Domain Visual-Contextual Speaker Extraction Network
Junjie Li, Meng Ge, Zexu Pan +2
Speaker extraction seeks to extract the target speech in a multi-talker scenario given an auxiliary reference. Such reference can be auditory, i.e., a pre-recorded speech, visual,…
Is Someone Speaking? Exploring Long-term Temporal Features for Audio-visual Active Speaker Detection
Ruijie Tao, Zexu Pan, Rohan Kumar Das +3
Active speaker detection (ASD) seeks to detect who is speaking in a visual scene of one or more speakers. The successful ASD depends on accurate interpretation of short-term and lo…
Multi-target DoA Estimation with an Audio-visual Fusion Mechanism
Xinyuan Qian, Maulik Madhavi, Zexu Pan +2
Most of the prior studies in the spatial \ac{DoA} domain focus on a single modality. However, humans use auditory and visual senses to detect the presence of sound sources. With th…
Muse: Multi-modal target speaker extraction with visual cues
Zexu Pan, Ruijie Tao, Chenglin Xu +1
Speaker extraction algorithm relies on the speech sample from the target speaker as the reference point to focus its attention. Such a reference speech is typically pre-recorded. O…
Multi-modal Attention for Speech Emotion Recognition
Zexu Pan, Zhaojie Luo, Jichen Yang +1
Emotion represents an essential aspect of human speech that is manifested in speech prosody. Speech, visual, and textual cues are complementary in human communication. In this pape…