activity
20202022
most citedIs Someone Speaking? Exploring Long-term Temporal Features for Audio-visual Active Speaker Detection

132 citations · 139 across the 3 of their papers we have counts for

collaborators

5 papers

cs.CV2022

VCSE: Time-Domain Visual-Contextual Speaker Extraction Network

Junjie Li, Meng Ge, Zexu Pan +2

Speaker extraction seeks to extract the target speech in a multi-talker scenario given an auxiliary reference. Such reference can be auditory, i.e., a pre-recorded speech, visual,…

eess.AS2021132 cited

Is Someone Speaking? Exploring Long-term Temporal Features for Audio-visual Active Speaker Detection

Ruijie Tao, Zexu Pan, Rohan Kumar Das +3

Active speaker detection (ASD) seeks to detect who is speaking in a visual scene of one or more speakers. The successful ASD depends on accurate interpretation of short-term and lo…

cs.SD2021

Multi-target DoA Estimation with an Audio-visual Fusion Mechanism

Xinyuan Qian, Maulik Madhavi, Zexu Pan +2

Most of the prior studies in the spatial \ac{DoA} domain focus on a single modality. However, humans use auditory and visual senses to detect the presence of sound sources. With th…

eess.AS2020

Muse: Multi-modal target speaker extraction with visual cues

Zexu Pan, Ruijie Tao, Chenglin Xu +1

Speaker extraction algorithm relies on the speech sample from the target speaker as the reference point to focus its attention. Such a reference speech is typically pre-recorded. O…

eess.AS20207 cited

Multi-modal Attention for Speech Emotion Recognition

Zexu Pan, Zhaojie Luo, Jichen Yang +1

Emotion represents an essential aspect of human speech that is manifested in speech prosody. Speech, visual, and textual cues are complementary in human communication. In this pape…