18 citations · 25 across the 7 of their papers we have counts for
Showing cs.SDShow all
2 papers · 1 filter
cs.SD2023★ 1 cited
PIAVE: A Pose-Invariant Audio-Visual Speaker Extraction Network
Qinghua Liu, Meng Ge, Zhizheng Wu +1
It is common in everyday spoken communication that we look at the turning head of a talker to listen to his/her voice. Humans see the talker to listen better, so do machines. Howev…
cs.SD2023
Rethinking the visual cues in audio-visual speaker extraction
Junjie Li, Meng Ge, Zexu pan +4
The Audio-Visual Speaker Extraction (AVSE) algorithm employs parallel video recording to leverage two visual cues, namely speaker identity and synchronization, to enhance performan…