1 citations · 1 across the 4 of their papers we have counts for
4 papers
Multi-Stage Face-Voice Association Learning with Keynote Speaker Diarization
Ruijie Tao, Zhan Shi, Yidi Jiang +4
The human brain has the capability to associate the unknown person's voice and face by leveraging their general relationship, referred to as ``cross-modal speaker verification''. T…
Target Speech Diarization with Multimodal Prompts
Yidi Jiang, Ruijie Tao, Zhengyang Chen +2
Traditional speaker diarization seeks to detect ``who spoke when'' according to speaker characteristics. Extending to target speech diarization, we detect ``when target event occur…
EEG-Derived Voice Signature for Attended Speaker Detection
Hongxu Zhu, Siqi Cai, Yidi Jiang +2
\textit{Objective:} Conventional EEG-based auditory attention detection (AAD) is achieved by comparing the time-varying speech stimuli and the elicited EEG signals. However, in ord…
Target Active Speaker Detection with Audio-visual Cues
Yidi Jiang, Ruijie Tao, Zexu Pan +1
In active speaker detection (ASD), we would like to detect whether an on-screen person is speaking based on audio-visual cues. Previous studies have primarily focused on modeling a…