1 citations · 1 across the 2 of their papers we have counts for
5 papers
Leveraging Language Information for Target Language Extraction
Mehmet Sinan Yıldırım, Ruijie Tao, Wupeng Wang +2
Target Language Extraction aims to extract speech in a specific language from a mixture waveform that contains multiple speakers speaking different languages. The human auditory sy…
Multi-Stage Face-Voice Association Learning with Keynote Speaker Diarization
Ruijie Tao, Zhan Shi, Yidi Jiang +4
The human brain has the capability to associate the unknown person's voice and face by leveraging their general relationship, referred to as ``cross-modal speaker verification''. T…
Target Speech Diarization with Multimodal Prompts
Yidi Jiang, Ruijie Tao, Zhengyang Chen +2
Traditional speaker diarization seeks to detect ``who spoke when'' according to speaker characteristics. Extending to target speech diarization, we detect ``when target event occur…
Hierarchical Emotion Prediction and Control in Text-to-Speech Synthesis
Sho Inoue, Kun Zhou, Shuai Wang +1
It remains a challenge to effectively control the emotion rendering in text-to-speech (TTS) synthesis. Prior studies have primarily focused on learning a global prosodic representa…
Prompt-driven Target Speech Diarization
Yidi Jiang, Zhengyang Chen, Ruijie Tao +3
We introduce a novel task named `target speech diarization', which seeks to determine `when target event occurred' within an audio signal. We devise a neural architecture called Pr…