62 citations · 305 across the 60 of their papers we have counts for
79 papers
Open Source MagicData-RAMC: A Rich Annotated Mandarin Conversational(RAMC) Speech Dataset
Zehui Yang, Yifan Chen, Lei Luo +9
This paper introduces a high-quality rich annotated Mandarin conversational (RAMC) speech dataset called MagicData-RAMC. The MagicData-RAMC corpus contains 180 hours of conversatio…
An Audio-Visual Attention Based Multimodal Network for Fake Talking Face Videos Detection
Ganglai Wang, Peng Zhang, Lei Xie +3
DeepFake based digital facial forgery is threatening the public media security, especially when lip manipulation has been used in talking face generation, the difficulty of fake vi…
Attention-Based Lip Audio-Visual Synthesis for Talking Face Generation in the Wild
Ganglai Wang, Peng Zhang, Lei Xie +2
Talking face generation with great practical significance has attracted more attention in recent audio-visual studies. How to achieve accurate lip synchronization is a long-standin…
Audio-visual speech separation based on joint feature representation with cross-modal attention
Junwen Xiong, Peng Zhang, Lei Xie +3
Multi-modal based speech separation has exhibited a specific advantage on isolating the target character in multi-talker noisy environments. Unfortunately, most of current separati…
Learn2Sing 2.0: Diffusion and Mutual Information-Based Target Speaker SVS by Learning from Singing Teacher
Heyang Xue, Xinsheng Wang, Yongmao Zhang +3
Building a high-quality singing corpus for a person who is not good at singing is non-trivial, thus making it challenging to create a singing voice synthesizer for this person. Lea…
Summary On The ICASSP 2022 Multi-Channel Multi-Party Meeting Transcription Grand Challenge
Fan Yu, Shiliang Zhang, Pengcheng Guo +13
The ICASSP 2022 Multi-channel Multi-party Meeting Transcription Grand Challenge (M2MeT) focuses on one of the most valuable and the most challenging scenarios of speech technologie…