activity
20172022
most citedDCCRN: Deep Complex Convolution Recurrent Network for Phase-Aware Speech Enhancement

62 citations · 305 across the 60 of their papers we have counts for

collaborators

79 papers

cs.CL2022

Open Source MagicData-RAMC: A Rich Annotated Mandarin Conversational(RAMC) Speech Dataset

Zehui Yang, Yifan Chen, Lei Luo +9

This paper introduces a high-quality rich annotated Mandarin conversational (RAMC) speech dataset called MagicData-RAMC. The MagicData-RAMC corpus contains 180 hours of conversatio…

cs.CV20224 cited

An Audio-Visual Attention Based Multimodal Network for Fake Talking Face Videos Detection

Ganglai Wang, Peng Zhang, Lei Xie +3

DeepFake based digital facial forgery is threatening the public media security, especially when lip manipulation has been used in talking face generation, the difficulty of fake vi…

cs.CV20229 cited

Attention-Based Lip Audio-Visual Synthesis for Talking Face Generation in the Wild

Ganglai Wang, Peng Zhang, Lei Xie +2

Talking face generation with great practical significance has attracted more attention in recent audio-visual studies. How to achieve accurate lip synchronization is a long-standin…

cs.SD20222 cited

Audio-visual speech separation based on joint feature representation with cross-modal attention

Junwen Xiong, Peng Zhang, Lei Xie +3

Multi-modal based speech separation has exhibited a specific advantage on isolating the target character in multi-talker noisy environments. Unfortunately, most of current separati…

cs.SD2022

Learn2Sing 2.0: Diffusion and Mutual Information-Based Target Speaker SVS by Learning from Singing Teacher

Heyang Xue, Xinsheng Wang, Yongmao Zhang +3

Building a high-quality singing corpus for a person who is not good at singing is non-trivial, thus making it challenging to create a singing voice synthesizer for this person. Lea…

cs.SD2022

Summary On The ICASSP 2022 Multi-Channel Multi-Party Meeting Transcription Grand Challenge

Fan Yu, Shiliang Zhang, Pengcheng Guo +13

The ICASSP 2022 Multi-channel Multi-party Meeting Transcription Grand Challenge (M2MeT) focuses on one of the most valuable and the most challenging scenarios of speech technologie…