activity
20212025
most citedMLCA-AVSR: Multi-Layer Cross Attention Fusion based Audio-Visual Speech Recognition

24 citations · 46 across the 48 of their papers we have counts for

collaborators

8 papers

cs.CL2022

The Conversational Short-phrase Speaker Diarization (CSSD) Task: Dataset, Evaluation Metric and Baselines

Gaofeng Cheng, Yifan Chen, Runyan Yang +9

The conversation scenario is one of the most important and most challenging scenarios for speech processing technologies because people in conversation respond to each other in a c…

cs.SD2022

Glow-WaveGAN 2: High-quality Zero-shot Text-to-speech Synthesis and Any-to-any Voice Conversion

Yi Lei, Shan Yang, Jian Cong +2

The zero-shot scenario for speech generation aims at synthesizing a novel unseen voice with only one utterance of the target speaker. Although the challenges of adapting new voices…

cs.SD2022

Minimizing Sequential Confusion Error in Speech Command Recognition

Zhanheng Yang, Hang Lv, Xiong Wang +2

Speech command recognition (SCR) has been commonly used on resource constrained devices to achieve hands-free user experience. However, in real applications, confusion among comman…

cs.SD2022

Cross-speaker Emotion Transfer Based On Prosody Compensation for End-to-End Speech Synthesis

Tao Li, Xinsheng Wang, Qicong Xie +3

Cross-speaker emotion transfer speech synthesis aims to synthesize emotional speech for a target speaker by transferring the emotion from reference speech recorded by another (sour…

eess.AS2022

Leveraging Acoustic Contextual Representation by Audio-textual Cross-modal Learning for Conversational ASR

Kun Wei, Yike Zhang, Sining Sun +2

Leveraging context information is an intuitive idea to improve performance on conversational automatic speech recognition(ASR). Previous works usually adopt recognized hypotheses o…

cs.SD2022

Learning Noise-independent Speech Representation for High-quality Voice Conversion for Noisy Target Speakers

Liumeng Xue, Shan Yang, Na Hu +2

Building a voice conversion system for noisy target speakers, such as users providing noisy samples or Internet found data, is a challenging task since the use of contaminated spee…