4 citations · 4 across the 3 of their papers we have counts for
3 papers
cs.CV2024
AniTalker: Animate Vivid and Diverse Talking Faces through Identity-Decoupled Facial Motion Encoding
Tao Liu, Feilong Chen, Shuai Fan +4
The paper introduces AniTalker, an innovative framework designed to generate lifelike talking faces from a single portrait. Unlike existing models that primarily focus on verbal cu…
cs.CL2023★ 4 cited
Improving Code-Switching and Named Entity Recognition in ASR with Speech Editing based Data Augmentation
Zheng Liang, Zheshu Song, Ziyang Ma +3
Recently, end-to-end (E2E) automatic speech recognition (ASR) models have made great strides and exhibit excellent performance in general speech recognition. However, there remain…
eess.AS2022
Unsupervised word-level prosody tagging for controllable speech synthesis
Yiwei Guo, Chenpeng Du, Kai Yu
Although word-level prosody modeling in neural text-to-speech (TTS) has been investigated in recent research for diverse speech synthesis, it is still challenging to control speech…