34 citations · 62 across the 10 of their papers we have counts for
10 papers
Towards Realistic Visual Dubbing with Heterogeneous Sources
Tianyi Xie, Liucheng Liao, Cheng Bi +7
The task of few-shot visual dubbing focuses on synchronizing the lip movements with arbitrary speech input for any talking head video. Albeit moderate improvements in current appro…
Towards Using Clothes Style Transfer for Scenario-aware Person Video Generation
Jingning Xu, Benlai Tang, Mingjie Wang +4
Clothes style transfer for person video generation is a challenging task, due to drastic variations of intra-person appearance and video scenarios. To tackle this problem, most rec…
Cross-speaker Emotion Transfer Based on Speaker Condition Layer Normalization and Semi-Supervised Training in Text-To-Speech
Pengfei Wu, Junjie Pan, Chenchang Xu +4
In expressive speech synthesis, there are high requirements for emotion interpretation. However, it is time-consuming to acquire emotional audio corpus for arbitrary speakers due t…
Towards High-fidelity Singing Voice Conversion with Acoustic Reference and Contrastive Predictive Coding
Chao Wang, Zhonghao Li, Benlai Tang +4
Recently, phonetic posteriorgrams (PPGs) based methods have been quite popular in non-parallel singing voice conversion systems. However, due to the lack of acoustic information in…
PPG-based singing voice conversion with adversarial representation learning
Zhonghao Li, Benlai Tang, Xiang Yin +4
Singing voice conversion (SVC) aims to convert the voice of one singer to that of other singers while keeping the singing content and melody. On top of recent voice conversion work…
Xiaomingbot: A Multilingual Robot News Reporter
Runxin Xu, Jun Cao, Mingxuan Wang +10
This paper proposes the building of Xiaomingbot, an intelligent, multilingual and multimodal software robot equipped with four integral capabilities: news generation, news translat…