activity
20182022
most citedTowards Realistic Visual Dubbing with Heterogeneous Sources

34 citations · 37 across the 5 of their papers we have counts for

collaborators

7 papers

cs.CV202234 cited

Towards Realistic Visual Dubbing with Heterogeneous Sources

Tianyi Xie, Liucheng Liao, Cheng Bi +7

The task of few-shot visual dubbing focuses on synchronizing the lip movements with arbitrary speech input for any talking head video. Albeit moderate improvements in current appro…

cs.CV2021

Towards Using Clothes Style Transfer for Scenario-aware Person Video Generation

Jingning Xu, Benlai Tang, Mingjie Wang +4

Clothes style transfer for person video generation is a challenging task, due to drastic variations of intra-person appearance and video scenarios. To tackle this problem, most rec…

cs.SD20212 cited

Towards High-fidelity Singing Voice Conversion with Acoustic Reference and Contrastive Predictive Coding

Chao Wang, Zhonghao Li, Benlai Tang +4

Recently, phonetic posteriorgrams (PPGs) based methods have been quite popular in non-parallel singing voice conversion systems. However, due to the lack of acoustic information in…

cs.SD20201 cited

PPG-based singing voice conversion with adversarial representation learning

Zhonghao Li, Benlai Tang, Xiang Yin +4

Singing voice conversion (SVC) aims to convert the voice of one singer to that of other singers while keeping the singing content and melody. On top of recent voice conversion work…

cs.CL2020

Improving Accent Conversion with Reference Encoder and End-To-End Text-To-Speech

Wenjie Li, Benlai Tang, Xiang Yin +6

Accent conversion (AC) transforms a non-native speaker's accent into a native accent while maintaining the speaker's voice timbre. In this paper, we propose approaches to improving…

eess.AS2020

ByteSing: A Chinese Singing Voice Synthesis System Using Duration Allocated Encoder-Decoder Acoustic Models and WaveRNN Vocoders

Yu Gu, Xiang Yin, Yonghui Rao +6

This paper presents ByteSing, a Chinese singing voice synthesis (SVS) system based on duration allocated Tacotron-like acoustic models and WaveRNN neural vocoders. Different from t…