activity
20212023
most citedText-to-Video: a Two-stage Framework for Zero-shot Identity-agnostic Talking-head Generation

2 citations · 2 across the 6 of their papers we have counts for

collaborators

6 papers

cs.SD2023

U-Style: Cascading U-nets with Multi-level Speaker and Style Modeling for Zero-Shot Voice Cloning

Tao Li, Zhichao Wang, Xinfa Zhu +4

Zero-shot speaker cloning aims to synthesize speech for any target speaker unseen during TTS system building, given only a single speech reference of the speaker at hand. Although…

eess.AS2023

VITS-Based Singing Voice Conversion Leveraging Whisper and multi-scale F0 Modeling

Ziqian Ning, Yuepeng Jiang, Zhichao Wang +2

This paper introduces the T23 team's system submitted to the Singing Voice Conversion Challenge 2023. Following the recognition-synthesis framework, our singing conversion model is…

eess.AS2023

MSM-VC: High-fidelity Source Style Transfer for Non-Parallel Voice Conversion by Multi-scale Style Modeling

Zhichao Wang, Xinsheng Wang, Qicong Xie +4

In addition to conveying the linguistic content from source speech to converted speech, maintaining the speaking style of source speech also plays an important role in the voice co…

cs.CV20232 cited

Text-to-Video: a Two-stage Framework for Zero-shot Identity-agnostic Talking-head Generation

Zhichao Wang, Mengyu Dai, Keld Lundgaard

The advent of ChatGPT has introduced innovative methods for information gathering and analysis. However, the information provided by ChatGPT is limited to text, and the visualizati…

cs.SD2022

Cross-speaker Emotion Transfer Based On Prosody Compensation for End-to-End Speech Synthesis

Tao Li, Xinsheng Wang, Qicong Xie +3

Cross-speaker emotion transfer speech synthesis aims to synthesize emotional speech for a target speaker by transferring the emotion from reference speech recorded by another (sour…

eess.AS2021

Multi-speaker Multi-style Text-to-speech Synthesis With Single-speaker Single-style Training Data Scenarios

Qicong Xie, Tao Li, Xinsheng Wang +4

In the existing cross-speaker style transfer task, a source speaker with multi-style recordings is necessary to provide the style for a target speaker. However, it is hard for one…