activity
20222024
most citedPromptStyle: Controllable Style Transfer for Text-to-Speech with Natural Language Descriptions

2 citations · 3 across the 6 of their papers we have counts for

collaborators
Showing cs.SDShow all

5 papers · 1 filter

cs.SD20231 cited

VITS-based Singing Voice Conversion System with DSPGAN post-processing for SVCC2023

Yiquan Zhou, Meng Chen, Yi Lei +2

This paper presents the T02 team's system for the Singing Voice Conversion Challenge 2023 (SVCC2023). Our system entails a VITS-based SVC model, incorporating three modules: a feat…

cs.SD2023

PromptSpeaker: Speaker Generation Based on Text Descriptions

Yongmao Zhang, Guanghou Liu, Yi Lei +4

Recently, text-guided content generation has received extensive attention. In this work, we explore the possibility of text description-based speaker generation, i.e., using text p…

cs.SD2023

Zero-Shot Emotion Transfer For Cross-Lingual Speech Synthesis

Yuke Li, Xinfa Zhu, Yi Lei +4

Zero-shot emotion transfer in cross-lingual speech synthesis aims to transfer emotion from an arbitrary speech reference in the source language to the synthetic speech in the targe…

cs.SD20232 cited

PromptStyle: Controllable Style Transfer for Text-to-Speech with Natural Language Descriptions

Guanghou Liu, Yongmao Zhang, Yi Lei +4

Style transfer TTS has shown impressive performance in recent years. However, style control is often restricted to systems built on expressive speech recordings with discrete style…

cs.SD2022

Glow-WaveGAN 2: High-quality Zero-shot Text-to-speech Synthesis and Any-to-any Voice Conversion

Yi Lei, Shan Yang, Jian Cong +2

The zero-shot scenario for speech generation aims at synthesizing a novel unseen voice with only one utterance of the target speaker. Although the challenges of adapting new voices…