activity
20232025
most citedPromptStyle: Controllable Style Transfer for Text-to-Speech with Natural Language Descriptions

2 citations · 3 across the 6 of their papers we have counts for

collaborators
Showing cs.SDShow all

5 papers · 1 filter

cs.SD20251 cited

Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Xinsheng Wang, Mingqi Jiang, Ziyang Ma +22

Recent advancements in large language models (LLMs) have driven significant progress in zero-shot text-to-speech (TTS) synthesis. However, existing foundation models rely on multi-…

cs.SD2024

CoDiff-VC: A Codec-Assisted Diffusion Model for Zero-shot Voice Conversion

Yuke Li, Xinfa Zhu, Hanzhao Li +6

Zero-shot voice conversion (VC) aims to convert the original speaker's timbre to any target speaker while keeping the linguistic content. Current mainstream zero-shot voice convers…

cs.SD2023

SponTTS: modeling and transferring spontaneous style for TTS

Hanzhao Li, Xinfa Zhu, Liumeng Xue +3

Spontaneous speaking style exhibits notable differences from other speaking styles due to various spontaneous phenomena (e.g., filled pauses, prolongation) and substantial prosody…

cs.SD2023

PromptSpeaker: Speaker Generation Based on Text Descriptions

Yongmao Zhang, Guanghou Liu, Yi Lei +4

Recently, text-guided content generation has received extensive attention. In this work, we explore the possibility of text description-based speaker generation, i.e., using text p…

cs.SD20232 cited

PromptStyle: Controllable Style Transfer for Text-to-Speech with Natural Language Descriptions

Guanghou Liu, Yongmao Zhang, Yi Lei +4

Style transfer TTS has shown impressive performance in recent years. However, style control is often restricted to systems built on expressive speech recordings with discrete style…