1 citations · 1 across the 5 of their papers we have counts for
6 papers
Single-Codec: Single-Codebook Speech Codec towards High-Performance Speech Generation
Hanzhao Li, Liumeng Xue, Haohan Guo +6
The multi-codebook speech codec enables the application of large language models (LLM) in TTS but bottlenecks efficiency and robustness due to multi-sequence prediction. To avoid t…
Accent-VITS:accent transfer for end-to-end TTS
Linhan Ma, Yongmao Zhang, Xinfa Zhu +4
Accent transfer aims to transfer an accent from a source speaker to synthetic speech in the target speaker's voice. The main challenge is how to effectively disentangle speaker tim…
SponTTS: modeling and transferring spontaneous style for TTS
Hanzhao Li, Xinfa Zhu, Liumeng Xue +3
Spontaneous speaking style exhibits notable differences from other speaking styles due to various spontaneous phenomena (e.g., filled pauses, prolongation) and substantial prosody…
Vec-Tok Speech: speech vectorization and tokenization for neural speech generation
Xinfa Zhu, Yuanjun Lv, Yi Lei +5
Language models (LMs) have recently flourished in natural language processing and computer vision, generating high-fidelity texts or images in various tasks. In contrast, the curre…
PromptSpeaker: Speaker Generation Based on Text Descriptions
Yongmao Zhang, Guanghou Liu, Yi Lei +4
Recently, text-guided content generation has received extensive attention. In this work, we explore the possibility of text description-based speaker generation, i.e., using text p…
Zero-Shot Emotion Transfer For Cross-Lingual Speech Synthesis
Yuke Li, Xinfa Zhu, Yi Lei +4
Zero-shot emotion transfer in cross-lingual speech synthesis aims to transfer emotion from an arbitrary speech reference in the source language to the synthetic speech in the targe…