most citedVec-Tok Speech: speech vectorization and tokenization for neural speech generation

1 citations · 1 across the 5 of their papers we have counts for

collaborators

6 papers

eess.AS2024

Single-Codec: Single-Codebook Speech Codec towards High-Performance Speech Generation

Hanzhao Li, Liumeng Xue, Haohan Guo +6

The multi-codebook speech codec enables the application of large language models (LLM) in TTS but bottlenecks efficiency and robustness due to multi-sequence prediction. To avoid t…

cs.SD2023

Accent-VITS:accent transfer for end-to-end TTS

Linhan Ma, Yongmao Zhang, Xinfa Zhu +4

Accent transfer aims to transfer an accent from a source speaker to synthetic speech in the target speaker's voice. The main challenge is how to effectively disentangle speaker tim…

cs.SD2023

SponTTS: modeling and transferring spontaneous style for TTS

Hanzhao Li, Xinfa Zhu, Liumeng Xue +3

Spontaneous speaking style exhibits notable differences from other speaking styles due to various spontaneous phenomena (e.g., filled pauses, prolongation) and substantial prosody…

cs.SD20231 cited

Vec-Tok Speech: speech vectorization and tokenization for neural speech generation

Xinfa Zhu, Yuanjun Lv, Yi Lei +5

Language models (LMs) have recently flourished in natural language processing and computer vision, generating high-fidelity texts or images in various tasks. In contrast, the curre…

cs.SD2023

PromptSpeaker: Speaker Generation Based on Text Descriptions

Yongmao Zhang, Guanghou Liu, Yi Lei +4

Recently, text-guided content generation has received extensive attention. In this work, we explore the possibility of text description-based speaker generation, i.e., using text p…

cs.SD2023

Zero-Shot Emotion Transfer For Cross-Lingual Speech Synthesis

Yuke Li, Xinfa Zhu, Yi Lei +4

Zero-shot emotion transfer in cross-lingual speech synthesis aims to transfer emotion from an arbitrary speech reference in the source language to the synthetic speech in the targe…